We are seeking a skilled and passionate AI/ML Engineer with a strong focus on large language model (LLM) fine-tuning and adaptation. The ideal candidate will bring expertise across the end-to-end ML lifecycle—from data preprocessing and model development to deployment, optimization, and monitoring.
You will play a critical role in shaping and refining advanced AI systems that our clients and medical experts will rely on for accurate, efficient, and trustworthy decision support.
This role requires not only a solid foundation in machine learning and data engineering, but also hands-on experience with LLM fine-tuning, parameter-efficient training methods (LoRA/QLoRA), and model evaluation techniques.
You will collaborate closely with cross-functional teams to deliver domain-specific AI solutions that balance innovation, performance, and reliability in high-stakes environments.
Key Responsibilities
- Data Preparation & Domain Adaptation: Collect, preprocess, and structure complex, unstructured medical and clinical data (e.g., records, reports, clinical notes) for training and fine-tuning, while ensuring compliance with privacy and security standards (HIPAA, GDPR).
- Model Development & LLM Fine-Tuning: Design, adapt, and fine-tune Large Language Models (LLMs) for specialized tasks such as summarization, information extraction, semantic search, and classification. Apply parameter-efficient approaches (LoRA/QLoRA, PEFT) to customize models for domain-specific applications.
- Model Exploration & Benchmarking: Stay ahead of the curve by evaluating cutting-edge open-source LLMs (e.g., Llama, Mistral, Gemma). Experiment with novel architectures, fine-tuning methods, and inference optimizations to identify the most effective solutions for real-world use cases.
- MLOps & Scalable Deployment: Build and maintain robust MLOps pipelines for training, deployment, and scaling of models. Use containerization, orchestration (Docker/Kubernetes), and CI/CD workflows to ensure reproducibility, reliability, and fast iteration.
- Performance Monitoring & Continuous Improvement: Establish monitoring frameworks to track model accuracy, drift, and robustness. Work with domain experts to validate outputs and drive iterative improvements based on feedback.
- Responsible & Ethical AI: Ensure fairness, transparency, and explainability in all deployed AI systems. Conduct bias audits and document decision processes to uphold ethical best practices.
- Cross-Functional Collaboration: Partner with engineers, product managers, and domain experts to translate requirements into impactful AI solutions. Communicate complex AI concepts clearly to both technical and non-technical stakeholders.
Preferred Skills & Experience
- Bachelor’s or master’s degree in computer science, Data Science, or a related quantitative field.
- Proven experience as a Machine Learning Engineer, Data Scientist, or similar role, with hands-on experience in building and deploying production-grade ML systems.
- Strong programming skills in Python, with proficiency in frameworks such as PyTorch or TensorFlow.
- Experience in LLM fine-tuning and generative AI models, including parameter-efficient approaches (LoRA, QLoR), instruction tuning, supervised fine-tuning (SFT), and reinforcement learning methods (RLHF, DPO, GRPO).
- Hands-on experience with LLM frameworks and libraries: Hugging Face Transformers, DeepSpeed, bitsandbytes, vLLM.
- Familiarity with retrieval-augmented generation (RAG), LangChain/LlamaIndex, and vector databases.
- Skilled in model evaluation and benchmarking, including task-specific metrics, human evaluation, and bias/fairness testing.
- Knowledge of distributed training, GPU optimization, and quantization (4-bit, 8-bit) for efficient model adaptation.
- Hands-on experience with MLOps practices and tools for deployment (e.g., vLLM, Ray Serve), monitoring (e.g., Evidently AI, Prometheus/Grafana), experiment tracking and model versioning (e.g., MLflow, Weights & Biases), and workflow orchestration (e.g., Apache Airflow).
- Familiarity with cloud platforms. Oracle Cloud Infrastructure (OCI) experience is a strong plus.
- Strong problem-solving skills and ability to collaborate with subject matter experts to translate complex domain knowledge into machine learning solutions.
- Understanding of data privacy and security principles, especially when working with sensitive or regulated data (e.g., PHI).