MLOps / LLMOps Engineer (GenAI Platform)
Location: Santa Clara, CA Client: Applied Materials (AMAT) Experience: 5–7 Years
Job Overview
We are seeking an experienced MLOps / LLMOps Engineer to design, deploy, and optimize production-grade Generative AI and Large Language Model (LLM) platforms. The ideal candidate will have strong expertise in Python, AI/ML platform engineering, model serving, Kubernetes, and cloud-native MLOps practices.
Required Skills
- 5–7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering, or Machine Learning Engineering .
- Strong proficiency in Python and software engineering best practices.
- Hands-on experience with open-source LLMs such as Llama, Mistral, Gemma, or Qwen .
- Expertise in LLM inference and model hosting using technologies such as:
- vLLM
- SGLang
- Hugging Face TGI
- NVIDIA Triton Inference Server
- Ray Serve
- Azure Machine Learning
- Databricks Model Serving
- Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow .
- Strong understanding of:
- Retrieval-Augmented Generation (RAG)
- Vector Databases
- GPU Optimization
- Model Quantization
- KV Cache
- PagedAttention
- Continuous/Dynamic Batching
- Proven experience building, deploying, troubleshooting, scaling, and optimizing production-grade GenAI and LLM applications .
- Experience implementing AI observability, governance, and Responsible AI best practices.
Preferred Qualifications
- Hands-on experience with LLM Fine-Tuning using:
- PEFT
- SFT
- CPT
- LoRA
- QLoRA
- Experience with:
- Azure AI Foundry
- Azure OpenAI
- Hugging Face
- DeepSpeed
- PEFT
- Knowledge of distributed training and multi-GPU environments.
- Experience with Agentic AI frameworks such as:
- LangGraph
- AutoGen
- CrewAI
- Familiarity with simulation platforms, digital twins, scientific computing, or modeling and simulation workflows.
What You'll Do
- Design, build, and maintain scalable AI/ML infrastructure for enterprise LLM applications.
- Deploy and optimize LLM inference workloads for high performance and low latency.
- Implement scalable model serving, monitoring, and observability solutions.
- Collaborate with AI researchers, data scientists, and software engineers to deliver production-ready GenAI solutions.
- Improve GPU utilization, model performance, and operational efficiency.
- Ensure AI governance, security, and Responsible AI compliance across deployments.
Why Join?
- Work on cutting-edge Generative AI and Large Language Model technologies.
- Build enterprise-scale AI platforms using modern cloud-native tools.