About the Role
Our client is seeking a Technical Lead, Machine Learning to lead the execution of its AI platform by translating research into scalable, production-ready machine learning systems. This role sits at the intersection of research, infrastructure, and product, with responsibility for ensuring models are trainable, deployable, observable, and optimized for real-world performance.
Working closely with research, engineering, and product teams, this position will drive the development of robust ML infrastructure while balancing performance, reliability, latency, and cost.
Key Responsibilities
- Lead the end-to-end execution of machine learning systems, including data pipelines, training workflows, evaluation frameworks, inference architecture, and production deployment.
- Fine-tune and optimize models using modern techniques such as LoRA, QLoRA, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and model distillation .
- Design, build, and operate scalable inference systems with a focus on latency, cost efficiency, and reliability.
- Develop and maintain data pipelines for both synthetic and real-world training datasets.
- Build evaluation frameworks to measure model performance, robustness, safety, and bias in collaboration with research teams.
- Optimize production deployments through GPU utilization, memory efficiency, inference optimization, and scaling strategies.
- Partner closely with application engineering teams to integrate machine learning systems into backend, desktop, and mobile products.
- Continuously improve production systems through rapid iteration, monitoring, and data-driven optimization.
Requirements
- Proven experience building and deploying production-grade machine learning systems used by real users.
- Strong expertise working with large language models and understanding model behavior, limitations, and failure modes.
- Experience developing scalable ML infrastructure, training pipelines, and inference systems.
- Strong software engineering skills with the ability to write maintainable, production-quality code.
- Experience balancing real-world production constraints, including latency, reliability, scalability, cost, and safety.
- Strong ownership mindset with the ability to independently drive technical initiatives from design through deployment.
- Excellent communication and collaboration skills, with experience working in cross-functional, high-performing engineering teams.
Preferred Technical Skills
Experience with the following technologies is preferred:
- Python
- PyTorch and/or JAX
- GPU-based model training and inference systems
Originally posted on Himalayas