remote
Training Infrastructure Engineer - mirelo
Devops Engineer
Seeking a Training Infrastructure Engineer to optimize GPU performance and debug training pipelines for generative AI models. Focus on the full training stack using Python, AWS, and Kubernetes.
About the role
Key Responsibilities
- Optimize and profile GPU behavior for large-scale AI model training.
- Debug and enhance complex training pipelines to ensure efficiency and stability.
- Develop and maintain infrastructure for distributed AI model training.
- Collaborate with ML engineers to implement and test new training methodologies.
- Ensure robust monitoring and logging for the training infrastructure.
- Contribute to the development of MLOps best practices for training.
Requirements
- Proven experience with Python for infrastructure and tooling.
- Strong understanding of machine learning training pipelines and their optimization.
- Experience with cloud platforms, particularly AWS.
- Familiarity with containerization technologies like Docker and orchestration with Kubernetes.
- Experience with GPU profiling and performance tuning is highly desirable.
Skills
pythonmachine learningawskubernetesdocker