remote
Technical Program Manager, Compute Infrastructure - openai
Product Manager
Seeking a Technical Program Manager for Compute Infrastructure to manage GPU fleets and large-scale compute clusters supporting AI models and training workloads. Focus on platform unification and responsible AI deployment.
About the role
Key Responsibilities
- Manage and optimize large-scale GPU fleets and compute clusters for AI model training and inference.
- Develop and maintain a unified platform for running production Applied AI and Research training workloads.
- Collaborate with engineering teams to ensure the seamless operation of critical compute infrastructure.
- Drive program execution, identify risks, and implement mitigation strategies for compute infrastructure projects.
- Contribute to the responsible and safe deployment of AI technologies.
Requirements
- Proven experience in Technical Program Management, specifically within compute infrastructure or cloud environments.
- Strong understanding of GPU hardware, large-scale distributed systems, and cloud computing principles.
- Experience managing complex technical projects from inception to completion.
- Excellent communication and collaboration skills, with the ability to work effectively with engineering teams.
- Familiarity with AI/ML workloads and their infrastructure requirements is a plus.
Skills
software developmentsystem designproblem solving