remote
Senior Software Engineer, Fleet-level ML Performance - Google
Software Engineer
We're looking for a Software Engineer focused on designing and building scalable technical solutions. This senior role requires 5+ years of relevant experience.
About the role
Minimum qualifications:
- Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
- 5 years of experience in systems architecture or computer architecture, power and performance trade-off analysis, or data center, cloud, infrastructure hardware optimization .
- Experience with Reliability, Availability, and Serviceability (RAS) features, paradigms, or architecture.
Preferred qualifications:
- Master's degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with an emphasis on computer architecture.
- Knowledge of deep learning workloads, including embedding architectures and their hardware execution characteristics.
Responsibilities
- Perform fleet-level performance analysis of key ML workloads (e.g., Gemini) on future TPU systems using advanced simulation tools to evaluate hardware/software trade-offs and guide next-generation chip architecture.
- Partner with teams across the ML stack including model researchers, compiler developers, systems engineers, and TPU architects to analyze and optimize performance across the design space.
- Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C++ and Python to enable scalable performance and TCO projections across Google.
- Collaborate cross-functionally with data center, hardware architecture, and framework teams to define key hardware and software requirements for future AI infrastructure.
Skills
machine learningdeep learningpythoncelectrical engineering