Staff Engineer, Foundation Model Serving
As a Staff Engineer, you will be crucial in designing and building scalable, reliable, and high-performance systems for Databricks' Foundation Model Serving API Product. This role involves working with cutting-edge AI models, optimizing GPU workloads for inference, and collaborating cross-functionally to deliver a world-class product experience.
At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business.
Foundation Model Serving is the API Product for hosting and serving frontier AI model inference for open source models like Llama, Qwen, and GPT OSS as well as proprietary models like Claude and OpenAI GPT. For this role, no prior ML or AI experience is necessary. We’re looking for engineers who have owned high scale operational sensitive systems like customer facing APIs, Edge Gateways, ML Inference, or similar services and have an interest in getting deep building LLM APIs and runtimes at scale.
As a Staff Engineer, you’ll play a critical role in shaping both the product experience and core infrastructure. You will design and build systems that enable high-throughput, low-latency inference on GPU workloads with frontier models, influence architectural direction, and collaborate closely across platform, product, infrastructure, and research teams to deliver a world-class foundation model API product.
Posted May 26, 2026