Research Engineer, Model Evaluation
As a Senior Research Engineer, Model Evaluation, you will develop next-generation evaluation methods and scalable infrastructure to measure LLM progress. This includes creating benchmarks, datasets, and environments, conducting research on LLM evaluation techniques, and building tools for analyzing model performance. The role is critical for advancing the state of the art in frontier model capabilities.
Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.
We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers.
Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must continue to develop new techniques to accurately measure our models' performance on frontier capabilities. In this role, you are responsible for creating next-generation evaluation methods and scalable infrastructure to measure LLM progress.
Posted June 9, 2026