OVERVIEW
Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking engineers to build platform infrastructure at the intersection of computational science, AI systems, and software engineering.
Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.
The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era.
We have one product: new physics, at scale.
ROLE AND
RESPONSIBILITIES
- Design and implement new runtime primitives for our AI platform. Each runtime encodes a programming model that researchers and engineers compose into agentic workflows for physics discovery. Example shapes include sequential pipelines and tree-search agents; we expect to add more as the science demands new patterns.
- Build and harden the multi-tenant durable workflow execution system that powers AI-driven physics research at scale: correctness under retries and replays, isolation between tenants, recovery from partial failures, and predictable behavior under load.
- Treat our AI platform as a library product. Design the programmatic interfaces that researchers and engineers across PSI extend, with clear architectural layers and explicit API contracts so that scientific workflows compose cleanly.
- Operate the platform that runs every research workflow and customer-facing AI product at PSI: define and meet SLOs, build instrumentation and alerting, plan capacity, lead incident response.
WHAT WE'RE LOOKING FOR
- Four or more years building and operating distributed systems in production at companies known for engineering rigor (e.g., Google, Netflix, Meta, Cloudflare, Datadog, or comparable), on major cloud platforms (GCP, AWS, or Azure) with Kubernetes or comparable container orchestration. You have written code that paying customers, internal teams, or large user bases depend on every day, and you are fluent in the operational realities of cloud-native infrastructure.
- A track record of designing and shipping a Python library or internal framework that other engineers extend, not just consume. You think about API ergonomics, type-driven contracts, composability, and backward-compatible evolution as first-order concerns.
- Real experience implementing or substantially extending orchestration primitives, workflow engines, dataflow systems, or agent runtimes. You understand the subtle bugs that come from retries, replays, and non-deterministic execution.