Senior Principal AI Scientist – Foundation Models & Generative AI
Remote (U.S.) | Full-Time | $185,000–$280,000 + Bonus + Equity
We're partnering with an industry-leading AI technology company that has deployed its technology at massive global scale and is building the next generation of foundation models powering intelligent, real-world systems. This is an opportunity to join a world-class research organization where you'll shape core model architecture, drive technical strategy, and solve some of the hardest problems in modern AI.
If you've trained large language models from scratch—not simply fine-tuned existing models—and enjoy solving deep optimization and scaling challenges, we'd love to connect.
What You'll Do
- Design and train large-scale transformer and foundation models from the ground up.
- Own architectural decisions across language, multimodal, and emerging model architectures.
- Lead research into optimization, scaling laws, and training stability for production-scale models.
- Develop novel approaches for model efficiency, convergence, and inference performance.
- Partner closely with ML systems engineers while maintaining ownership of model architecture and research direction.
- Influence the long-term technical roadmap for next-generation generative AI systems.
What You'll Bring
- Extensive experience building and training large-scale transformer models from scratch.
- Deep expertise in modern deep learning, optimization theory, and representation learning.
- Strong understanding of transformer internals, including attention mechanisms, RoPE, ALiBi, GQA, and Mixture-of-Experts (MoE).
- Experience with distributed training frameworks including FSDP, ZeRO, tensor parallelism, and pipeline parallelism.
- Hands-on knowledge of mixed precision training (BF16, FP8), gradient checkpointing, and large-scale GPU clusters.
- Expertise with optimization algorithms including AdamW, Lion, Adafactor, learning-rate scheduling, and debugging training instability.
- Strong understanding of scaling laws, compute vs. data tradeoffs, and efficient model scaling.
- Experience with modern alignment techniques such as RLHF, DPO, or GRPO is highly desirable.
You'll Solve Problems Like
- Preventing training divergence at extreme model scale.
- Designing architectures that improve convergence and generalization.
- Optimizing compute efficiency without sacrificing model quality.
- Building scalable multimodal and hybrid foundation model architectures.
- Advancing state-of-the-art generative AI capabilities for production environments.
Compensation & Benefits
- Base Salary: $185,000–$280,000
- Annual performance bonus
- Equity opportunity
- Comprehensive medical, dental, and vision coverage
- Life and disability insurance