Job Description
Position Summary
Vertex is seeking a Sr. Principal AI Engineer, Agentic AI Platform to design, build, and optimize the shared platform capabilities that power AI-enabled products and intelligent workflows across the enterprise. This role will focus on delivering production-grade platform services for model integration, prompt and workflow orchestration, evaluation, observability, performance optimization, and agent lifecycle management.
A key focus of this role will be enabling a build/bring-your-own-agents capability within the Agentic AI Platform, allowing teams across Vertex to create, integrate, customize, and operationalize their own agents using shared platform standards, tooling, and governance controls.
The ideal candidate combines strong software engineering fundamentals with deep experience in applied AI systems. This individual will be comfortable operating across rapid experimentation and engineering rigor, translating emerging AI capabilities into scalable, reliable, secure, and reusable platform components. The Senior Principal AI Engineer will play a critical leadership role in accelerating AI adoption across Vertex by enabling product teams to build and deploy AI solutions faster and more effectively.
Key Responsibilities
- Architect and develop shared AI/agentic platform services that support enterprise AI products and internal workflows
- Design and implement a build/bring-your-own-agents capability that enables teams to create, register, integrate, deploy, and manage their own agents within the enterprise agentic platform
- Establish reusable frameworks, SDKs, templates, interfaces, and guardrails that standardize how custom agents are built and onboarded onto the platform
- Define agent lifecycle capabilities including agent registration, configuration, testing, deployment, monitoring, versioning, and retirement
- Build and maintain robust integrations with foundation models, model gateways, APIs, enterprise tools, and related AI infrastructure
- Design and implement systems for prompt orchestration, workflow execution, tool use, memory patterns, and agentic task coordination
- Develop reusable frameworks and services for evaluation, benchmarking, and validation of AI model, agent, and workflow performance
- Establish platform capabilities for observability, monitoring, tracing, logging, and alerting across AI workloads and autonomous agent interactions
- Optimize platform performance, scalability, latency, reliability, and cost efficiency for production AI and agentic systems
- Partner with product, data, engineering, security, and architecture teams to enable enterprise-ready AI solutions
- Translate prototypes and experimental concepts into hardened, maintainable, production-grade services