remote
VP of Site Reliability - Titan Ai
Software Engineer
Seeking a VP of Site Reliability to lead the scaling of AI software deployments for financial institutions. This role requires expertise in cloud infrastructure, DevOps, and ensuring high availability and performance across diverse environments.
About the role
Key Responsibilities
- Lead and scale the Site Reliability Engineering (SRE) function to support rapid growth in customer deployments.
- Develop and implement strategies for ensuring the high availability, performance, and scalability of AI software across various cloud and on-premise environments.
- Oversee the architecture and management of cloud infrastructure, including Kubernetes, to meet the demanding needs of banking clients.
- Establish and maintain robust DevOps practices, CI/CD pipelines, and incident response protocols.
- Collaborate with engineering and product teams to define reliability standards and best practices.
- Ensure compliance with stringent banking industry standards for security, audit, and model risk.
Requirements
- Proven experience in a senior Site Reliability Engineering or DevOps leadership role.
- Deep understanding of cloud platforms (e.g., Azure) and container orchestration (Kubernetes).
- Expertise in building and managing scalable, resilient, and secure infrastructure.
- Strong background in system architecture, performance tuning, and troubleshooting complex distributed systems.
- Excellent communication and leadership skills, with the ability to mentor and grow a team.