remote
Vice President - Site Reliability Engineering - Innovapptive
Software Engineer
Seeking a Vice President of Site Reliability Engineering to lead and scale SRE operations. This role focuses on building robust, scalable, and highly available cloud infrastructure through automation and best practices.
About the role
Key Responsibilities
- Lead and mentor a team of Site Reliability Engineers, fostering a culture of operational excellence.
- Design, implement, and maintain scalable, highly available, and fault-tolerant cloud infrastructure.
- Develop and drive automation strategies for deployment, monitoring, and incident response.
- Establish and enforce best practices for system architecture, performance tuning, and capacity planning.
- Oversee incident management processes, conduct post-mortems, and implement preventative measures.
- Collaborate with engineering teams to ensure reliability and scalability are built into new features and services.
Requirements
- Proven experience in a senior SRE or leadership role, managing complex cloud environments.
- Deep understanding of cloud platforms (e.g., AWS, Azure, GCP) and their services.
- Strong expertise in automation tools and scripting languages (e.g., Python, Bash).
- Experience with CI/CD pipelines, containerization (Docker, Kubernetes), and infrastructure as code.
- Excellent problem-solving, communication, and interpersonal skills.
Skills
pythonbashawsgcpazurekubernetesdocker