onsite
VPII, Head of Site Reliability Engineering - lplfinancial
Software Engineer
Seeking a VPII, Head of Site Reliability Engineering to lead SRE initiatives. Expertise in cloud infrastructure, DevOps, and system architecture is crucial for ensuring high availability and performance.
About the role
Key Responsibilities
- Lead and mentor a team of Site Reliability Engineers.
- Develop and implement strategies for system reliability, scalability, and performance.
- Oversee the design, implementation, and maintenance of cloud infrastructure.
- Drive adoption of DevOps best practices and automation across engineering teams.
- Establish and monitor key performance indicators (KPIs) for system health and availability.
- Manage and resolve production incidents, conducting post-mortems to prevent recurrence.
Requirements
- Proven experience in a Site Reliability Engineering leadership role.
- Deep understanding of cloud platforms (e.g., AWS, Azure, GCP) and infrastructure as code.
- Strong knowledge of CI/CD pipelines, containerization (Docker, Kubernetes), and automation tools.
- Experience with performance monitoring, logging, and alerting systems.
- Excellent problem-solving and incident management skills.
Skills
awsgcpazurekubernetesdocker