Summary
We are seeking a highly skilled and experienced Site Reliability Manager to join our team to ensure the reliability, scalability, and performance of our systems and services. You will lead a team of engineers focusing on three core pillars: Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate must reside in DMV area and be available to work on site in office or customer locations in this area. Must have demonstrated experience in having performed this role for at least 3 years.
What You'll Be Doing:
- Team Leadership: Lead an 8-20 person service delivery team (Service Support specialist, DevSecOps, and Site Reliability engineers), mentoring them to foster a culture of learning and innovation.
- Pillar 1: Application Reliability (Core Focus): Take joint ownership of production reliability, standardize observability and error handling using Datadog, develop SLOs and KPIs to measure system performance, and conduct incident post-mortems and root cause analyses.
- Pillar 2: DevSecOps: Define and implement best practices for infrastructure as code, deployment automation, and drive vulnerability management and resolution.
- Pillar 3: Platform Lifecycle Management: Oversee the end-to-end platform lifecycle, collaborate with cross-functional teams to design scalable and fault-tolerant architectures, and drive continuous improvement initiatives to enhance system efficiency.
Required Qualifications:
- Bachelor’s degree in Computer Science, Engineering, or a related field; Master's degree preferred.
- 10+ years of experience in a similar role managing a team of site reliability engineers and delivering in the AWS cloud platform.
- 5+ years of experience supporting operations and maintenance for cloud-native applications in production that are fault-tolerant, self-healing, scalable, and highly available.
- Deep understanding of the AWS cloud computing platform and containerization technologies (e.g., Docker, Kubernetes).
- Strong knowledge of infrastructure as code tools (e.g., Terraform, Ansible, ArgoCD) and CI/CD pipelines.
- Experience with Datadog as the primary logging, monitoring, and observability platform (alongside AWS Cloudwatch).
- Excellent communication and interpersonal skills, with the ability to collaborate effectively with cross-functional teams.
- Strong problem-solving and analytical skills, with a keen attention to detail.
- Ability to obtain and maintain a Public Trust clearance.
- Certifications such as AWS Certified DevOps Engineer are a plus.