Overview
- We are seeking a highly skilled Staff Software Engineer with deep expertise in DevOps, Site Reliability Engineering (SRE), Cloud Platforms, Kubernetes, and GitOps practices. This role will architect, scale, and operate enterprise-grade Kubernetes platforms while driving reliability, scalability, automation, and operational excellence across critical enterprise applications and infrastructure.
- The ideal candidate will provide technical leadership on Kubernetes platform strategy, influence engineering best practices, and partner with cross-functional teams to build resilient, secure, and highly available cloud-native platforms at scale. This position requires deep hands-on experience with Kubernetes architecture, ArgoCD, GitOps methodologies, Infrastructure Automation, and Production Engineering.
What You'll Do
DevOps & Platform Engineering
- Design, build, and optimize scalable CI/CD pipelines and Kubernetes-native platform solutions.
- Drive Infrastructure as Code (IaC), automation, and platform standardization initiatives.
- Improve developer experience through self-service infrastructure and deployment automation.
- Lead architecture discussions and establish Kubernetes and DevOps platform standards across engineering teams.
Site Reliability Engineering (SRE)
- Define and champion reliability standards, SLAs, SLOs, and operational excellence frameworks.
- Lead incident response, root cause analysis, and reliability improvement programs.
- Drive performance optimization, scalability enhancements, capacity planning, and disaster recovery strategies.
- Build proactive monitoring, observability, and alerting capabilities to improve system health and availability.
GitOps & Automation
- Architect and manage GitOps practices using ArgoCD for multi-cluster Kubernetes application delivery.
- Automate application deployment, configuration management, and environment provisioning.
- Establish deployment governance, release management processes, and operational controls.
- Ensure secure, consistent, and auditable deployments across all environments.
Kubernetes & Cloud Infrastructure
- Define and evolve enterprise Kubernetes platform architecture, standards, and multi-cluster/multi-region deployment strategies.
- Design, deploy, and operate production Kubernetes clusters at scale across cloud environments (EKS, AKS, GKE, or equivalent).
- Build and optimize containerized platform solutions using Docker and Kubernetes for high availability and performance.
- Lead adoption of Kubernetes-native tooling including Helm, Kustomize, operators, and service mesh technologies (Istio, Linkerd, or equivalent).
- Drive cluster lifecycle management including upgrades, autoscaling, capacity planning, disaster recovery, and cost optimization.