Overview
We are seeking a Lead Software Engineer with strong expertise in DevOps, Site Reliability Engineering (SRE), Cloud Platforms, Kubernetes, and GitOps practices. The ideal candidate will design, operate, and scale production Kubernetes platforms while driving reliability, automation, observability, and operational excellence across enterprise cloud-native applications.
This role will partner closely with Engineering, Product, Infrastructure, and Security teams to build resilient, Kubernetes-based cloud-native solutions and champion modern DevOps practices using ArgoCD, GitOps, and container orchestration best practices.
What You'll Do
DevOps & Platform Engineering
- Design, implement, and optimize CI/CD pipelines and Kubernetes-based deployment strategies.
- Build and maintain scalable, secure, and highly available cloud infrastructure.
- Drive infrastructure automation, configuration management, and platform standardization.
- Improve developer productivity through platform engineering and self-service capabilities.
Site Reliability Engineering (SRE )
- Establish and maintain reliability standards, SLAs, SLOs, and operational best practices.
- Lead incident management, root cause analysis, and reliability improvement initiatives.
- Enhance system performance, scalability, availability, and disaster recovery capabilities.
- Implement proactive monitoring, alerting, and observability solutions.
GitOps & Automation
- Implement and manage GitOps practices using ArgoCD for Kubernetes application delivery.
- Automate application deployments, environment provisioning, and configuration management.
- Ensure consistent, secure, and auditable deployments across all environments.
- Promote DevOps and GitOps best practices across engineering teams.
Kubernetes & Cloud Infrastructure
- Design, deploy, and manage production Kubernetes clusters across cloud environments (EKS, AKS, GKE, or equivalent).
- Build and optimize containerized application deployments using Docker and Kubernetes.
- Implement Kubernetes-native tooling including Helm, Kustomize, and operators for application lifecycle management.
- Configure and troubleshoot Kubernetes networking, ingress, storage, autoscaling, and resource management.
- Drive cluster reliability, upgrade strategies, capacity planning, and performance optimization.
- Partner with Security teams to implement Kubernetes security best practices (RBAC, network policies, pod security standards).
Collaboration & Leadership
- Partner with Engineering, Product, Infrastructure, and Security teams to deliver reliable solutions.
- Mentor engineers and provide technical leadership on DevOps and SRE best practices.
- Collaborate with global teams across EMEA and US regions.