Responsibilities
CORE RESPONSIBILITIES
- Define the platform SRE, DevSecOps, and security strategy - spanning architecture, toolchain selection, and
operational standards.
- Design for high availability, scalability, resilience, and disaster recovery across all platform tiers and deployment
regions.
- Establish observability, monitoring, logging, and incident management - own the full telemetry stack from
instrumentation to alerting to post-mortems.
- Govern security architecture, IAM, secrets management, and compliance - ensure the platform meets SOC 2, ISO
27001, and customer-specific security requirements.
- Define CI/CD pipelines, infrastructure automation (IaC), and operational runbooks; drive adoption of GitOps and policy as-code practices.
- Drive performance engineering and capacity planning; own SLOs, SLIs, and error budgets across all platform services.
- Partner with the Technical Architect and FDE leads to embed reliability and security earlier in the development lifecycle
- Lead incident response, blameless post-mortems, and reliability reviews; translate learnings into platform
improvements
Requirements
MUST-HAVE SKILLS & EXPERIENCE
- 12+ years in SRE, Platform Engineering, Cloud Infrastructure, or Cloud Security roles at enterprise or hyperscale scale.
- Expert-level Kubernetes: multi-cluster operations, operator patterns, network policies, pod security, and cluster
hardening.
- Deep cloud-native expertise across AWS, Azure, or GCP: VPC design, IAM, KMS, secrets management (Vault/AWS
Secrets Manager), and cloud security posture management.
- Strong DevSecOps practice: SAST/DAST tooling, container image scanning, SBOM, supply chain security, and secure
CI/CD pipeline design.
- Observability stack mastery: Prometheus, Grafana, OpenTelemetry, distributed tracing (Jaeger/Tempo), log
aggregation (ELK/Loki), and AIOps-ready alerting.
- Infrastructure-as-Code proficiency: Terraform, Pulumi, or CDK; GitOps with ArgoCD or Flux; policy-as-code with
OPA/Kyverno.
- Proven experience defining and operating SLOs, error budgets, and on-call practices in a high-stakes production
environment
Nice-to-have Skills
- Experience operating AI/ML platforms at scale - GPU cluster management, model serving infrastructure (vLLM, Triton),
and LLM observability.
- Compliance and audit experience: SOC 2 Type II, ISO 27001, GDPR, or telecom-specific security frameworks (NIST
Chaos engineering practice: Chaos Monkey, LitmusChaos, or equivalent - running Game Days and failure injection exercises.