remote
Weekend Site Reliability Engineer - Sporty Group
Site Reliability Engineer
We're looking for a Site Reliability Engineer focused on designing and building scalable technical solutions. This mid level role requires 3+ years of relevant experience.
About the role
What you’ll be doing
- Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
- Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
- Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
- Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
- Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
- Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
- Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
- Take ownership and responsibility for our cloud operation activities
- Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
- Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
- Mentoring less experienced team members
What you’ll bring
- 3+ years DevOps / platform engineering experience
- Must be based in Europe or Asia or LatAM
- Experience independently leading the planning and deployment of a project
- Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
- Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
- Experience with Infrastructure-as-Code, particularly Terraform
- Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
- Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
- Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
- Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
- Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
- Experience defining SLIs and SLOs and using them to inform reliability work