Role Reasonability
MetLife is seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of critical applications and platforms.
The SRE Engineer will monitor production systems, respond to incidents, improve observability, maintain runbooks, and automate operational tasks. Working closely with engineering, cloud, and infrastructure teams, the role supports SRE practices, operational readiness, and service reliability through SLIs, SLOs, and Error Budget management.
Core Responsibilities
- Monitoring: Monitor service health, dashboards, alerts, and key reliability indicators for assigned applications and platforms.
- Incident Response: Respond to alerts, support bridge calls, gather evidence, execute runbooks, communicate status, and escalate when required.
- Observability Support: Create and maintain dashboards, log queries, telemetry checks, alert validation, and actionable monitoring signals.
- Runbook Management: Document operational procedures, update recovery steps, validate readiness with service owners, and support knowledge sharing.
- Automation: Create scripts for repetitive checks, data collection, remediation, operational reporting, and toil reduction.
- Problem Follow-up: Support root cause analysis, postmortem documentation, and closure of assigned corrective/preventive action items.
- Continuous Improvement: Identify alert noise, toil, monitoring gaps, and preventive improvements for senior SRE review.
- SRE Alignment: Support adoption of SLOs, SLIs, SLAs, error budgets, operational readiness reviews, and production support standards.
- AI Readiness: Use or help improve AI-assisted tools for anomaly detection, incident correlation, root cause hints, and operational knowledge retrieval.
- Collaboration: Work with engineering, infrastructure, cloud, and application teams to align service performance with business goals.
Skills and Experience
- Foundations: Linux, networking fundamentals, application support, cloud fundamentals, production operations, and ITIL-style incident/change processes.
- Scripting: Python, PowerShell, Bash, or equivalent scripting for automation, diagnostics, evidence collection, and reporting.
- Tools: Git, CI/CD basics, ServiceNow or equivalent ticketing; exposure to Elastic/ELK, Grafana, Prometheus, Splunk, APM, and Azure Monitor preferred.
- Cloud & Containers: Azure services, Docker, Kubernetes, and hybrid cloud operations exposure; Terraform or infrastructure-as-code awareness preferred.
- Reliability: Basic understanding of SLIs, SLOs, SLAs, error budgets, alerting, incident response, postmortems, and operational runbooks.