Wells Fargo is seeking a Senior Lead Platform Reliability Engineer to join the CTO Platform organization. This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database, or Storage) and have demonstrated experience collaborating across at least one additional infrastructure domains (Enterprise Tools, Cloud, Observability Tools). The expectation is that this engineer will elevate themselves in looking for trends and patterns that are not limited to these streams and investigate systemic issues that span multiple streams or need a deeper troubleshooting.
As part of our Platform Reliability Engineering (PRE) team, you will apply modern Site Reliability Engineering (SRE) practices to improve the availability, resiliency, observability, scalability, and operational excellence of critical enterprise platforms. You will leverage your domain expertise to identify systemic issues, drive automation, and deliver engineering solutions that strengthen platform stability at scale.
In This Role You Will
- Serve as the reliability engineering expert for your primary domain (Network, Middleware, Database, or Storage) while partnering across adjacent technology disciplines
- Lead the investigation and resolution of complex production incidents, identifying root causes and implementing long-term corrective actions
- Apply SRE principles including service level indicators (SLIs), service level objectives (SLOs), error budgets, and reliability engineering practices to improve platform health
- Lead capacity analysis, forecasting, and utilization reviews to identify future scaling risks and prevent service degradation before customer impact occurs
- Perform deep performance analysis across infrastructure layers, identifying bottlenecks, contention points, latency drivers, and resource inefficiencies
- Identify and remediate configuration drift, operational debt, and platform hygiene issues that impact long-term reliability
- Drive proactive reliability improvements through observability, automation, performance optimization, and resiliency engineering
- Design and implement automation solutions that eliminate operational toil, reduce manual intervention, and improve recovery capabilities
- Define and enhance enterprise observability standards through metrics, logging, tracing, alerting, and service health monitoring
- Partner closely with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availability
- Lead blameless post-incident reviews and convert recurring operational issues into measurable engineering improvements
- Identify reliability risks and communicate technical recommendations to engineering leaders and senior stakeholders