Lead, Site Reliability Engineering Application Support - OMERS
Software Engineer
We're looking for a Software Engineer focused on delivering exceptional customer experiences. This lead role requires 5+ years of relevant experience.
About the role
Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments.
Respond to incidents, lead triage activities, and work closely with SRE, Platform, Network, Security, and application teams to restore service and resolve issues.
Support deployments, release activities, change management, and CI/CD pipelines, including GitHub Actions workflows.
Check, troubleshoot, and resolve user and developer access issues, including Azure AD groups, SSO, application permissions, and firewall rules.
Configure and support platform components such as Azure Container Apps, App Registrations, Key Vault, DNS, certificates, networking, and shared cloud services.
Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, Log Analytics, and related observability tools.
Support onboarding of new applications and teams to the DEV platform by helping with setup, access, deployment readiness, monitoring, and operational handover.
Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation to improve support effectiveness and knowledge sharing.
Contribute to automation and continuous improvement initiatives that reduce manual effort, improve reliability, and strengthen operational processes.
Provide technical guidance to team members and stakeholders while promoting Site Reliability Engineering and platform support best practices.
5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support.
Strong hands-on experience with Microsoft Azure services, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management (APIM), and Azure Functions.
Experience supporting production environments, including incident response, troubleshooting, problem management, and operational support processes.
Experience with container technologies and cloud-native application architectures.
Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms.
Strong understanding of identity, networking, and access management concepts, including SSO, OAuth, application registrations, and security groups.
Experience with observability and monitoring platforms such as Datadog, Azure Monitor, and Log Analytics.
Understanding of cloud networking concepts, including DNS, certificates, firewalls, private endpoints, and network security controls.
Experience with scripting and automation using technologies such as PowerShell, Bash, Azure CLI, Python, or similar tools.
Strong knowledge of operating systems and cloud infrastructure concepts.