Become a part of our caring community
Key Responsibilities
Event Correlation & Observability Engineering
- Design, configure, and continuously improve event correlation rules and alerting strategies across platforms such as Splunk ITSI and Dynatrace
- Integrate data from multiple monitoring, application, and infrastructure sources to create meaningful, actionable events
- Normalize and enrich event data using standardized fields and metadata to improve correlation accuracy and reduce noise
- Drive reduction of false positives and duplicate alerts through correlation, aggregation, and suppression strategies
Dashboarding & Data Visualization
- Develop and maintain operational and executive dashboards in Splunk and other reporting tools
- Translate technical telemetry into clear, business-aligned insights , highlighting service health, degradation, and emerging risks
- Partner with command center, TOC, and incident teams to ensure dashboards support real-time decision making and escalation
Incident Detection & Escalation
- Leverage correlated event data and observability insights to trigger proactive incident identification prior to user-reported impact
- Apply criticality tiering and CMDB data to assess business impact and drive proper prioritization and escalation paths
ServiceNow Integration & ITSM Enablement
- Partner with ServiceNow stakeholders to improve workflows, reporting, and automation capabilities
CMDB & Data-Driven Decisioning
- Leverage CMDB relationships and service mapping where available to enrich event data with application, infrastructure, and business context
- Utilize service ownership, business criticality, and operational hours data to inform prioritization decisions
- Partner with CMDB and service mapping teams to improve data quality and completeness
Trend Analysis & Continuous Improvement
- Analyze patterns across incidents, alerts, and events to identify systemic issues and opportunities for improvement
- Partner with Problem Management to eliminate recurring issues through structural fixes
- Drive improvements in monitoring coverage, alert quality, and detection speed
- Contribute to a shift toward predictive, AIOps-driven operations
Use your skills to make an impact
Required Qualifications
- 3–5+ years of experience in Incident, Event, or Problem Management
- Hands-on experience with Splunk (preferably ITSI) and Dynatrace or similar observability platforms
- Experience building dashboards, reports, and analytics to support operational decision-making