Site Reliability Engineer
MX is a fintech company on a mission to empower the world to be financially strong.
MX is a fintech company on a mission to empower the world to be financially strong. We build technology that helps banks, credit unions, and fintechs deliver smarter, more intuitive financial experiences to millions of people.
Like many startups, we’ve navigated real growth challenges — and we’ve come out stronger on the other side. Today, MX is in a phase of renewed momentum and scale, with a solid foundation and a clear vision for what’s next. This is a place where thoughtful execution matters, innovation is encouraged, and individuals have real ownership over their work.
Our culture values curiosity, accountability, and impact. We give people the space to question assumptions, design better solutions, and help shape how the company grows. If you’re looking to do meaningful work, influence outcomes, and grow alongside a company that’s ready to move fast, you’ll feel at home at MX.
At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.
We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.
We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.
This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.
Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.
What you'll do:
Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
Posted July 31, 2026