ABOUT THE TEAM AND ROLE:
At Talon.One, we're building the infrastructure behind modern promotions, loyalty, and gamification for some of the world's most forward-thinking brands.
As our engineering organization scales, we are introducing a new Technical Governance and Enablement layer that sits right between Product Engineering, Architecture, and Platform/Infra. We are looking for a Technical Program Manager to become the operational engine, stakeholder shield, and planning partner for this critical new function.
Production Engineering defines how production is run across R&D - from incidents and releases to operational workflows, reliability initiatives, and the systems that help engineers operate safely at scale. We treat this domain as an internal product.
This is not a customer-facing role. It is about bringing structure, focus, and predictability to a highly technical team, turning ambitious reliability and architectural strategies into realistic, data-driven execution plans. You will remove the coordination and operational overhead for engineering leaders so they can focus on scaling and mentoring the team.
ONCE YOU ARE HERE YOU WILL:
- Work hand-in-hand with leadership to design, document, and roll out foundational engineering programs from scratch (e.g., modernizing our incident management framework, on-call training, blameless post-mortem culture, and release governance)
- Standardize incoming channels (Slack, Shortcut, Incident Follow-ups, GTM feedback etc.) to manage ad-hoc chaos, triage operational requests, and protect the focus of SREs and TechOps so they can deliver deep project work. Establish a clear operational intake process ("front door") for the team
- Drive quarterly planning, backlog refinement, sprint readiness, and execution tracking for reliability and Ops initiatives. Own the Production Engineering backlog and continuously improve incident response and release processes
- Coordinate work across Engineering, Platform, Architecture, Product, Support, and Post-sales, managing dependencies and surfacing delivery risks early
- Configure and optimize project management tools (like Jira/Trello) to build clean tracking dashboards and automate intake workflows
- Design the communication loops, tracking dashboards, and engineering reviews that give senior stakeholders clear insight into operational health, incident follow-ups, and tech-debt remediation
- Evolve the metrics (SLOs/SLIs, MTTR, deployment velocity, tech-debt ratios) that guide how we balance shipping new business features with maintaining absolute system reliability
- Enable Architects and engineering leadership with the operational data and planning insights needed for effective capacity planning and prioritization
WHAT WE NEED YOU TO BRING TO THE TABLE:
- Direct experience defining or revamping incident management, post-incident review practices, or release gove