Description
At Rocket.net, reliability, performance, and customer experience are at the center of everything we build. We are looking for a Site Reliability Engineer to help maintain the health, stability, and performance of our hosting platform while providing advanced technical support to our customers.
The Platform Operations team acts as a critical escalation layer between WordPress Support and Engineering. This role combines infrastructure operations, platform monitoring, troubleshooting, and advanced customer support.
As a Site Reliability Engineer, you will help ensure Rocket.net's servers, services, and customer environments are operating at the highest standards. You will assist WordPress Support Engineers with complex issues, support VIP customers with advanced technical requests, investigate platform-level problems, and work with internal teams to deliver fast and effective solutions.
Responsibilities
Platform Monitoring & Reliability
- Monitor the health, availability, and performance of Rocket.net servers, services, and customer environments.
- Proactively identify infrastructure issues, performance degradation, and potential service disruptions.
- Investigate alerts and operational events to maintain platform stability.
- Perform regular platform health checks and ensure critical systems are operating correctly.
- Participate in incident response and coordinate troubleshooting during customer-impacting events.
- Communicate platform issues, updates, and resolutions to relevant internal teams.
Advanced Technical Support & Escalations
- Provide advanced technical support for VIP customers and customers with complex hosting-related issues.
- Act as a senior escalation point for WordPress Support Engineers when issues require deeper technical investigation.
- Troubleshoot complex issues involving servers, websites, networking, DNS, performance, caching, and hosting infrastructure.
- Assist customers with advanced technical problems beyond standard WordPress troubleshooting.
- Investigate and resolve issues involving server resources, application performance, connectivity, and platform behavior.
- Work directly with customers when required to provide expert-level technical assistance.
- Ensure escalated customer issues are handled with urgency, ownership, and clear communication.
Infrastructure Operations
- Troubleshoot and maintain Linux-based production environments.
- Investigate issues related to NGINX, Apache, PHP-FPM, MySQL/MariaDB, Redis, and other platform services.
- Assist with server maintenance, configuration changes, and operational improvements.
- Support security updates, system hardening, and infrastructure best practices.
- Monitor resource usage and identify capacity or performance concerns.