Location : Ghent, Belgium/ Bucharest, Romania, hybrid model 2 days on-site
About the position
We’re looking for a Principal Cloud Ops Engineer to help lead and strengthen our cloud infrastructure environment. This role is suited for someone with deep AWS expertise, strong technical judgment, and the ability to quickly understand complex systems. You will play a key role in supporting and evolving cloud infrastructure, improving operational resilience, and helping the team manage critical technical issues. This is a hands-on principal-level role for someone who combines strong infrastructure depth with a collaborative working style.
Key Responsibilities
- Design, implement and improve core AWS-based infrastructure to support accelerated expansion of platform capabilities
- Contribute to infrastructure as code practices, including work with CDK and related tooling
- Diagnose, troubleshoot, and resolve complex infrastructure and platform issues
- Investigate systems deeply and reverse-engineer problems when documentation or context is limited
- Help maintain reliable operational support, including participation in on-call and escalation rotations
- Partner closely with teammates to share context, unblock work, and strengthen team effectiveness
- Ramp quickly on an existing environment and build understanding of product-specific infrastructure over time
- Support knowledge transfer and continuity in a complex operational setup.
Are you our next Cloud Engineer?
Must haves
- AI-Native Engineering: Comfortable using AI agents (e.g., Claude Code, OpenAI Codex/Cursor) daily to accelerate development, refactoring, and infrastructure workflows
- Principal-level AWS Infrastructure Experience: Deep hands-on expertise provisioning, scaling, and managing complex AWS ecosystems
- Strong Infrastructure as Code (IaC): Proven experience leveraging AWS CDK or closely related cloud provisioning frameworks to define and deploy infrastructure cleanly
- Strong Scripting & Coding Skills: Proficient in writing scalable, maintainable code to automate workflows and integrate platform components
- Production Debugging & Troubleshooting: Excellent track record of identifying, diagnosing, and resolving incidents in live production environments
- Autonomy & Rapid Adaptability: Ability to self-direct, pick up unfamiliar systems quickly, and deliver impact with minimal initial ramp-up time
- Communication & Ownership: Clear, collaborative communicator comfortable driving initiatives and participating in on-call/escalation rotation.
Nice to haves
- AI Infrastructure Operations: Opportunity to build, scale, and optimize specialized AI/ML infrastructure and workload pipelines (great if you have experience, but we’re also happy to help you get hands-on here)