This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Engineer, Infrastructure based in Germany.
Join an innovative engineering team building the infrastructure that powers advanced AI-driven products and agent-based workflows. In this role, you will design, operate, and improve highly scalable cloud platforms that enable reliable software delivery and efficient production systems. You will work across Kubernetes, cloud architecture, security, networking, data infrastructure, and developer tooling while solving complex technical challenges in a rapidly evolving environment. This position combines hands-on software engineering with infrastructure ownership, giving you the opportunity to shape platform foundations, improve operational excellence, and directly influence how engineering teams build and deploy products. It is an ideal opportunity for an experienced infrastructure engineer who enjoys automation, solving ambiguous problems, and creating systems that scale.
Accountabilities
- Design, build, and operate reliable infrastructure platforms supporting AI-powered products and large-scale production workloads.
- Own and improve Kubernetes environments, including cluster architecture, networking, ingress, service communication, workload isolation, autoscaling, and deployment practices.
- Develop and maintain cloud infrastructure across areas such as networking, security, identity management, secrets, firewalls, content delivery, and access controls.
- Improve system reliability through enhanced observability, monitoring, alerting, incident response processes, disaster recovery planning, and resilience improvements.
- Build and optimize infrastructure automation using infrastructure-as-code, CI/CD pipelines, GitOps workflows, and reusable platform tooling.
- Investigate and resolve complex production issues across application, infrastructure, networking, and data system layers.
- Improve infrastructure efficiency through cost optimization, capacity planning, resource management, and operational improvements.
- Support scalable data systems and distributed technologies, ensuring strong performance and reliability.
- Identify architectural bottlenecks and proactively develop solutions that prepare the platform for future growth and increasing AI workloads.
- Reduce operational complexity by replacing manual processes with automated, scalable engineering solutions.
Requirements
- Strong software engineering background with experience writing and maintaining production-quality code.
- Proven experience designing, building, and operating infrastructure on major cloud platforms such as GCP, AWS, or similar environments.
- Hands-on production experience with Kubernetes and containerized infrastructure.