AI Engineer
Senior MLOps / AI Platform Engineer Role Overview Aivar Innovations is looking for a Senior MLOps / AI Platform Engineer to help design and build an enterprise-grade MLOps and AIOps software platform running on Kubernetes.
Senior MLOps / AI Platform Engineer Role
Overview
Aivar Innovations is looking for a Senior MLOps / AI Platform Engineer to help design and build an enterprise-grade MLOps and AIOps software platform running on Kubernetes. You will be responsible for developing the infrastructure and platform capabilities required to deploy, manage, scale, and observe machine learning, deep learning, and generative AI workloads across cloud and on-premises environments. You will work extensively with Kubernetes, GPUs, model-serving frameworks, distributed systems, and cloud-native technologies. You will collaborate closely with Product, Engineering, AI/ML, DevOps, and Customer Delivery teams to transform complex AI infrastructure requirements into reliable, secure, and easy-to-use platform capabilities. This role requires strong hands-on engineering experience and a practical understanding of how machine learning models move from experimentation into reliable production environments.
Requirements
5–8 years of experience in software engineering, platform engineering, DevOps, SRE, MLOps, or related infrastructure roles. Strong hands-on experience with Kubernetes, including writing Kubernetes Operators and Custom Resource Definitions using frameworks such as Kubebuilder, Operator SDK, or equivalent. Experience designing and operating cloud-native infrastructure on AWS, particularly Amazon EKS, EC2, S3, ECR, IAM, VPC, and CloudWatch. Experience deploying and operating machine learning, deep learning, or generative AI models in production. Experience running and troubleshooting GPU-accelerated workloads on Kubernetes, with an understanding of GPU scheduling, utilization, memory constraints, and performance. Familiarity with model-serving frameworks such as KServe, NVIDIA Triton Inference Server, vLLM, Ray Serve, TorchServe, or equivalent technologies. Strong programming experience in Go or Python, with experience building production-grade APIs, controllers, or distributed backend services. Experience with containers, Helm, CI/CD, infrastructure as code, and observability tools such as Prometheus, OpenTelemetry, and Grafana. Strong understanding of Linux, networking, storage, security, and distributedsystem fundamentals. Strong debugging, problem-solving, communication, and cross-functional collaboration skills.
Preferred Qualifications
Posted July 24, 2026