Senior AI Platform Engineer

name · via Himalayas ·

TypeFull-time job
LocationUnited States
Posted1 hour ago
Our client is a technology consulting company providing operational and engineering services to the high-tech sector. They support platform and infrastructure teams across cloud and distributed environments, helping operate scalable, reliable, and production-ready technology platforms.
The Role
Our client is looking for a Senior AI Platform Engineer to operate, maintain, and continuously improve production AI platforms running on Kubernetes across on-premise, AWS, and GCP environments.
The role combines AI platform engineering, MLOps, Kubernetes, Python, observability, and production operations, working with environments similar to AI on EKS and Kubeflow-based machine learning platforms.
You will also play a senior role in improving engineering practices, mentoring team members, and helping shape the platform roadmap.
Key Responsibilities

Deploy platform releases and configuration changes using GitOps and DevOps practices.

Monitor AI platform and service health through logs, metrics, monitoring, and observability tools.

Improve platform reliability through automation, operational tooling, observability, and self-service capabilities.

Participate in incident response, root cause analysis, and 24/7 operational rotations.

Investigate and resolve user, platform, integration, and configuration-related issues.

Promote strong standards across platform security, reliability, and operational engineering.

Mentor junior engineers in Python fundamentals and help develop their MLOps capabilities.

Drive the adoption of MLOps best practices across the engineering team.

Identify gaps in tooling, technical capabilities, and processes required to support production-grade AI systems.

Contribute to the technical direction and ongoing development of the AI platform.

Qualifications

3+ years of experience supporting production AI, ML, or data platforms using technologies such as Ray, Jupyter, AWS SageMaker, Kubeflow, or similar platforms.

5+ years of experience across the AI/ML lifecycle, including development, deployment, DevOps, or MLOps.

5+ years of hands-on Python experience supporting AI/ML workflows, applications, or data engineering pipelines.

Strong practical experience with Kubernetes, including managed platforms such as AWS EKS or Google GKE.

Good understanding of microservices architectures and service communication patterns.

Strong troubleshooting skills across application crashes, resource contention, service latency, performance, and scaling issues.

Experience analysing logs, metrics, monitoring systems, and service-level KPIs within production environments.

Nice to Have

Exposure to additional AI and data platforms such as Flyte, Hugging Face, Vertex AI, LangChain, Claude Code, or other AI agent platforms.

Hands-on automation or scripting experience using Bash or Python.

Relevant Kubernetes or cloud certifications such as CKAD or AWS certifications.

Originally posted on Himalayas
ai-platform-engineering mlops cloud-platform-engineering site-reliability-engineering devops senior-ai-platform-engineer senior-ai-platform-developer senior-ml-platform-engineer senior-ai-infrastructure-engineer senior-ai-software-engineer
Apply on Himalayas →

Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.