xelys jobs xelys jobs

Senior Platform/MLOps Engineer

Jobgether

seniorpermanentdevopsbackenddata United States 52 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsPlatform EngineeringKubernetesGPUCI/CDGitOpsTerraformAnsiblePrometheusGrafana

About the role

Role Overview

Senior Platform/MLOps Engineer (United States) responsible for building and scaling MLOps and platform infrastructure for AI-driven manufacturing systems. The work focuses on reliable end-to-end movement of models from training to production, supporting computer vision, deep learning, and robotics workloads in factory environments.

Responsibilities

  • Build and evolve production infrastructure for scalable ML and platform systems with a focus on reliability, performance, and developer productivity.
  • Design and maintain end-to-end MLOps pipelines (training, deployment, and monitoring) for computer vision and robotics models.
  • Design, implement, and maintain scalable ML/AI infrastructure, including training pipelines, deployment systems, and inference services.
  • Build and optimize GPU-enabled workloads on Kubernetes.
  • Develop CI/CD and GitOps workflows to support continuous delivery of ML and platform services.
  • Collaborate with cross-functional teams to define architecture, evaluate tradeoffs, and prototype platform capabilities.
  • Improve reliability using observability tooling, incident response practices, and performance optimization.
  • Partner with applied AI and robotics teams to ensure infrastructure meets real-world production needs.
  • Produce documentation and contribute to engineering best practices.

Requirements

  • 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering.
  • Strong programming skills in Python, Go, JavaScript, C#, or similar.
  • Proven experience designing and operating MLOps pipelines in production.
  • Deep knowledge of Kubernetes (including CNCF ecosystem; managed and self-hosted).
  • Hands-on experience running and optimizing GPU workloads in Kubernetes clusters.
  • Experience with Infrastructure as Code (e.g., Terraform) and configuration management (e.g., Ansible).
  • Experience with CI/CD pipelines and GitOps delivery workflows.
  • Familiarity with observability tools such as Prometheus, Grafana, and OpenTelemetry.
  • Strong software engineering practices across the SDLC.
  • Strong communication and cross-team collaboration.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.

Preferred

  • Experience with highly secure or air-gapped environments.
  • Mentoring engineers.
  • Contributing to architectural decisions in complex distributed systems.

About Jobgether

Jobgether lists this role on behalf of a partner company. The partner builds infrastructure for next-generation AI-driven manufacturing systems, focusing on computer vision, deep learning, and robotics deployments in real-world factory environments.

Scraped 6/18/2026