xelys jobs xelys jobs

Staff Machine Learning Systems Engineer (MLOps)

Jobgether

leadbackenddevopssecurity United States 49 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Machine LearningMLOpsKubernetesAWS EKSTerraformGitOpsCI/CDPythonObservabilityOpenTelemetry

About the role

Role overview

Staff Machine Learning Systems Engineer (MLOps) for a partner company, supporting the design and operation of production ML/AI infrastructure powering large-scale AI services. This is a hands-on senior technical role at the intersection of platform engineering, DevOps/SRE, and applied AI—focused on deploying, observing, securing, and scaling machine learning workloads across cloud-native environments.

Responsibilities

  • Lead the design, evolution, and operation of core ML infrastructure platforms for production AI workloads.
  • Own and optimize Kubernetes-based infrastructure (e.g., AWS EKS), including autoscaling, workload orchestration, and cluster lifecycle management.
  • Build and maintain GitOps-based CI/CD pipelines for safe, repeatable deployments of AI services across environments.
  • Design and implement model serving and inference infrastructure (LLM routing, API gateways, and multi-provider integrations).
  • Develop observability, tracing, and monitoring for AI workloads using tools such as OpenTelemetry and Datadog (including LLM tracing platforms).
  • Define SLOs, incident response processes, and production reliability standards for ML systems.
  • Implement infrastructure-as-code and platform tooling to improve developer velocity and consistency (Terraform, CLIs, internal frameworks).
  • Drive security architecture: IAM, secrets management, compliance, least-privilege access, and data protection.
  • Collaborate with ML, product, and data teams to move prototypes into production systems.
  • Identify and improve platform bottlenecks for performance, cost efficiency, and deployment speed.
  • Provide technical leadership, mentorship, and architectural guidance.

Requirements

  • 8+ years of experience in platform engineering, DevOps, SRE, or infrastructure roles, including hands-on ML/AI systems experience.
  • Deep expertise in Kubernetes (preferably EKS): cluster operations, autoscaling, and workload orchestration.
  • Proficiency with Terraform and experience designing secure cloud architectures.
  • Strong programming skills in Python for infrastructure tooling/automation.
  • Hands-on experience operating LLM/ML inference in production (routing, serving, observability).
  • Experience with observability stacks (e.g., Datadog, OpenTelemetry, logging/tracing systems).
  • Strong understanding of CI/CD, GitOps, and developer platform engineering.
  • Experience designing IAM, OIDC, and secrets management systems.
  • Systems-thinking mindset focused on failure modes, reliability, and maintainability.
  • Ability to collaborate across engineering, ML, security, and product in fast-paced environments.

Nice to have

  • Experience in regulated/high-compliance environments (e.g., healthcare, fintech).

Benefits

  • Competitive salary with equity opportunities.
  • Comprehensive health benefits.

Scraped 6/18/2026