xelys jobs xelys jobs

Senior MLOps Engineer

Franklin Fitch

full-remoteseniorpermanentbackenddevops United States 126 days ago via LinkedIn
160,000 - 220,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsPythonKubernetesDockerAWSGCPAzureCI/CDMLflowModel Monitoring

About the role

Role Overview

Senior MLOps Engineer (Remote, U.S.) for a production AI/ML platform team. You’ll be a senior technical contributor responsible for designing and building scalable, reliable infrastructure and automation that power the end-to-end ML lifecycle.

Responsibilities

  • Build and maintain end-to-end ML pipelines (training, deployment, monitoring)
  • Develop scalable model-serving systems for batch and real-time use cases
  • Implement ML CI/CD workflows
  • Set standards for observability, reliability, and model governance
  • Automate retraining and model promotion workflows
  • Collaborate with Data Science, Platform Engineering, and Software Engineering to improve platform performance and engineering velocity

Requirements

  • Strong Python engineering background
  • Hands-on experience with Docker and Kubernetes
  • Cloud experience with AWS, GCP, or Azure
  • Experience with ML workflow tools such as (any of): MLflow, Kubeflow, SageMaker, Vertex, Airflow, Dagster, Prefect
  • Strong understanding of model deployment, distributed systems, and data pipelines
  • Practical experience building production ML systems

Nice to Have

  • Familiarity with feature stores or model registries
  • Experience with monitoring/observability tooling
  • Exposure to streaming platforms such as Kafka or Kinesis
  • Experience with Terraform or other IaC tools
  • Experience with LLM/GenAI pipelines

About Franklin Fitch

Franklin Fitch is a recruiting and staffing firm partnering with a high-growth technology company. The client builds and operates a production Machine Learning Platform, focusing on scalable infrastructure for continuous training, deployment, and monitoring of ML models.

Scraped 5/20/2026