xelys jobs xelys jobs

Senior MLOps Engineer

Franklin Fitch

full-remoteseniorpermanentbackenddevops United States 75 days ago via LinkedIn
160,000 - 220,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsPythonDockerKubernetesAWSGCPCI/CDMLflowAirflowKubernetes

About the role

Role Overview

Senior MLOps Engineer for a remote (U.S.) role, owning key parts of a production AI/ML environment. You’ll design scalable, reliable infrastructure and automation that power the end-to-end ML lifecycle—continuous training, deployment, and monitoring—working closely with Data Science and engineering teams.

Responsibilities

  • Build and maintain end-to-end ML pipelines (training, deployment, monitoring)
  • Develop scalable model-serving systems for batch and real-time use cases
  • Implement CI/CD workflows for ML
  • Define standards for observability, reliability, and model governance
  • Automate retraining and model promotion workflows
  • Collaborate cross-functionally to improve platform performance and engineering velocity

Requirements

  • Strong Python engineering background
  • Hands-on experience with Docker and Kubernetes
  • Cloud experience with AWS, GCP, or Azure
  • Experience with ML workflow tools such as:
    • MLflow, Kubeflow, SageMaker, Vertex AI, Airflow, Dagster, Prefect
  • Strong understanding of model deployment, distributed systems, and data pipelines
  • Practical experience building production ML systems

Nice to Have

  • Feature stores or model registries
  • Monitoring/observability tooling
  • Streaming platforms such as Kafka or Kinesis
  • Terraform or other Infrastructure as Code (IaC) tools
  • Experience with LLM/GenAI pipelines

About Franklin Fitch

Franklin Fitch is a technology-focused recruitment partner supporting high-growth engineering teams. They are working with a well-funded company building a production Machine Learning Platform, including scalable infrastructure for training, deployment, and monitoring of ML models.

Scraped 5/12/2026