xelys jobs xelys jobs

MLOps Engineer

Evlo AI

midpermanentbackenddevops New York, NY 8 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsCI/CDKubernetesKubeFlowApache AirflowPythonDockerAWSTerraformMLflow

About the role

Role Overview

Own the infrastructure, pipelines, and deployment frameworks that power real-time AI models. You will build robust MLOps platforms to automate the ML lifecycle—training, deployment, monitoring, and scaling—while ensuring reliable, cost-effective, low-latency performance in production cloud environments.

Responsibilities

  • Design, build, and maintain CI/CD pipelines for ML models to automate the path from research to production
  • Deploy and orchestrate ML workflows using Kubernetes, Kubeflow, and/or Apache Airflow (manage complex DAGs and training pipelines)
  • Implement model monitoring, logging, and alerting for data drift, concept drift, and performance degradation in real time
  • Optimize model serving infrastructure using Triton Inference Server, TorchServe, or TensorFlow Serving to minimize latency and infrastructure costs
  • Build and maintain a centralized feature store (e.g., Feast or Tecton) to standardize data definitions across training and serving
  • Work with security and compliance teams on data governance, model lineage, and access controls across the ML lifecycle

Requirements

  • 3–6 years of software engineering/DevOps/data platform experience; at least 2 years in MLOps in production
  • Strong Python and shell scripting skills
  • Experience with Docker and Kubernetes (containerization and orchestration)
  • Hands-on cloud infrastructure experience, preferably AWS or GCP
  • Terraform for Infrastructure as Code (IaC)
  • Familiarity with ML tracking/registry tools such as MLflow and Weights & Biases (or cloud-native model registries)
  • Solid software engineering practices: Git workflows, unit testing, and automated integration testing

Bonus

  • Experience with large-scale distributed training (e.g., Ray, Horovod)
  • Experience deploying LLMs using vLLM or Hugging Face TGI

About Evlo AI

Evlo AI develops AI products and platforms that rely on deploying machine learning models in production environments. The role described focuses on building production-grade MLOps infrastructure, pipelines, and deployment systems that enable scalable, low-latency AI services.

Scraped 7/18/2026