xelys jobs xelys jobs

MLOps Engineer

Evlo AI

midpermanentbackenddevops New York, NY Today via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsCI/CDKubernetesTerraformAWSFeature StoresMLflowML MonitoringTritonLLM Serving

About the role

Role Overview

Own the infrastructure, orchestration, and deployment pipelines that move machine learning models from experimental notebooks to resilient, low-latency production services.

Responsibilities

  • Architect, build, and maintain ML-focused CI/CD pipelines using tools such as GitHub Actions, Argo Workflows, or Kubeflow Deploy.
  • Deploy and operate production ML models, scaling and monitoring with Kubernetes, Docker, and serving engines like NVIDIA Triton or TorchServe.
  • Implement automated monitoring for data drift, concept drift, and system performance (latency, throughput).
  • Optimize infrastructure cost and utilization for GPU and high-compute workloads via auto-scaling.
  • Establish security, reproducibility, and lineage standards for training data, feature stores, and model artifacts.
  • Integrate ML models into backend services via gRPC and REST APIs.

Requirements

  • 3–6 years experience in MLOps, DevOps, or machine learning engineering focused on production infrastructure.
  • Strong proficiency in containerization, Kubernetes, and Infrastructure as Code (Terraform or CloudFormation).
  • Hands-on experience with feature stores (Feast, Tecton) and model registries (MLflow, Weights & Biases).
  • Strong software engineering fundamentals in Python and Bash for scalable backend services.
  • Solid understanding of cloud infrastructure on AWS, GCP, or Azure, including networking, IAM, and GPU management.

Bonus

  • Experience deploying and serving large language models (LLMs) using vLLM, TGI, or TensorRT-LLM.

About Evlo AI

Evlo AI is an AI-focused company building and deploying machine learning models into production systems. The role emphasizes reliable, low-latency ML infrastructure, orchestration, and deployment pipelines, indicating an industry focused on AI/ML engineering and cloud-native production services.

Scraped 8/6/2026