xelys jobs xelys jobs

MLOps Engineer

Scale.jobs

middevopsdata Chicago, IL 82 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsPythonKubernetesDockerTerraformMLflowFeature StoresPrometheusGrafanaCI/CD

About the role

Role Overview

You will bridge the gap between machine learning development and production operations by designing, building, and maintaining automated infrastructure and pipelines for deploying, monitoring, and scaling machine learning models and large language models across the enterprise.

Responsibilities

  • Design and implement automated ML CI/CD pipelines using GitLab CI, GitHub Actions, or Jenkins to deploy models to Kubernetes clusters
  • Develop and maintain central model registries and feature stores for standardized asset management, versioning, and feature reuse
  • Build real-time monitoring and alerting for model latency, throughput, data drift, and concept drift using Prometheus, Grafana, or similar observability tools
  • Optimize inference serving using Triton Inference Server, TorchServe, or vLLM to reduce latency and cloud costs
  • Collaborate with data engineers on data validation pipelines to ensure data quality before training and evaluation
  • Implement automated containerization, Infrastructure as Code (e.g., Terraform), and orchestration strategies for reproducible ML environments in multi-tenant cloud setups

Requirements

  • 3–6 years of experience in software engineering, MLOps, or DevOps, focused on deploying ML workloads to production cloud environments
  • Proficiency in Python and shell scripting
  • Hands-on experience with Docker, Kubernetes, and Helm charts
  • Experience with production ML infrastructure such as MLflow, Kubeflow, Feast, or AWS SageMaker pipelines
  • Strong Infrastructure-as-Code practices using Terraform or CloudFormation on AWS, GCP, or Azure
  • Strong cross-functional communication and collaboration with Data Science, Data Engineering, and Infrastructure teams

Nice to Have

  • Experience deploying large language models
  • Knowledge of quantization techniques (e.g., AWQ, GPTQ)
  • Experience managing distributed training clusters (e.g., Ray, Spark)

About Scale.jobs

Scale.jobs is an engineering-focused hiring platform that connects companies with technical talent. This role is for an organization seeking to strengthen its MLOps and production ML infrastructure, bridging machine learning development with reliable deployment and monitoring.

Scraped 7/3/2026