MLOps Engineer
Evlo AI
middevopsbackend Atlanta, GA Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
MLOpsKubernetesDockerKubeflowMLflowML MonitoringCI/CDAWSGCPTerraform
About the role
Role Overview
Own the infrastructure, orchestration, and scaling of machine learning systems so models move smoothly from research notebooks to high-availability production services.
Responsibilities
- Design and implement end-to-end MLOps pipelines for automated model training, validation, and deployment.
- Build and operate MLOps tooling using Kubernetes, Docker, Kubeflow, and MLflow.
- Provision and manage scalable cloud infrastructure on AWS or GCP for distributed training and low-latency inference.
- Build reliable data and feature pipelines to ensure low-latency data access and integrity across training and serving.
- Implement model monitoring to detect data drift, concept drift, and performance degradation.
- Optimize inference latency, throughput, and resource usage using quantization, pruning, and hardware acceleration.
- Collaborate with ML engineers and data scientists to set up standardized CI/CD for model artifacts and code.
Requirements
- 3–7 years of experience in MLOps, DevOps, or machine learning engineering, focused on production infrastructure.
- Deep expertise in containerization and orchestration: Docker, Kubernetes, Helm.
- Hands-on experience with ML lifecycle platforms/model registries such as MLflow, Weights & Biases, Kubeflow, or SageMaker.
- Strong proficiency in Python and Bash.
- Infrastructure-as-code experience with Terraform or CloudFormation.
- Solid understanding of CI/CD principles and automated testing for software and ML models.
Bonus
- Experience deploying and scaling LLMs and vector databases in production.
About Evlo AI
Evlo AI is an AI-focused company building machine learning systems that move from research to production. It operates in the machine learning and cloud infrastructure space, with an emphasis on deploying scalable, reliable AI services and supporting high-throughput workloads.
Scraped 7/29/2026