MLOps Engineer
Scale.jobs
midpermanentdevopsdata Miami, FL 2 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
MLOpsPythonKubernetesKubeflowMLflowCI/CDTerraformDockerAWS SageMakerTriton Inference Server
About the role
Role Overview
Own the infrastructure and deployment frameworks that make machine learning models operational at scale. You’ll build highly automated CI/CD and continuous training (CT) pipelines to ensure models are reliably served and monitored in production.
Responsibilities
- Design and maintain scalable ML infrastructure and deployment pipelines using Kubernetes, Kubeflow, and/or MLflow for automated serving and monitoring.
- Build and maintain feature stores and automated ETL/ELT data pipelines using dbt, Apache Spark, and/or Feast.
- Develop real-time and batch model serving systems using Docker, FastAPI, and Triton Inference Server with a focus on low latency and high throughput.
- Implement monitoring for model performance, data drift, and system health using Prometheus and Grafana (plus ML monitoring tools).
- Set up automated testing, versioning, and CI/CD for models and infrastructure code to enable safe, repeatable deployments.
- Collaborate with security/compliance to enforce data governance, access controls, and secure model execution across cloud environments.
Requirements
- 3–6 years of software engineering or data engineering experience, including at least 2 years in MLOps/production ML infrastructure.
- Advanced Python skills; strong shell scripting experience.
- Experience with Infrastructure as Code (Terraform) and containerization (Docker, Kubernetes).
- Hands-on ML pipeline experience on AWS (SageMaker, EKS, S3) or GCP (Vertex AI, GKE).
- CI/CD experience with GitHub Actions and/or GitLab CI.
- MLOps platform experience with MLflow, Kubeflow, Weights & Biases, and/or Prefect.
- Bachelor’s or Master’s degree in a relevant quantitative field.
Bonus
- Experience with Triton Inference Server, production fine-tuning of LLMs, or vector databases (Pinecone, Milvus).
Scraped 7/26/2026