Machine Learning Engineer
Evlo AI
midpermanentbackenddata Seattle, WA Yesterday via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine LearningPythonPyTorchApache SparkDockerKubernetesAWSONNX RuntimeCI/CDFeature Store
About the role
Role Overview
Own the end-to-end lifecycle of production machine learning systems—from scalable training pipelines to deploying low-latency models that support core product features.
Responsibilities
- Architect and implement distributed machine learning pipelines using Python, PyTorch, and Apache Spark.
- Deploy, monitor, and scale production models using cloud infrastructure such as AWS, Docker, and Kubernetes.
- Optimize inference latency, throughput, and memory via quantization, pruning, and ONNX Runtime.
- Build automated monitoring to detect feature drift, data quality issues, and performance degradation in real time.
- Collaborate with data engineering teams to define feature stores and ensure training/inference data consistency.
- Write maintainable code, perform thorough peer code reviews, and contribute to system architecture documentation.
Requirements
- 3–6 years of professional software engineering; at least 3 years focused on machine learning engineering.
- Strong Python skills and deep hands-on experience with production-grade ML frameworks (PyTorch or TensorFlow).
- Experience deploying and maintaining containerized ML models in cloud environments (AWS, GCP, or Azure).
- Solid software engineering practices: CI/CD, automated testing, and Infrastructure-as-Code.
- BS or MS in Computer Science, Machine Learning, Statistics, or related technical field.
Nice to Have
- Experience with LLM fine-tuning and/or RAG architectures.
- Experience contributing to major open-source ML projects.
About Evlo AI
Evlo AI is an AI-focused company building production machine learning systems to power core product features. The role emphasizes end-to-end ML lifecycle ownership, including scalable training, deployment, monitoring, and performance/cost optimization in production.
Scraped 7/28/2026