Machine Learning Engineer
Evlo AI
midpermanentbackenddata Dallas, TX Yesterday via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine LearningGenAILLMsPythonPyTorchRAGLangChainLlamaIndexAWSVector Databases
About the role
Role Overview
Own the end-to-end design, scaling, and operationalization of machine learning and GenAI systems in high-throughput production. You’ll work at the intersection of applied research and scalable infrastructure to improve latency, accuracy, and reliability.
Responsibilities
- Design, train, and deploy production-grade ML models and LLM-based systems using Python and PyTorch
- Build robust data pipelines for feature extraction, data cleaning, and vector embedding generation
- Implement Retrieval-Augmented Generation (RAG) pipelines and orchestrate LLM workflows using LangChain and LlamaIndex
- Monitor deployed models for data drift, concept drift, and performance degradation with automated logging/observability
- Optimize inference latency, throughput, and compute costs using techniques like quantization and pruning plus hardware acceleration
- Write clean, testable, well-documented code and participate in peer code reviews and architecture discussions
Requirements
- 3–6 years of professional experience in software engineering and ML engineering with production deployments
- Strong Python skills and hands-on deep learning experience (PyTorch or TensorFlow)
- Experience with cloud ML infrastructure on AWS, GCP, or Azure
- Solid understanding of vector databases, embedding generation, and semantic search architectures
- Degree in Computer Science, Statistics, Mathematics, or related technical field
Bonus
- Experience fine-tuning open-source LLMs using LoRA/QLoRA
- Contributions to open-source ML projects
About Evlo AI
Evlo AI is building applied AI systems, including machine learning and GenAI solutions, designed for high-throughput production environments. The work combines applied research with scalable infrastructure to deliver reliable, low-latency model performance.
Scraped 8/4/2026