xelys jobs xelys jobs

Machine Learning Engineer

Evlo AI

midpermanentbackenddata Dallas, TX Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Machine LearningGenAILLMsPythonPyTorchRAGLangChainLlamaIndexAWSVector Databases

About the role

Role Overview

Own the end-to-end design, scaling, and operationalization of machine learning and GenAI systems in high-throughput production. You’ll work at the intersection of applied research and scalable infrastructure to improve latency, accuracy, and reliability.

Responsibilities

  • Design, train, and deploy production-grade ML models and LLM-based systems using Python and PyTorch
  • Build robust data pipelines for feature extraction, data cleaning, and vector embedding generation
  • Implement Retrieval-Augmented Generation (RAG) pipelines and orchestrate LLM workflows using LangChain and LlamaIndex
  • Monitor deployed models for data drift, concept drift, and performance degradation with automated logging/observability
  • Optimize inference latency, throughput, and compute costs using techniques like quantization and pruning plus hardware acceleration
  • Write clean, testable, well-documented code and participate in peer code reviews and architecture discussions

Requirements

  • 3–6 years of professional experience in software engineering and ML engineering with production deployments
  • Strong Python skills and hands-on deep learning experience (PyTorch or TensorFlow)
  • Experience with cloud ML infrastructure on AWS, GCP, or Azure
  • Solid understanding of vector databases, embedding generation, and semantic search architectures
  • Degree in Computer Science, Statistics, Mathematics, or related technical field

Bonus

  • Experience fine-tuning open-source LLMs using LoRA/QLoRA
  • Contributions to open-source ML projects

About Evlo AI

Evlo AI is building applied AI systems, including machine learning and GenAI solutions, designed for high-throughput production environments. The work combines applied research with scalable infrastructure to deliver reliable, low-latency model performance.

Scraped 8/4/2026