xelys jobs xelys jobs

MLOps Engineer

Bright Vision Technologies

full-remoteseniorpermanentbackenddevopsdata Tempe, AZ Today via LinkedIn
100,000 - 150,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

PythonGoRustC++LLM InferencevLLMTensorRT-LLMKubernetesObservabilityDistributed Systems

About the role

Role Overview

MLOps Engineer (100% Remote, U.S.) to design, build, and operate high-performance, reliable ML inference platforms for serving large machine learning models in production.

Responsibilities

  • Design and operate model serving platforms for diverse workloads (e.g., LLMs, vision models, recommendation systems)
  • Optimize inference performance using techniques such as:
    • continuous batching, paged attention, speculative decoding, request multiplexing
  • Implement multi-tenant routing, rate limiting, and QoS policies across model endpoints
  • Build autoscaling and capacity management balancing latency, throughput, and cost
  • Tune GPU utilization, memory management, and KV cache strategies for LLM workloads
  • Integrate serving with API gateways, identity systems, and observability platforms
  • Implement caching, prompt deduplication, and response reuse strategies
  • Deliver end-to-end observability (latency histograms, queue dynamics, GPU utilization, error tracking)
  • Create deployment workflows: canary releases, shadow testing, automated rollback
  • Operate incident response and drive durable reliability improvements
  • Collaborate with ML and product teams to support model releases
  • Add serving-layer security controls (request signing, content filtering, abuse detection)
  • Document operational procedures and performance/tuning guidance
  • Translate AI serving research advances into production improvements

Requirements

  • BS or MS in Computer Science (or related)
  • 6+ years experience in distributed systems, infrastructure, or ML platform engineering
  • Strong Python and a systems language such as Go, Rust, or C++
  • Proven experience running high-throughput, low-latency production services
  • Hands-on LLM inference experience with frameworks like vLLM or TensorRT-LLM
  • Deep understanding of GPU architecture and accelerator utilization
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms
  • Experience with observability stacks (metrics, tracing, structured logging)
  • Strong performance engineering and capacity planning background
  • Communication and incident response skills

Preferred Qualifications

  • Open-source contributions to model serving infrastructure
  • Multi-region/global distributed AI serving experience
  • Model quantization/distillation/compression exposure
  • FinOps for AI workloads / cost-efficient serving design
  • Experience supporting external-facing AI APIs at scale

About Bright Vision Technologies

Bright Vision Technologies is a technology consulting and software development company. It delivers cloud, AI, data, and enterprise solutions across the United States.

Scraped 7/25/2026