xelys jobs xelys jobs

MLOps Engineer

Bright Vision Technologies

full-remoteseniorpermanentbackenddevops Scottsdale, AZ Today via LinkedIn
100,000 - 150,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsLLM InferenceDistributed SystemsPythonGoRustC++KubernetesGPU ArchitectureObservability

About the role

Role Overview

Design, build, and operate high-performance, reliable ML inference (MLOps) platforms for serving large machine learning models in production. You’ll focus on the systems engineering aspects of AI deployment—scaling, routing, batching, caching, GPU efficiency, and end-to-end observability.

Key Responsibilities

  • Build and operate model serving platforms for diverse workloads (LLMs, vision models, recommendation systems)
  • Improve inference performance using techniques like:
    • continuous batching
    • paged attention
    • speculative decoding
    • request multiplexing
  • Implement multi-tenant routing, rate limiting, and quality-of-service (QoS) policies
  • Create autoscaling and capacity management systems balancing latency, throughput, and cost
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving
  • Integrate model serving with API gateways, identity systems, and observability platforms
  • Add caching, prompt deduplication, and response reuse where applicable
  • Own end-to-end observability (latency histograms, queue dynamics, GPU utilization, error tracking)
  • Implement deployment workflows such as canary releases, shadow testing, and automated rollback
  • Lead incident response and drive reliability improvements for high-availability AI services
  • Partner with ML and product teams to support model releases and rollout of new capabilities
  • Add serving-layer security controls (request signing, content filtering, abuse detection)
  • Document operational procedures, performance characteristics, and tuning guidance
  • Stay current with AI serving research and translate advances into production

Required Qualifications

  • Bachelor’s or Master’s in Computer Science (or related field)
  • 6+ years experience in distributed systems, infrastructure, or ML platform engineering
  • Strong Python skills plus a systems language such as Go, Rust, or C++
  • Proven production experience with high-throughput, low-latency services
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms
  • Experience with observability stacks (metrics, tracing, structured logging)
  • Solid performance engineering and capacity planning background
  • Strong communication and incident response skills

Preferred Qualifications

  • Open-source contributions to model serving infrastructure
  • Multi-region / globally distributed AI serving experience
  • Familiarity with model quantization, distillation, and compression
  • Exposure to FinOps for AI workloads and cost-efficient serving
  • Experience supporting external-facing AI APIs at scale

About Bright Vision Technologies

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. The role is positioned as a full-time, remote opportunity within an established organization focused on AI and cloud delivery.

Scraped 7/28/2026