xelys jobs xelys jobs

MLOps Engineer

Bright Vision Technologies

full-remoteseniorpermanentbackenddevops Gilbert, AZ 11 days ago via LinkedIn
100,000 - 150,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

MLOpsModel ServingInference PlatformsDistributed SystemsPerformance EngineeringLLM ServingAutoscalingGPU UtilizationObservabilityCI/CD

About the role

Role Overview

MLOps Engineer to design, build, and operate high-performance, highly reliable ML inference platforms for production-scale large models. You’ll focus on the systems engineering side of AI deployment, including routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability.

Responsibilities

  • Design and operate model serving platforms for diverse workloads (e.g., LLMs, vision models, recommendation systems)
  • Improve inference performance using techniques such as:
    • Continuous batching
    • Paged attention
    • Speculative decoding
    • Request multiplexing
  • Implement multi-tenant routing, rate limiting, and QoS policies across endpoints
  • Build autoscaling and capacity management systems balancing latency/throughput/cost
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving
  • Integrate serving with API gateways, identity systems, and observability platforms
  • Apply caching and optimization strategies such as prompt deduplication and response reuse
  • Deliver end-to-end observability (latency histograms, queue dynamics, GPU utilization, error tracking)
  • Build deployment workflows: canary releases, shadow testing, automated rollback
  • Operate incident response for high-availability AI services; drive lasting reliability improvements
  • Partner with ML and product teams for new model releases and rollouts
  • Implement serving-layer security controls (e.g., request signing, content filtering, abuse detection)
  • Document operational procedures and performance/tuning guidance
  • Stay current with AI serving research and translate advances into production

Requirements

  • 6+ years of experience (Bachelor’s or Master’s in Computer Science or related field)
  • Hands-on experience shipping ML serving/inference systems at scale
  • Strong distributed systems and performance engineering background
  • Understanding of trade-offs between latency, throughput, cost, and quality in ML serving

Additional Notes

  • 100% remote (Continental United States)
  • Full-time W2 in-house SOW engagement; no third-party clients.
  • No C2C/1099 arrangements.
  • Technical coding assessment mandatory.
  • No new H1B sponsorship available (H1B transfers for qualified candidates supported).

About Bright Vision Technologies

Bright Vision Technologies is a software development company focused on building innovative solutions that automate and optimize business operations. It leverages cutting-edge technologies to create scalable, secure, and user-friendly applications, and is expanding its engineering team to support production AI initiatives.

Scraped 7/15/2026