MLOps Engineer
Bright Vision Technologies
full-remoteseniorpermanentbackenddevopsdata Tempe, AZ Today via LinkedIn
100,000 - 150,000 USD/annual
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
PythonGoRustC++LLM InferencevLLMTensorRT-LLMKubernetesObservabilityDistributed Systems
About the role
Role Overview
MLOps Engineer (100% Remote, U.S.) to design, build, and operate high-performance, reliable ML inference platforms for serving large machine learning models in production.
Responsibilities
- Design and operate model serving platforms for diverse workloads (e.g., LLMs, vision models, recommendation systems)
- Optimize inference performance using techniques such as:
- continuous batching, paged attention, speculative decoding, request multiplexing
- Implement multi-tenant routing, rate limiting, and QoS policies across model endpoints
- Build autoscaling and capacity management balancing latency, throughput, and cost
- Tune GPU utilization, memory management, and KV cache strategies for LLM workloads
- Integrate serving with API gateways, identity systems, and observability platforms
- Implement caching, prompt deduplication, and response reuse strategies
- Deliver end-to-end observability (latency histograms, queue dynamics, GPU utilization, error tracking)
- Create deployment workflows: canary releases, shadow testing, automated rollback
- Operate incident response and drive durable reliability improvements
- Collaborate with ML and product teams to support model releases
- Add serving-layer security controls (request signing, content filtering, abuse detection)
- Document operational procedures and performance/tuning guidance
- Translate AI serving research advances into production improvements
Requirements
- BS or MS in Computer Science (or related)
- 6+ years experience in distributed systems, infrastructure, or ML platform engineering
- Strong Python and a systems language such as Go, Rust, or C++
- Proven experience running high-throughput, low-latency production services
- Hands-on LLM inference experience with frameworks like vLLM or TensorRT-LLM
- Deep understanding of GPU architecture and accelerator utilization
- Familiarity with Kubernetes, autoscaling, and modern cloud platforms
- Experience with observability stacks (metrics, tracing, structured logging)
- Strong performance engineering and capacity planning background
- Communication and incident response skills
Preferred Qualifications
- Open-source contributions to model serving infrastructure
- Multi-region/global distributed AI serving experience
- Model quantization/distillation/compression exposure
- FinOps for AI workloads / cost-efficient serving design
- Experience supporting external-facing AI APIs at scale
About Bright Vision Technologies
Bright Vision Technologies is a technology consulting and software development company. It delivers cloud, AI, data, and enterprise solutions across the United States.
Scraped 7/25/2026