xelys jobs xelys jobs

Engineering Manager, Inference Benchmarking — AI Perf

NVIDIA

leadpermanentengineering-managementbackend Colorado, United States 35 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Engineering ManagementLLM InferenceBenchmarkingKubernetesPrometheusZMQGPU TelemetryDCGMvLLMOpen Source

About the role

Role Overview

NVIDIA is hiring an Engineering Manager (Technical Lead Manager) for the Inference Benchmarking — AI Perf (AIPerf) effort within the Dynamo organization. You will lead an engineering team building and evolving NVIDIA’s open-source benchmarking platform used to measure LLM serving performance across inference frameworks and deployment environments (datacenter, local, and edge).

Responsibilities

  • Drive the technical roadmap for AIPerf’s core infrastructure, including:
    • load generation
    • ZMQ-based microservices
    • GPU telemetry and performance measurement
    • Prometheus metrics and statistical confidence intervals
    • Kubernetes-native deployment
  • Ensure benchmark accuracy and statistical soundness, so teams across the industry can rely on results for production inference decisions.
  • Advise and support upstream engine integrations in partnership with NVIDIA teams, including vLLM, TRT-LLM, and SGLang.
  • Lead, mentor, and grow a team of senior engineers in a fast-moving, high-external-visibility open-source environment.

Requirements

  • Bachelor’s degree in Computer Science/EE or related field, or equivalent experience.
  • 8+ years of software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
  • 3+ years in engineering leadership (Tech Lead/TLM/Engineering Manager).
  • Deep understanding of LLM inference mechanics and measurement rigor, including:
    • TTFT, ITL
    • KV caching
    • Prefill/Decode
    • speculative decoding
    • reasoning about measurement correctness and reproducibility.
  • Proven ability collaborating cross-functionally and delivering production-quality outputs in high-velocity environments.

Nice to Have / Ways to Stand Out

  • Extensive experience with vLLM, TRT-LLM, and/or SGLang internals and contributions to upstream projects.
  • Experience building Kubernetes-native infrastructure (operators, Helm charts) and GPU observability tooling (DCGM, dcgm-exporter, PyNVML).
  • Background with competitive benchmarking frameworks such as MLPerf or equivalent evaluation systems.
  • Track record of meaningful open-source contributions.

Location

  • Colorado, United States

About NVIDIA

NVIDIA is a technology company focused on computer graphics, PC gaming, and accelerated computing. It uses GPU and AI platforms to help enable modern computing for data centers, enterprise workloads, and real-world applications like robotics and self-driving cars.

Scraped 6/23/2026