xelys jobs xelys jobs

Principal LLM Inference Engineer

d-Matrix

full-remoteleadpermanentbackenddevops Full remote 20 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

LLM InferencevLLMTensorRT-LLMSGLangONNX RuntimeCUDATritonDistributed InferenceQuantizationProduction Serving

About the role

Role Overview

Join d-Matrix as a Principal LLM Inference Engineer. You will develop and optimize end-to-end LLM inference systems, working across kernel-level performance improvements and high-level serving APIs. You’ll identify emerging inference use cases, build proof-of-concept systems, and contribute to distributed inference.

Key Missions

  • Concept and prototype emerging LLM inference use cases tailored to heterogeneous hardware deployments.
  • Develop and optimize custom kernels/operators to maximize throughput and minimize latency.
  • Build distributed inference systems, including:
    • Tensor parallelism / pipeline parallelism
    • Disaggregated serving concepts such as prefill/decoding separation.

Responsibilities

  • Own problems end-to-end (from prototype → optimization → handoff).
  • Collaborate with hardware architects and partner with product/business development.
  • Explore new heterogeneous deployment topologies and create early-stage POCs.
  • Improve inference performance by squeezing maximum efficiency from modern generative AI architectures.

Requirements

  • Strong bias for action; can take ownership from prototype to optimization.
  • Deep intuition for modern generative AI architectures and inference-time performance.
  • Familiarity with open-source inference frameworks and ability to extend/replace them, including:
    • vLLM, SGLang, TensorRT-LLM (and similar)
  • Experience optimizing LLM inference, including:
    • Attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4)
  • Strong Python and C/C++ proficiency.
  • Experience with heterogeneous compute deployments (scheduling across CPUs/GPUs/accelerators).
  • GPU kernel programming and profiling knowledge (CUDA/Triton and performance profiling tools).
  • Distributed inference experience (tensor/pipeline parallelism; disaggregated serving).
  • Production inference serving at scale (latency SLOs, continuous batching, multi-model serving).
  • 6+ years relevant experience with a Master’s/PhD preferred; 10+ years with a Bachelor’s degree acceptable (or equivalent).

Nice to Have

  • Familiarity with custom silicon / ASIC-based inference beyond GPU-only.
  • Experience contributing to open-source inference/ML systems.
  • Knowledge of speculative decoding, Mixture-of-Experts (MoE) routing, and long-context serving.
  • Familiarity with the JAX Scaling Book or equivalent systems-level understanding.

Benefits / Work Model

  • Equity, healthcare, and flexible time off
  • Remote/hybrid working model (listed as full remote).

About d-Matrix

d-Matrix is a company focused on building high-performance inference systems for large language models. It works across the LLM inference stack, spanning kernel/operator optimization to distributed serving, with close collaboration across hardware and product teams.

Scraped 7/16/2026