Principal LLM Inference Engineer
d-Matrix
full-remoteleadpermanentbackenddevops Full remote 20 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
LLM InferencevLLMTensorRT-LLMSGLangONNX RuntimeCUDATritonDistributed InferenceQuantizationProduction Serving
About the role
Role Overview
Join d-Matrix as a Principal LLM Inference Engineer. You will develop and optimize end-to-end LLM inference systems, working across kernel-level performance improvements and high-level serving APIs. You’ll identify emerging inference use cases, build proof-of-concept systems, and contribute to distributed inference.
Key Missions
- Concept and prototype emerging LLM inference use cases tailored to heterogeneous hardware deployments.
- Develop and optimize custom kernels/operators to maximize throughput and minimize latency.
- Build distributed inference systems, including:
- Tensor parallelism / pipeline parallelism
- Disaggregated serving concepts such as prefill/decoding separation.
Responsibilities
- Own problems end-to-end (from prototype → optimization → handoff).
- Collaborate with hardware architects and partner with product/business development.
- Explore new heterogeneous deployment topologies and create early-stage POCs.
- Improve inference performance by squeezing maximum efficiency from modern generative AI architectures.
Requirements
- Strong bias for action; can take ownership from prototype to optimization.
- Deep intuition for modern generative AI architectures and inference-time performance.
- Familiarity with open-source inference frameworks and ability to extend/replace them, including:
- vLLM, SGLang, TensorRT-LLM (and similar)
- Experience optimizing LLM inference, including:
- Attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4)
- Strong Python and C/C++ proficiency.
- Experience with heterogeneous compute deployments (scheduling across CPUs/GPUs/accelerators).
- GPU kernel programming and profiling knowledge (CUDA/Triton and performance profiling tools).
- Distributed inference experience (tensor/pipeline parallelism; disaggregated serving).
- Production inference serving at scale (latency SLOs, continuous batching, multi-model serving).
- 6+ years relevant experience with a Master’s/PhD preferred; 10+ years with a Bachelor’s degree acceptable (or equivalent).
Nice to Have
- Familiarity with custom silicon / ASIC-based inference beyond GPU-only.
- Experience contributing to open-source inference/ML systems.
- Knowledge of speculative decoding, Mixture-of-Experts (MoE) routing, and long-context serving.
- Familiarity with the JAX Scaling Book or equivalent systems-level understanding.
Benefits / Work Model
- Equity, healthcare, and flexible time off
- Remote/hybrid working model (listed as full remote).
About d-Matrix
d-Matrix is a company focused on building high-performance inference systems for large language models. It works across the LLM inference stack, spanning kernel/operator optimization to distributed serving, with close collaboration across hardware and product teams.
Scraped 7/16/2026