Senior Staff AI/ML System Software Engineer
d-Matrix
full-remoteleadpermanentbackendfullstack Full remote 73 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
CC++PythonLinuxDistributed SystemsComputer ArchitectureMLIRPyTorchTensorRTMLOps
About the role
Role overview
As a Senior Staff AI/ML System Software Engineer at d-Matrix, you will develop and maintain next-generation AI deployment software. You’ll work on system software and the end-to-end toolchain needed to deploy ML workloads efficiently, collaborating closely with hardware and software experts, including compiler specialists.
Key missions
- Help commercialize the AI software stack for the AI compute engine.
- Develop, enhance, and maintain next-generation AI deployment software.
- Partner with compiler/infrastructure experts to build compiler infrastructure.
Responsibilities
- Design and implement distributed, high-performance system software for AI deployment.
- Optimize the full-stack toolchain and ensure scaling of software deliverables.
- Collaborate cross-functionally with experts in compilers, hardware/software systems, and ML.
Requirements
- 7+ years industry software development (or MS preferred with 5+ years).
- Strong foundations in computer architecture, data structures, system software, and ML fundamentals.
- C/C++/Python development in a Linux environment with standard development tools.
- Degree in Computer Science/Engineering/Math/Physics (BS required; MS/PhD preferred).
- Experience with distributed systems and high-performance software design.
Nice-to-haves / preferred experience
- SIMD algorithms on vector processors.
- Open-source ML compiler frameworks such as MLIR.
- Deep learning frameworks: PyTorch, TensorFlow.
- ML runtimes: ONNX Runtime, TensorRT.
- Inference/model serving frameworks: Triton, TensorFlow Serving, KubeFlow.
- Distributed systems collectives such as NCCL and OpenMPI.
- Deploying ML workloads on distributed, multi-tenant systems.
- MLOps from training to deployment (training, quantization, sparsity, preprocessing).
- Training/tuning/deploying models for CV (ResNet), NLP (BERT/GPT), and/or recommendation (DLRM).
About d-Matrix
d-Matrix develops AI deployment software and tooling for next-generation AI compute. The team focuses on building logic and compiler/infrastructure components that enable efficient, scalable deployment of machine-learning workloads.
Scraped 5/12/2026