Developer
NLB Services
full-remotemidpermanentbackenddevops United States Yesterday via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
GPU Kernel DevelopmentPyTorchTritonGEMM OptimizationCUDAHPCParallel ProgrammingGPU ArchitectureTransformer ModelsBERT
About the role
Role Overview
GPU Kernel Developer (AI/ML) working on high-performance optimization and scaling of large language models (LLMs), with a focus on BERT and transformer workloads. The role is remote and full-time, centered on developing and tuning GPU kernels for modern ML frameworks.
Responsibilities
- Design, implement, and optimize GPU kernels for AI/ML workloads
- Concentrate on GEMM (matrix multiplication) and other performance-critical operations
- Collaborate with researchers and engineers to accelerate BERT and transformer-based models
- Integrate optimized kernels into PyTorch and Triton
- Perform profiling, benchmarking, and tuning across diverse GPU hardware/architectures
- Work with system engineers to support production deployment
Requirements (Must-Have)
- Strong GPU kernel development expertise
- Hands-on experience with PyTorch and Triton (mandatory)
- Proficiency in Python, C, and C++
- Deep understanding of system-level programming and GPU architecture
- Experience in AI/ML model optimization, especially BERT
- Strong knowledge of HPC and parallel programming
Nice-to-Haves
- Experience deploying large-scale AI/ML models
- Familiarity with CUDA (or other GPU programming frameworks)
- Background in distributed training systems
- Strong problem-solving and debugging skills
Scraped 7/28/2026