xelys jobs xelys jobs

Developer

NLB Services

full-remotemidpermanentbackenddevops United States Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

GPU Kernel DevelopmentPyTorchTritonGEMM OptimizationCUDAHPCParallel ProgrammingGPU ArchitectureTransformer ModelsBERT

About the role

Role Overview

GPU Kernel Developer (AI/ML) working on high-performance optimization and scaling of large language models (LLMs), with a focus on BERT and transformer workloads. The role is remote and full-time, centered on developing and tuning GPU kernels for modern ML frameworks.

Responsibilities

  • Design, implement, and optimize GPU kernels for AI/ML workloads
  • Concentrate on GEMM (matrix multiplication) and other performance-critical operations
  • Collaborate with researchers and engineers to accelerate BERT and transformer-based models
  • Integrate optimized kernels into PyTorch and Triton
  • Perform profiling, benchmarking, and tuning across diverse GPU hardware/architectures
  • Work with system engineers to support production deployment

Requirements (Must-Have)

  • Strong GPU kernel development expertise
  • Hands-on experience with PyTorch and Triton (mandatory)
  • Proficiency in Python, C, and C++
  • Deep understanding of system-level programming and GPU architecture
  • Experience in AI/ML model optimization, especially BERT
  • Strong knowledge of HPC and parallel programming

Nice-to-Haves

  • Experience deploying large-scale AI/ML models
  • Familiarity with CUDA (or other GPU programming frameworks)
  • Background in distributed training systems
  • Strong problem-solving and debugging skills

Scraped 7/28/2026