xelys jobs xelys jobs

Site Reliability Engineer, Inference Infrastructure

Cohere

seniorpermanentdevopsbackend New York, NY 3 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringKubernetesSLOsOn-callObservabilityResilience EngineeringDistributed SystemsGPU WorkloadsLinuxGolang

About the role

Role Overview

Site Reliability Engineer (Inference Infrastructure) on the Model Serving team at Cohere. The team develops, deploys, and operates the AI platform that delivers Cohere’s large language models through API endpoints.

Responsibilities

  • Build self-service systems to automate managing, deploying, and operating services.
  • Develop and maintain custom Kubernetes operators for language model deployments.
  • Automate observability and improve resilience across environments.
  • Enable developers to troubleshoot and resolve issues efficiently.
  • Ensure targets are met for SLOs, including participation in an on-call rotation.
  • Collaborate with internal developers and influence the Infrastructure roadmap using feedback.
  • Grow the team via knowledge sharing and an active review process.

Requirements

  • 5+ years of engineering experience running production infrastructure at large scale.
  • Experience designing highly available distributed systems with Kubernetes and GPU workloads.
  • Kubernetes development and production coding/support experience.
  • Experience with clouds/hybrid/multi-cloud: GCP, Azure, AWS, OCI, and/or on-prem hybrid serving.
  • Ability to design, deploy, support, and troubleshoot complex Linux-based computing environments.
  • Strong knowledge of compute/storage/network resource and cost management.
  • Strong collaboration and troubleshooting skills for mission-critical systems.
  • Understanding of accelerator characteristics (GPUs, TPUs, and/or custom accelerators) and their impact on latency and throughput.
  • Working knowledge of distributed systems.
  • Experience with high-performance server languages, including Golang and/or C++ (or similar).

Nice-to-Haves

  • Familiarity with computational characteristics of accelerators and inference-specific performance optimization (latency/throughput).
  • Experience optimizing and operating inference infrastructure at scale.

About Cohere

Cohere is a security-first enterprise AI company that builds foundation AI models and end-to-end products. It trains and deploys frontier models for enterprises building real-world AI systems, with a focus on delivering large language models via production APIs. The company is a global technology organization with offices in major cities worldwide.

Scraped 7/28/2026