Site Reliability Engineer, Inference Infrastructure
Cohere
seniorpermanentdevopsbackend New York, NY 3 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability EngineeringKubernetesSLOsOn-callObservabilityResilience EngineeringDistributed SystemsGPU WorkloadsLinuxGolang
About the role
Role Overview
Site Reliability Engineer (Inference Infrastructure) on the Model Serving team at Cohere. The team develops, deploys, and operates the AI platform that delivers Cohere’s large language models through API endpoints.
Responsibilities
- Build self-service systems to automate managing, deploying, and operating services.
- Develop and maintain custom Kubernetes operators for language model deployments.
- Automate observability and improve resilience across environments.
- Enable developers to troubleshoot and resolve issues efficiently.
- Ensure targets are met for SLOs, including participation in an on-call rotation.
- Collaborate with internal developers and influence the Infrastructure roadmap using feedback.
- Grow the team via knowledge sharing and an active review process.
Requirements
- 5+ years of engineering experience running production infrastructure at large scale.
- Experience designing highly available distributed systems with Kubernetes and GPU workloads.
- Kubernetes development and production coding/support experience.
- Experience with clouds/hybrid/multi-cloud: GCP, Azure, AWS, OCI, and/or on-prem hybrid serving.
- Ability to design, deploy, support, and troubleshoot complex Linux-based computing environments.
- Strong knowledge of compute/storage/network resource and cost management.
- Strong collaboration and troubleshooting skills for mission-critical systems.
- Understanding of accelerator characteristics (GPUs, TPUs, and/or custom accelerators) and their impact on latency and throughput.
- Working knowledge of distributed systems.
- Experience with high-performance server languages, including Golang and/or C++ (or similar).
Nice-to-Haves
- Familiarity with computational characteristics of accelerators and inference-specific performance optimization (latency/throughput).
- Experience optimizing and operating inference infrastructure at scale.
About Cohere
Cohere is a security-first enterprise AI company that builds foundation AI models and end-to-end products. It trains and deploys frontier models for enterprises building real-world AI systems, with a focus on delivering large language models via production APIs. The company is a global technology organization with offices in major cities worldwide.
Scraped 7/28/2026