xelys jobs xelys jobs

Cloud Site Reliability Engineer

SambaNova

middevopsbackend San Jose, CA 54 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSREAI InferencingKubernetesPrometheusGrafanaDatadogMonitoringIncident ManagementAuto-scaling

About the role

Role Overview

SambaNova is seeking a Cloud Site Reliability Engineer (SRE) to own reliability, performance, and scalability for its AI inferencing service. You will bridge software development and operations to ensure inference endpoints deliver high uptime, low-latency responses, and efficient resource utilization. The role includes participating in a shared on-call rotation for 24/7 service reliability.

What You’ll Do

  • Service Ownership & On-Call

    • Take shared production ownership of the inferencing service, including:
      • availability, latency, performance, efficiency
      • change management, monitoring, emergency response, capacity planning
    • Support deployment of AI infrastructure in new regions (e.g., Asia, Europe, Latin America)
    • Participate in a balanced 24/7 on-call rotation (primary/secondary / follow-the-sun style)
  • Operational Excellence & Incident Management

    • Lead incident response for issues impacting the inferencing service
    • Run blameless post-mortems and implement corrective actions
    • Emphasize prevention via automation, robust testing, and resilient system design
    • Maintain actionable alerts with low false-positive rates (avoid alert fatigue)
  • Monitoring, Performance, and Scalability

    • Build and maintain advanced monitoring, alerting, and dashboards using tools such as:
      • Prometheus, Grafana, Datadog
    • Monitor service health and model performance (latency, throughput, error rates) and accelerator utilization
    • Identify performance bottlenecks and implement cost-effective auto-scaling for variable inference load

On-Call Principles

  • Balanced rotation across the team to prevent disproportionate burden
  • Prevention-focused approach (automation/system design)
  • Actionable alerts requiring human intervention
  • Blameless post-mortems and corrective action to prevent recurrence

About SambaNova

SambaNova is building AI computing for enterprise and government customers, combining integrated hardware and software to run generative AI at scale. Its SambaNova Suite is a full-stack generative AI platform delivered on-premises or in the cloud, powered by the SN40L chip and supported by open-source models that can be securely fine-tuned with customer data.

Scraped 6/15/2026