xelys jobs xelys jobs

Site Reliability Engineer

Five9

midpermanentdevopsbackendsecurity United States 115 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)LinuxTerraformAnsibleCI/CDObservabilitySLIs/SLOsIncident ResponseSecurity ScanningCost Optimization

About the role

Role Overview

Five9 is hiring a Site Reliability Engineer (SRE) to build and maintain highly reliable, scalable systems. The role is ~50% software engineering and ~50% operations, emphasizing automation, monitoring, and reliability engineering over manual operations.

Responsibilities

  • Observability & Monitoring
    • Design and implement dashboards and metrics for both OS/platform and application monitoring
    • Use primary (RED) and secondary (USE) indicators
  • Availability & Reliability Engineering
    • Define and manage SLIs (Service Level Indicators), SLOs (Service Level Objectives), and error budgets
  • Performance Monitoring & Alerting
    • Build proactive alerting and performance monitoring to prevent user impact
  • Incident Management
    • Participate in on-call rotations and lead incident response
    • Run post-mortems and drive remediation
    • Maintain the official on-call routing
  • CI/CD & Deployment Pipeline Management
    • Maintain continuous integration and deployment pipelines across cloud and on-prem deployments
  • Infrastructure as Code & Configuration Management
    • Develop and maintain infrastructure using Terraform, Ansible, or similar tools
    • Automate configuration and ensure consistency across environments
  • Security & Compliance
    • Ensure security scanning exists and review escalated vulnerabilities
    • Maintain authentication, authorization, and audit logging
    • Support regulatory and industry compliance reporting
    • Participate in security incident response and remediation
  • Cost Optimization & Capacity Planning
    • Monitor and optimize cloud resource usage and costs (planned and unplanned changes)
    • Analyze usage patterns and plan future capacity
    • Recommend cost-effective architecture and right-sizing strategies (including automated scaling)
  • Platform / Common Services Engineering
    • Build and maintain shared services (e.g., notification systems, caching layers, message queues)
    • Support database reliability/performance/scaling where not handled by dedicated DB teams
    • Implement service discovery, load balancing, and network policies
    • Create developer tools to improve developer productivity and system reliability

Required Qualifications

  • 3+ years managing large-scale production environments
  • Comfortable with 24/7 on-call and incident response
  • Strong Linux/Unix system administration skills
  • Understanding of TCP/IP, DNS, and load balancing

Nice-to-haves / Implied Skills

  • Experience with observability using RED/USE metrics
  • Hands-on Infrastructure as Code with Terraform and/or Ansible
  • Experience with CI/CD pipelines for cloud and on-prem deployments
  • Familiarity with security scanning, access control, and audit logging
  • Background in cost optimization and capacity planning

About Five9

Five9 is a provider of cloud contact center software, delivering cloud-based customer experience and supporting organizations worldwide. The company emphasizes innovation and a team-first, inclusive culture to build and maintain reliable customer-facing services.

Scraped 4/1/2026