Site Reliability Engineer
Five9
midpermanentdevopsbackendsecurity United States 115 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability Engineering (SRE)LinuxTerraformAnsibleCI/CDObservabilitySLIs/SLOsIncident ResponseSecurity ScanningCost Optimization
About the role
Role Overview
Five9 is hiring a Site Reliability Engineer (SRE) to build and maintain highly reliable, scalable systems. The role is ~50% software engineering and ~50% operations, emphasizing automation, monitoring, and reliability engineering over manual operations.
Responsibilities
- Observability & Monitoring
- Design and implement dashboards and metrics for both OS/platform and application monitoring
- Use primary (RED) and secondary (USE) indicators
- Availability & Reliability Engineering
- Define and manage SLIs (Service Level Indicators), SLOs (Service Level Objectives), and error budgets
- Performance Monitoring & Alerting
- Build proactive alerting and performance monitoring to prevent user impact
- Incident Management
- Participate in on-call rotations and lead incident response
- Run post-mortems and drive remediation
- Maintain the official on-call routing
- CI/CD & Deployment Pipeline Management
- Maintain continuous integration and deployment pipelines across cloud and on-prem deployments
- Infrastructure as Code & Configuration Management
- Develop and maintain infrastructure using Terraform, Ansible, or similar tools
- Automate configuration and ensure consistency across environments
- Security & Compliance
- Ensure security scanning exists and review escalated vulnerabilities
- Maintain authentication, authorization, and audit logging
- Support regulatory and industry compliance reporting
- Participate in security incident response and remediation
- Cost Optimization & Capacity Planning
- Monitor and optimize cloud resource usage and costs (planned and unplanned changes)
- Analyze usage patterns and plan future capacity
- Recommend cost-effective architecture and right-sizing strategies (including automated scaling)
- Platform / Common Services Engineering
- Build and maintain shared services (e.g., notification systems, caching layers, message queues)
- Support database reliability/performance/scaling where not handled by dedicated DB teams
- Implement service discovery, load balancing, and network policies
- Create developer tools to improve developer productivity and system reliability
Required Qualifications
- 3+ years managing large-scale production environments
- Comfortable with 24/7 on-call and incident response
- Strong Linux/Unix system administration skills
- Understanding of TCP/IP, DNS, and load balancing
Nice-to-haves / Implied Skills
- Experience with observability using RED/USE metrics
- Hands-on Infrastructure as Code with Terraform and/or Ansible
- Experience with CI/CD pipelines for cloud and on-prem deployments
- Familiarity with security scanning, access control, and audit logging
- Background in cost optimization and capacity planning
About Five9
Five9 is a provider of cloud contact center software, delivering cloud-based customer experience and supporting organizations worldwide. The company emphasizes innovation and a team-first, inclusive culture to build and maintain reliable customer-facing services.
Scraped 4/1/2026