Senior Engineering Manager (Site Reliability)
Horizon3
full-remoteseniorpermanentengineering-managementdevops Full remote Today via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
AWSSRESite Reliability EngineeringObservabilityIncident ManagementPagerDutyFireHydrantSLO/SLARunbooksAPM
About the role
Role overview
Senior Engineering Manager (Site Reliability) at Horizon3 (full remote). You will build and lead the Site Reliability function, professionalize incident management, and drive engineering reliability and quality through observability and operational excellence.
Key missions
- Build and lead the Site Reliability function, including recruiting and managing a team of 4–6 SRE/Infrastructure engineers.
- Professionalize incident management by defining and documenting incident processes and practices for SRE and application feature teams.
- Balance incident response with execution of a roadmap across observability and reliability engineering.
- Lead horizontally with peer management and senior engineering leaders; mentor and develop your team.
Responsibilities
- Establish and evolve SRE practices including on-call programs, SLO/SLA definition, and operational runbooks.
- Select and deploy incident management tooling (e.g., PagerDuty, FireHydrant) and define decision-making frameworks for vendor selection.
- Partner with Product and Engineering leaders to define SLOs and drive a culture of ownership.
- Produce durable process documentation and support adoption across teams.
- Provide training/education on runbooks, on-call expectations, investigations, and effective postmortems.
- Manage strategic reliability roadmap initiatives alongside operational needs and capacity constraints.
Requirements
- Strong working knowledge of a major cloud provider (AWS preferred, or GCP/Azure): infrastructure, networking, managed services, IAM, and cost management.
- Deep hands-on observability expertise: APM, logs and traces, golden signals, and service-specific metrics.
- Proven experience leading hiring and growing SRE/Infrastructure teams; experience building SRE functions and incident management processes.
- Previous career experience as a Site Reliability Engineer; comfortable being hands-on while scaling the team.
- Ability to engage with engineering/product leadership to define SLOs and drive ownership.
- Strong skills in clear process documentation.
- Experience with architecture decision records and RFC-style practices (in infra/platform context).
- Ability to train others and run effective postmortems.
- Demonstrated ability to build, scale, and retain high-performing distributed/remote teams.
Nice to have
- Experience selecting and deploying additional incident management tooling and frameworks beyond the examples.
Location / work model
- Full remote.
About Horizon3
Horizon3 is a technology company building and operating software products that rely on reliable cloud infrastructure. The role focuses on establishing Site Reliability Engineering (SRE) capabilities, including incident management, observability, and reliability engineering processes.
Scraped 7/30/2026