Lead Site Reliability Engineer
Empower
leadpermanentdevopssecurity United States Today via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
AWSSite Reliability Engineering (SRE)SLO/SLI/Error BudgetsTerraformKubernetesDatadogSplunkCI/CDChaos EngineeringFinOps
About the role
Role Overview
Lead Site Reliability Engineer to drive reliability across Empower’s financial services platform. You’ll combine deep SRE expertise with technical leadership—setting standards, guiding incident response and recovery planning, and partnering with engineering, product, and security to align reliability initiatives with business and compliance needs.
Responsibilities
- Lead cross-functional reliability initiatives across multiple value streams and coordinate execution across teams.
- Define and evolve organization-wide SRE best practices, tools, and methodologies.
- Architect enterprise-scale, multi-region AWS infrastructure balancing reliability, cost, performance, and security.
- Establish and operate SLOs, SLIs, and error budgets to drive service prioritization.
- Serve as incident commander for major incidents and lead postmortems with completed action items and organizational learning.
- Lead disaster recovery planning for critical financial services infrastructure.
- Build shared Infrastructure as Code foundations using Terraform (reusable modules, standards, patterns adopted across teams).
- Design and implement production-scale Kubernetes patterns (multi-tenancy, security policies, advanced scheduling).
- Establish observability standards using Datadog and Splunk (metrics, logging, tracing, dashboards, alerting).
- Set CI/CD standards using pipeline-as-code and progressive delivery at scale.
- Lead chaos engineering, game days, and systematic reliability testing.
- Drive FinOps initiatives to optimize cloud spend while maintaining reliability targets.
- Lead a functional SRE team through projects and operational initiatives (without direct reports) and mentor SREs via coaching and reviews.
- Partner with Engineering, Product, and Security leadership on reliability work, zero-trust architecture, and compliance controls.
Requirements
- Bachelor’s degree in Computer Science/IT (or equivalent practical experience).
- 7–10 years of Site Reliability Engineering experience with demonstrated technical leadership.
- Proven ability to lead complex projects to completion and lead technical teams.
- Expert AWS knowledge, including large-scale multi-region architecture.
- Deep Kubernetes expertise, including security and production operations.
- Mastery of Terraform for shared platforms/frameworks.
- Strong software engineering background with production experience in Python and/or Go.
- Extensive observability experience with Datadog and Splunk.
- Deep understanding of CI/CD principles and enterprise-grade pipeline implementation.
- Track record leading major incidents and conducting effective postmortems.
- Strong understanding of security, networking, and infrastructure design patterns.
Nice-to-Haves
- Strong experience aligning reliability initiatives with zero-trust architecture and compliance requirements.
- Demonstrated mentoring through design/code reviews and training sessions.
Additional Notes
- Applicants must be authorized to work in the U.S.; the company is unable to sponsor or take over sponsorship of employment visas (including CPT/OPT).
About Empower
Empower is a financial services company focused on helping customers transform their financial lives. It promotes internal mobility, purpose, well-being, and an inclusive, flexible work environment while supporting employees’ community involvement.
Scraped 7/29/2026