xelys jobs xelys jobs

Senior Site Reliability Engineer

Block

seniorpermanentdevopsbackend United States 53 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)KotlinJavaTerraformKubernetesIstioEnvoyAWSObservabilityCI/CD

About the role

Role Overview

As a Senior Site Reliability Engineer (SRE) on Block’s SRE team, you will proactively and reactively improve the reliability of Block’s platform and critical infrastructure. The role is metrics-driven and systems-oriented, with a focus on distributed platforms and building AI-enabled tooling for observability, incident detection/response, and automation.

Responsibilities

  • Improve reliability of platform and Tier 0 (most critical) services through proactive and reactive work.
  • Lead incident command during high-severity events (sev 0–1), coordinating mitigation and escalation.
  • Serve as primary oncall for Tier 0 services (12 hours/day, about one week every few weeks depending on team size).
  • Build and extend reliability platforms and standardize reliability tools across teams.
  • Triage, coordinate, and lead stabilization during sev 0–1 incidents.
  • Drive platform-wide reliability improvements, including shared operational tooling and deploy-safety patterns.
  • Use AI-driven systems to improve signal detection, reduce alert noise, and accelerate root cause analysis.
  • Design and implement safe deployment patterns (e.g., progressive delivery, automated rollback, guardrails).
  • Create and maintain evidence-based maturity assessments using trailing 90-day data windows.
  • Manage vendor/dependency escalation contacts with fast response expectations (reachable within ≤ 5 minutes).

Requirements

  • 5+ years of software development experience.
  • Experience with production oncall for high-availability systems.
  • Strong incident management skills: structured triage, mitigation under pressure, and blameless postmortems.
  • CI/CD fluency, progressive rollout strategies, and rollback automation.
  • Monitoring & observability expertise, including alert tuning for uptime, error rates, latency regressions, and resource exhaustion.
  • Demonstrated initiative and leadership on prior projects, especially with backend/platform focus.
  • Comfort reaching for AI to accelerate problem-solving and reduce toil.

Nice to Have

  • Familiarity with AI-driven tooling for observability, incident analysis, or automation.
  • Technical initiative with systems that have many moving parts and strong root-cause ownership.
  • Experience with evidence-based reliability program management.

On-Call

  • Primary platform oncall is 12 hours per day, roughly one week every few weeks (team-dependent).

About Block

Block is a technology company focused on economic empowerment. It builds and supports foundational teams across People, Finance, Counsel, Hardware, Information Security, and Platform Infrastructure Engineering to safeguard systems and enable scalable product development.

Scraped 6/12/2026