xelys jobs xelys jobs

Staff Site Reliability Engineer

BlinkRx

leadpermanentbackenddevops San Francisco, CA 53 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)Error BudgetsObservabilitySLIs/SLOsIncident ResponseLinuxPythonGoNetworkingAutomation

About the role

Role Overview

Staff Site Reliability Engineer (SRE) at BlinkRx. You will establish and lead reliability practices across the organization, drive observability and operational excellence, and take ownership of large, ambiguous infrastructure initiatives.

Responsibilities

  • Establish and evolve SRE best practices: reliability principles, error budgets, incident response, postmortems, and operational readiness standards.
  • Define and drive an observability strategy for health, performance, and reliability:
    • SLIs/SLOs, alert quality, dashboards, and service health indicators.
  • Design and implement software-driven infrastructure solutions to automate manual processes and reduce operational complexity and toil.
  • Serve as a technical leader/force multiplier across core cloud infrastructure, reliability tooling, and platform architecture.
  • Take ownership of large, ambiguous initiatives from concept to delivery while aligning engineering, security, and product stakeholders.
  • Use deep knowledge across software development, infrastructure, and security to improve:
    • resilience, scalability, performance, and compliance.
  • Proactively identify systemic reliability risks and lead upgrades/architectural improvements before incidents.
  • Partner with engineering teams to improve developer workflows, tooling, and operational maturity.
  • Provide mentorship, architecture guidance, and high-quality design/code reviews across infrastructure and product teams.
  • Lead documentation and knowledge sharing to avoid single-person ownership.
  • Participate in and mature incident response, escalation practices, and post-incident learning.

Requirements

  • Bachelor’s or Master’s degree in Computer Science (or equivalent practical experience).
  • 7+ years in SRE, infrastructure engineering, or platform engineering with impact at scale.
  • Expert, methodical troubleshooting across the full stack (application through OS and networking).
  • Strong command-line proficiency; deep expertise in Linux.
  • Advanced networking knowledge: load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication.
  • Software & automation:
    • Experience across multiple languages (e.g., Python, Go, Bash).
    • Familiarity with troubleshooting application stacks (e.g., React).
    • Strong track record of automating operational work to reduce toil.
    • Ability to design/build internal tools (Python or Go) to standardize and scale practices.
  • Comfortable working in an Agile environment with disciplined testing and quality practices.

Nice-to-haves

  • Cloud and platform depth (the posting begins a “Cloud & Platform Engineering” section but is truncated in the provided text).

About BlinkRx

BlinkRx (Blink Health) is a healthcare technology company building products that make prescriptions more accessible and affordable. Its platform streamlines the prescription supply chain by offering transparent pricing, home delivery, and support for patients and branded medications. The company operates cloud-based “pharma-to-patient” services and focuses on reliability and innovation in a traditionally slow-to-change industry.

Scraped 6/15/2026