xelys jobs xelys jobs

Site Reliability Engineer, Team Lead

Omnicell

hybridleaddevopsengineering-management Austin, TX 10 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)SLOsSLIsError BudgetsIncident ManagementObservabilityHIPAASOC 2AIOpsML-assisted Observability

About the role

Role Overview

Site Reliability Engineer, Team Lead (Player‑Coach) Omnicell is establishing a Global Cloud Operations organization and hiring its first senior SRE to design and run the Site Reliability Engineering (SRE) practice end-to-end. This is a hands-on role where you will define “good” reliability for Omnicell, set standards, make foundational technology decisions, and operate in production alongside coaching responsibilities.

Location / Eligibility

  • Preferred local: Austin, TX (U.S.) or Cranberry Woods, PA
  • Open to remote: fully remote in mainland USA
  • Visa sponsorship: not offered
  • Must be: U.S. citizen or Permanent Resident

Responsibilities

  • Define and operate Omnicell’s SRE function, balancing engineering with practice design and cross-functional leadership.
  • Ensure Tier‑1 cloud services are observable, resilient, and dependable.
  • Reliability Practice & Operating Model
    • Define and publish SLIs, SLOs, and error budgets for top Tier‑1 customer-facing services.
    • Design incident command: severity definitions, incident declaration criteria, war-room protocols, and stakeholder communications.
    • Shape the on-call rotation experience.
    • Standardize on an observability platform.
    • Prioritize reliability investment against feature velocity.
  • Regulated operations & security/auditability alignment in HIPAA, SOC 2, and sometimes FedRAMP.
  • AIOps / ML-assisted operations (technical ownership): introduce and validate ML-assisted observability and AI-driven operations (e.g., anomaly detection, intelligent alert correlation, LLM-assisted runbook generation), prioritizing against foundational reliability work.
  • Team building (initially small team): run the plays yourself until the team grows, and coach an Engineer III SRE.

Requirements

  • Strong experience designing and operating reliability/operations practices in production.
  • Ability to define SLO/SLI/error budget frameworks and incident management processes.
  • Comfort operating in a regulated healthcare environment where reliability, security, and auditability must be integrated.

Nice-to-Haves

  • Experience introducing or running AIOps / ML-assisted observability.
  • Experience standardizing observability across services.
  • Familiarity with environments spanning both private connectivity and public internet and with hybrid product architectures (hospital hardware communicating with cloud services).

About Omnicell

Omnicell provides cloud-native, SaaS-delivered healthcare technology that hospitals depend on for 24/7 medication dispensing and patient care. The company is building its Global Cloud Operations organization as it transitions from on-prem, hardware-centric products to a regulated cloud platform serving healthcare environments.

Scraped 7/19/2026