xelys jobs xelys jobs

Senior Site Reliability Engineer

Alpaca

seniorpermanentdevopsbackend New York, NY Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)KubernetesGitOpsPostgreSQLSLIs/SLOsError BudgetsObservabilityIncident ResponseCDCLinux

About the role

Role Overview

Senior Site Reliability Engineer (SRE) at Alpaca, helping keep the brokerage platform reliable, observable, and operable as the company scales. The role spans cloud infrastructure, Kubernetes, observability, messaging, and data layers, with a strong emphasis on PostgreSQL reliability.

Responsibilities

  • Operate production day-to-day: on-call, incident response, postmortems, and follow-ups to close the loop.
  • Own reliability practice: define/refine SLIs/SLOs and error budgets; collaborate with product teams to live within them.
  • Improve observability: strengthen metrics, logs, traces, and alerting.
  • Ship infrastructure as code in a GitOps workflow for cloud resources and Kubernetes workloads.
  • Own PostgreSQL reliability posture:
    • performance tuning
    • schema and migration review
    • online migrations on large tables
    • HA/DR
    • CDC pipelines
  • Mentor engineers through code reviews, design reviews, and pairing focused on reliability and database fundamentals.

Requirements (Must-Haves)

  • 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.
  • Hands-on operating production services on Kubernetes.
  • Experience shipping infrastructure as code using GitOps.
  • Strong PostgreSQL production knowledge:
    • query plans and pg_stat_*
    • indexing and schema trade-offs
    • safe online migrations for non-trivial tables
  • Cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and comfort debugging cross-service connectivity.
  • Proficient with Linux at the operator level; comfortable with a modern observability stack.
  • Incident response experience: calm under pressure; structured debugging and postmortems that drive change.
  • Working proficiency in Go or Python.
  • Strong written and verbal communication.
  • Genuine interest in databases and growing PostgreSQL/DBA expertise.

Nice-to-Haves

  • Deeper PostgreSQL experience (e.g., large OLTP clusters), HA/DR ownership, and online migrations on big tables.

About Alpaca

Alpaca is a US-headquartered, global brokerage infrastructure company focused on agent-first trading support across stocks, ETFs, options, crypto, and fixed income. It serves hundreds of financial institutions worldwide with institutional-grade APIs and operates a trading-critical, developer-friendly platform backed by significant funding and open-source community involvement.

Scraped 8/6/2026