Senior Site Reliability Engineer (SRE) - Remote Work
BairesDev
full-remoteseniorpermanentdevopssecurity Austin, TX 2 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
KubernetesTerraformHelmCI/CDObservabilityPrometheusGrafanaDatadogSLOs/SLIsIAM Hardening
About the role
Role Overview
As a Senior Site Reliability Engineer (SRE), you will ensure critical systems are reliable, observable, and resilient at scale. You’ll define reliability standards, improve incident learnings, and work across multi-cluster Kubernetes and modern observability stacks.
Responsibilities
- Manage and scale Kubernetes environments across multiple clusters.
- Build and maintain CI/CD pipelines and Infrastructure as Code (IaC).
- Implement and maintain observability for proactive issue detection.
- Define and track reliability standards and drive continuous improvement using incident learnings.
- Lead post-incident review practices and reliability management processes.
Requirements
- 5+ years in Site Reliability Engineering or infrastructure engineering.
- Strong Kubernetes expertise (operators, autoscaling, multi-cluster management).
- Experience engineering CI/CD pipelines.
- Proficiency with Infrastructure as Code using Terraform and Helm.
- Hands-on observability experience with Prometheus, Grafana, Datadog, and/or OpenTelemetry.
- Ability to define SLOs/SLIs, manage error budgets, and run post-incident reviews.
- Background in cloud security tooling and IAM hardening.
- Advanced English proficiency.
Nice to Have
- Not explicitly stated (focus areas above are the detailed expectations).
About BairesDev
BairesDev is a technology services company that has delivered large-scale and innovative engineering projects for over 15 years. It works with major industry players and startups, leveraging a large global remote team of highly skilled engineers.
Scraped 8/6/2026