Senior Site Reliability Engineer
Fieldguide
full-remoteseniorpermanentdevopsbackend Full remote 74 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability EngineeringTerraformAWSObservabilitySLOs/SLIsPrometheusGrafanaDatadogOpenTelemetryIncident Response
About the role
Role overview
Join Fieldguide as a Senior Site Reliability Engineer (SRE) to ensure the reliability, scalability, and observability of production systems. You’ll partner with product and platform engineering teams to define reliability standards, improve system performance, and build robust monitoring and operational practices.
Responsibilities
- Design and operate highly scalable, fault-tolerant systems for production workloads in a distributed cloud environment
- Define and implement SLOs/SLIs and error budgets to guide reliability decisions
- Build and improve observability systems, including:
- Metrics, logs, and tracing
- Deep visibility into system behavior and performance
- Drive continuous improvement of system resilience and operational practices
Requirements
- Proficiency with Infrastructure as Code (especially Terraform)
- Deep understanding of system performance, reliability patterns, and distributed system failure modes
- Strong experience operating and scaling distributed systems in cloud environments (AWS preferred)
- Hands-on experience with observability platforms such as:
- Datadog, Prometheus, Grafana, CloudWatch
- Experience defining SLOs/SLIs and using them to influence engineering priorities
- Proficiency in at least one programming/scripting language for automation and tooling
- 5+ years of experience in SRE, infrastructure, or related software engineering
- Experience supporting production systems via on-call and incident response
- Strong communication and collaboration skills across engineering and product teams
Nice-to-haves
- Experience with distributed tracing (e.g., OpenTelemetry)
- Capacity planning and performance benchmarking at scale
- Familiarity with database performance tuning and observability
- Exposure to compliance/regulation-heavy environments (e.g., SOC 2, FedRAMP)
- Experience with chaos engineering to proactively strengthen resilience
Location / remote
- Full remote
About Fieldguide
Fieldguide is a technology company building and operating production software systems for its customers. The role focuses on reliability engineering across distributed cloud infrastructure, emphasizing scalability, observability, and resilience practices.
Scraped 5/12/2026