xelys jobs xelys jobs

Senior Site Reliability Engineer

Fieldguide

full-remoteseniorpermanentdevopsbackend Full remote 74 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringTerraformAWSObservabilitySLOs/SLIsPrometheusGrafanaDatadogOpenTelemetryIncident Response

About the role

Role overview

Join Fieldguide as a Senior Site Reliability Engineer (SRE) to ensure the reliability, scalability, and observability of production systems. You’ll partner with product and platform engineering teams to define reliability standards, improve system performance, and build robust monitoring and operational practices.

Responsibilities

  • Design and operate highly scalable, fault-tolerant systems for production workloads in a distributed cloud environment
  • Define and implement SLOs/SLIs and error budgets to guide reliability decisions
  • Build and improve observability systems, including:
    • Metrics, logs, and tracing
    • Deep visibility into system behavior and performance
  • Drive continuous improvement of system resilience and operational practices

Requirements

  • Proficiency with Infrastructure as Code (especially Terraform)
  • Deep understanding of system performance, reliability patterns, and distributed system failure modes
  • Strong experience operating and scaling distributed systems in cloud environments (AWS preferred)
  • Hands-on experience with observability platforms such as:
    • Datadog, Prometheus, Grafana, CloudWatch
  • Experience defining SLOs/SLIs and using them to influence engineering priorities
  • Proficiency in at least one programming/scripting language for automation and tooling
  • 5+ years of experience in SRE, infrastructure, or related software engineering
  • Experience supporting production systems via on-call and incident response
  • Strong communication and collaboration skills across engineering and product teams

Nice-to-haves

  • Experience with distributed tracing (e.g., OpenTelemetry)
  • Capacity planning and performance benchmarking at scale
  • Familiarity with database performance tuning and observability
  • Exposure to compliance/regulation-heavy environments (e.g., SOC 2, FedRAMP)
  • Experience with chaos engineering to proactively strengthen resilience

Location / remote

  • Full remote

About Fieldguide

Fieldguide is a technology company building and operating production software systems for its customers. The role focuses on reliability engineering across distributed cloud infrastructure, emphasizing scalability, observability, and resilience practices.

Scraped 5/12/2026