Reliability Monitoring Engineer
Jobgether
seniorpermanentbackend United States Yesterday via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
ObservabilitySREReliability EngineeringOpenTelemetryDistributed TracingPrometheusGrafanaDatadogDatadogStructured Logging
About the role
Role Overview
Design, build, and maintain observability and reliability monitoring solutions for cloud-native systems. Turn operational telemetry (metrics, logs, traces) into actionable insights to improve system reliability, performance, and incident response.
Responsibilities
- Own the development and operation of observability platforms and workflows.
- Design, implement, and maintain observability across metrics, logging, tracing, dashboards, and alerting.
- Build and run scalable telemetry collection pipelines and monitoring infrastructure.
- Configure and optimize monitoring platforms such as Prometheus and Grafana, plus at least one commercial observability stack.
- Implement distributed tracing, structured logging, and OpenTelemetry-based solutions.
- Manage high-volume/high-cardinality metrics and log environments while maintaining performance and scalability.
- Create and maintain dashboards, alerts, and reporting for engineering teams.
- Support SRE practices using SLOs, SLIs, and error budgets.
- Integrate observability with CI/CD, incident management tools, and operational workflows.
- Analyze monitoring data to identify trends, improve performance, and reduce infrastructure costs.
- Collaborate with engineering and business stakeholders to translate telemetry into operational outcomes.
Requirements
- 5+ years of experience in SRE, platform engineering, reliability engineering, or observability roles.
- Hands-on experience with Prometheus and Grafana, and a commercial observability platform (e.g., Datadog, New Relic, Splunk).
- Strong understanding of OpenTelemetry, distributed tracing, telemetry pipelines, and structured logging.
- Proficiency in a programming language such as Go, Python, or Java.
- Experience managing high-throughput metrics and log processing systems.
- Solid knowledge of SRE principles including SLOs/SLIs and error budgets.
- Experience integrating observability with CI/CD and incident response processes.
- Strong Linux, networking, and container fundamentals.
- Excellent troubleshooting, analytical, communication, and collaboration skills.
Nice to Have
- Experience with observability components such as Thanos, Mimir, Cortex, Loki, or Tempo.
Scraped 7/29/2026