xelys jobs xelys jobs

Reliability Monitoring Engineer

Jobgether

seniorpermanentbackend United States Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

ObservabilitySREReliability EngineeringOpenTelemetryDistributed TracingPrometheusGrafanaDatadogDatadogStructured Logging

About the role

Role Overview

Design, build, and maintain observability and reliability monitoring solutions for cloud-native systems. Turn operational telemetry (metrics, logs, traces) into actionable insights to improve system reliability, performance, and incident response.

Responsibilities

  • Own the development and operation of observability platforms and workflows.
  • Design, implement, and maintain observability across metrics, logging, tracing, dashboards, and alerting.
  • Build and run scalable telemetry collection pipelines and monitoring infrastructure.
  • Configure and optimize monitoring platforms such as Prometheus and Grafana, plus at least one commercial observability stack.
  • Implement distributed tracing, structured logging, and OpenTelemetry-based solutions.
  • Manage high-volume/high-cardinality metrics and log environments while maintaining performance and scalability.
  • Create and maintain dashboards, alerts, and reporting for engineering teams.
  • Support SRE practices using SLOs, SLIs, and error budgets.
  • Integrate observability with CI/CD, incident management tools, and operational workflows.
  • Analyze monitoring data to identify trends, improve performance, and reduce infrastructure costs.
  • Collaborate with engineering and business stakeholders to translate telemetry into operational outcomes.

Requirements

  • 5+ years of experience in SRE, platform engineering, reliability engineering, or observability roles.
  • Hands-on experience with Prometheus and Grafana, and a commercial observability platform (e.g., Datadog, New Relic, Splunk).
  • Strong understanding of OpenTelemetry, distributed tracing, telemetry pipelines, and structured logging.
  • Proficiency in a programming language such as Go, Python, or Java.
  • Experience managing high-throughput metrics and log processing systems.
  • Solid knowledge of SRE principles including SLOs/SLIs and error budgets.
  • Experience integrating observability with CI/CD and incident response processes.
  • Strong Linux, networking, and container fundamentals.
  • Excellent troubleshooting, analytical, communication, and collaboration skills.

Nice to Have

  • Experience with observability components such as Thanos, Mimir, Cortex, Loki, or Tempo.

Scraped 7/29/2026