xelys jobs xelys jobs

Site Reliability Engineer

Aalyria

seniorpermanentbackend United States 33 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability Engineering (SRE)ObservabilityPrometheusOpenTelemetryGrafanaLokiTempoKubernetesTerraformGitOps

About the role

Role Overview

Build a strategic, production-grade observability and reliability platform for satellite and deep-space network operations. This is a greenfield/brownfield SRE role focused on creating the “nervous system” for a platform that orchestrates interconnected satellite, ground, and fleet networks.

Key Responsibilities

  • Design and build a centralized observability platform integrating and scaling:
    • Metrics (e.g., Prometheus)
    • Logging (e.g., Loki)
    • Distributed tracing (e.g., Tempo/OpenTelemetry)
  • Define, implement, and manage SLOs, SLIs, and error budgets to ensure services are launch-ready.
  • Partner with software engineers to apply observability best practices, including:
    • standard templates and documentation
    • configuring OpenTelemetry libraries
  • Automate deployment, scaling, and management of the observability stack using:
    • Infrastructure as Code (e.g., Terraform)
    • GitOps (e.g., ArgoCD)
  • Work with the infrastructure team to deliver deep visibility into:
    • Kubernetes clusters
    • underlying GCP and AWS environments
  • Own and lead monitoring, alerting, and incident response strategy; drive proactive reliability and blameless post-mortems.

Requirements

  • 4+ years in SRE or platform engineering with a focus on observability for large-scale distributed compute/network systems.
  • Deep, hands-on experience building and scaling observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb).
  • Production experience with GCP and Kubernetes.
  • Experience with Infrastructure as Code and GitOps (e.g., ArgoCD).
  • Strong systems programming skills with a preference for Go and Python.
  • Proven experience defining and managing SLOs/SLIs/error budgets for high-availability distributed production services.

Preferred Qualifications

  • Multi-cloud operations experience across GCP and AWS.
  • Hands-on GitLab CI experience for CI/CD pipelines.
  • Knowledge of service mesh technologies (e.g., Istio or Linkerd).
  • Familiarity with instrumenting apps written in Go and C++.
  • Active Secret clearance or higher.
  • Experience with JVM observability (tuning/monitoring) for Java applications.

Additional Notes

  • Includes on-call responsibilities.

About Aalyria

Aalyria is a technology company providing laser communications technology and temporospatial software-defined networking platforms for the aerospace industry. It focuses on satellite and airborne mesh networks, including cislunar and deep-space communications, and is building orchestration and observability capabilities for planetary mesh networks. The company was formed around technology acquired from Google.

Scraped 6/23/2026