xelys jobs xelys jobs

Site Reliability Engineer

Offchain

nullmidpermanentdevops United States 50 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringKubernetesTerraformGitOpsArgoCDPrometheusLokiGrafanaLinuxAWS

About the role

Role: Site Reliability Engineer

You’ll help operate and improve production systems that power blockchain scaling infrastructure, with a strong focus on reliability, security, and automation.

Responsibilities

  • Operate production Kubernetes clusters and maintain scalable, declarative infrastructure.
  • Build and improve deployment automation and GitOps-style delivery workflows.
  • Design and run observability for reliability (metrics, logs, dashboards), and use it to troubleshoot issues.
  • Participate in on-call rotation: respond to incidents, troubleshoot under pressure, and drive postmortems to improve reliability.
  • Diagnose complex networking and storage issues across distributed systems.
  • Implement secure-by-default infrastructure and contribute to architecture reviews and threat modeling.

Requirements

  • Experience with GitOps-style systems and treating infrastructure/application delivery as code.
  • Comfortable operating within cloud platforms (AWS, GCP, or Azure) and understanding underlying components.
  • Strong Linux and shell scripting skills; productive in Python or Go.
  • Ability to work with YAML, logs, and low-level debugging.
  • On-call experience with incident response and postmortem-driven improvements.

Nice-to-haves

  • Experience using tools like k9s and ArgoCD (e.g., ArgoCD ApplicationSets or similar).
  • Building CI/CD workflows with tools such as ArgoCD, GitHub Actions, or CodeBuild.
  • Observability experience with Prometheus, Loki, Mimir, Grafana, and CloudWatch.
  • Use of Terraform (or similar) for infrastructure as code.

What You’ve Done (signals)

  • Operated Kubernetes in production and built infrastructure with Terraform or equivalent.
  • Deployed and maintained Kubernetes environments.
  • Designed CI/CD workflows spanning both infrastructure and application deployments.
  • Implemented observability and diagnosed challenging networking/storage issues.
  • Built secure-by-default systems and contributed to threat modeling/architecture reviews.

About Offchain

Offchain is a blockchain infrastructure company focused on scalability and security, helping power decentralized applications at scale. It is closely associated with the Arbitrum ecosystem, including the Arbitrum stack that supports Arbitrum One, and has been adopted by many projects and teams across Ethereum. Offchain is backed by substantial funding and operates infrastructure that processes millions of transactions.

Scraped 6/11/2026