xelys jobs xelys jobs

Site Reliability Engineer | $70/hr Remote

Crossing Hurdles

full-remotemidcontractdevopsbackend United States 89 days ago via LinkedIn
40,000 - 84,000 USD/annual

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringSREDockerKubernetesPythonBashCI/CDInfrastructureAutomationTroubleshooting

About the role

Role Overview

Site Reliability Engineer (Hourly Contract) — Remote (10–40 hours/week). You will deploy, monitor, and recover containerized AI training environments, ensuring stability and performance.

Responsibilities

  • Deploy, monitor, and recover containerized AI training environments
  • Troubleshoot infrastructure bottlenecks and resolve system failures in real time
  • Build resilient systems for stability and performance optimization
  • Collaborate with engineering teams to improve CI/CD pipelines and automation
  • Manage filesystem structures, storage, and process scheduling in containerized environments
  • Execute dynamic replanning during runtime issues and system failures
  • Document system processes, solutions, and best practices

Requirements

  • Strong experience with terminal-based system administration and troubleshooting
  • Expertise with containerized environments (Docker and/or Kubernetes)
  • Strong Python skills for scripting, automation, and debugging
  • Proficiency in Bash and familiarity with additional programming languages
  • Strong understanding of infrastructure, build systems, and version control
  • Ability to perform dynamic infrastructure recovery under high pressure
  • Excellent written and verbal communication skills

Scraped 4/28/2026