Site Reliability Engineer | $70/hr Remote
Crossing Hurdles
full-remotemidcontractdevopsbackend United States 89 days ago via LinkedIn
40,000 - 84,000 USD/annual
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Site Reliability EngineeringSREDockerKubernetesPythonBashCI/CDInfrastructureAutomationTroubleshooting
About the role
Role Overview
Site Reliability Engineer (Hourly Contract) — Remote (10–40 hours/week). You will deploy, monitor, and recover containerized AI training environments, ensuring stability and performance.
Responsibilities
- Deploy, monitor, and recover containerized AI training environments
- Troubleshoot infrastructure bottlenecks and resolve system failures in real time
- Build resilient systems for stability and performance optimization
- Collaborate with engineering teams to improve CI/CD pipelines and automation
- Manage filesystem structures, storage, and process scheduling in containerized environments
- Execute dynamic replanning during runtime issues and system failures
- Document system processes, solutions, and best practices
Requirements
- Strong experience with terminal-based system administration and troubleshooting
- Expertise with containerized environments (Docker and/or Kubernetes)
- Strong Python skills for scripting, automation, and debugging
- Proficiency in Bash and familiarity with additional programming languages
- Strong understanding of infrastructure, build systems, and version control
- Ability to perform dynamic infrastructure recovery under high pressure
- Excellent written and verbal communication skills
Scraped 4/28/2026