xelys jobs xelys jobs

Director of Cloud Operations

Firstup

full-remoteleadpermanentdevopsengineering-managementsecurity Full remote 74 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

AWSTerraformKubernetesEKSSREDevOpsSLO/SLIObservabilityCI/CDReliability Engineering

About the role

Role Overview

You will lead and evolve Firstup’s cloud infrastructure and operational practices for a globally distributed SaaS platform. This is a hands-on leadership role focused on reliability, scalability, and efficiency across multiple AWS regions in the US and Europe.

Key Missions

  • Lead cloud infrastructure and CloudOps/SRE operational practices across a global SaaS environment.
  • Improve system reliability using SLIs/SLOs, error budgets, and proactive engineering practices.
  • Reduce MTTR and strengthen incident response effectiveness.
  • Enhance operational excellence through better observability and continuous improvement in how services are built and run.
  • Lead and mentor a distributed engineering team across the US and UK, fostering a high-performing, collaborative, growth-oriented culture.

Responsibilities

  • Partner with Engineering, Security, and Product to strengthen operational excellence.
  • Drive reliability engineering practices across services, including incident management and on-call operations.
  • Balance deep technical work with leadership, guidance, and mentoring.

Requirements

  • 10+ years in cloud infrastructure, SRE, or DevOps, including recent experience leading CloudOps/SRE teams.
  • Proven ability to lead operational/platform transformations in a SaaS environment.
  • Experience operating multi-region, customer-facing systems at scale.
  • Strong understanding of microservices and distributed systems design.
  • Infrastructure as Code (preferred: Terraform).
  • CI/CD pipeline experience (e.g., CircleCI or similar).
  • Hands-on experience with:
    • AWS multi-region architectures
    • Serverless and modern cloud-native patterns
    • Kubernetes (EKS) and containerized environments
    • Observability platforms (e.g., Datadog)
    • SLO/SLI frameworks, monitoring strategies, and performance optimization
  • Deep experience with incident management, on-call operations, and reliability engineering.
  • Ability to influence across teams and functions with a pragmatic, measurable-outcomes mindset.

Nice to Have

  • Additional experience with modern observability and performance optimization tooling beyond Datadog-equivalents.

About Firstup

Firstup is a SaaS company building a globally distributed platform for customer-facing services. The company focuses on reliable operations and continuous improvement across cloud infrastructure, security, and engineering practices.

Scraped 5/12/2026