xelys jobs xelys jobs

DevOps Team Lead

Alpaca

full-remoteleadpermanentengineering-managementdevopsbackend Full remote 73 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

DevOpsSREEngineering ManagementKubernetesTerraformPostgreSQLPrometheusGrafanaChange ManagementIncident Management

About the role

Role overview

Join Alpaca as a DevOps Team Lead to lead a globally distributed engineering team responsible for critical infrastructure initiatives. This leadership role combines people leadership with technical planning and operational leadership, including roadmap delivery, change management, and incident/on-call management.

Key missions

  • Lead and manage a globally distributed team to ensure high performance and effective collaboration.
  • Own delivery of complex yearly roadmap items for foundational infrastructure upgrades, focused on high availability and platform resilience.
  • Design and refine robust support workflows, agile planning methodologies, and deployment/rollout strategies to ensure operational excellence.

Responsibilities & focus areas

  • People and tech leadership (not a hands-on daily engineering role).
  • Prioritization and planning for multi-stage infrastructure initiatives.
  • Change Management lifecycle ownership.
  • Building and maintaining support frameworks and methodologies.
  • On-call management and incident management lifecycle leadership.

Requirements

  • Strong communication and organizational skills; ability to coordinate complex multi-stage rollouts and deployments.
  • Proven experience as an Engineering Manager, DevOps Lead, or SRE Lead, including managing globally distributed teams.
  • Strategic mindset to navigate shifting priorities and act as an organizational anchor for core infrastructure.
  • Deep expertise in engineering support frameworks, roadmap planning, and team prioritization methodologies.
  • Proven experience owning Change Management lifecycles.
  • Extensive experience managing Incident Management lifecycles and running sustainable global on-call rotations.
  • Exceptional people management skills: coaching, mentoring, and fostering culture across time zones.
  • Solid technical background in modern DevOps/SRE ecosystems (must fluently understand operational realities).

Nice-to-have / technical fluency (required to understand)

  • Kubernetes (GKE)
  • Infrastructure as Code (Terraform)
  • Relational Databases (PostgreSQL)
  • Observability: Prometheus, Grafana, Thanos

About Alpaca

Alpaca is a technology company operating in the infrastructure/DevOps space, focused on building and running critical systems. The role emphasizes leading engineering teams responsible for platform resilience, high availability, and operational excellence in a globally distributed environment.

Scraped 5/13/2026