xelys jobs xelys jobs

Staff DevOps Engineer

webAI

full-remoteleadpermanentdevopsbackend Full remote 73 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

DevOpsStaff EngineerKubernetesTerraformAWSAzureGCPMLOpsCI/CDZero Trust

About the role

Role overview

Staff DevOps Engineer (full remote) at webAI. You will architect, build, and scale secure infrastructure for deploying AI workloads across cloud and edge environments. This is a staff-level individual contributor role acting as a subject matter expert in cloud architecture, security best practices, and platform reliability.

Responsibilities

  • Architect, design, and operate secure infrastructure for deploying AI workloads across cloud and edge
  • Design and run production Kubernetes clusters optimized for AI/ML workloads (GPU support)
    • Implement container security, multi-tenancy, and resource optimization
  • Lead MLOps infrastructure initiatives
    • Model deployment pipelines, model/versioning and lifecycle management
    • Feature stores, experiment tracking, and monitoring for performance and drift
  • Drive platform reliability and infrastructure strategy through technical initiatives
  • Build CI/CD pipelines with integrated security controls
  • Apply security-by-design patterns (e.g., encryption, secrets management, IAM, network security)
  • Use GitOps and declarative infrastructure management

Requirements

  • Strong scripting/programming skills: Python (preferred) plus Bash or Go for automation
  • 5+ years implementing Infrastructure as Code with Terraform, Ansible, or Pulumi, managing 50+ cloud resources
  • Production observability/monitoring experience: Prometheus, Grafana, ELK, CloudWatch, Datadog (or similar)
  • Deep cloud experience: AWS, Azure, or GCP (compute, networking, storage, managed services)
  • MLOps workflows experience: deployment automation, versioning, lifecycle management
  • Security best practices: encryption, secrets management, IAM, and network security
  • Expert-level Docker and Kubernetes (CKA/CKAD preferred)
  • GitOps experience and declarative infrastructure management
  • 7+ years hands-on experience in DevOps/SRE/Infrastructure engineering with production system ownership
  • Excellent communication for technical documentation and cross-functional collaboration

Nice-to-haves / Additional skills

  • Experience with GitHub Actions, GitLab CI, Jenkins, ArgoCD (secure CI/CD)
  • Multi-cloud/hybrid cloud architecture with portability and interoperability
  • Experience deploying large language models (LLMs)/transformer models at scale
  • Zero Trust architecture and modern security patterns
  • Service mesh experience: Istio or Linkerd
  • AI/ML infrastructure: feature stores, model registries, A/B testing infrastructure, monitoring
  • Edge computing and distributed systems architectures
  • Cost optimization / FinOps (rightsizing and cost management)
  • Mentoring or leading technical initiatives
  • Certifications: CKA/CKAD, Terraform Associate, AWS Solutions Architect, Azure Administrator, or GCP Professional Cloud Architect

About webAI

webAI builds an AI platform and supporting infrastructure for deploying AI workloads. The role focuses on secure, scalable operations across cloud and edge environments, including Kubernetes and MLOps capabilities.

Scraped 5/13/2026