Staff DevOps Engineer
webAI
full-remoteleadpermanentdevopsbackend Full remote 73 days ago via WTTJ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
DevOpsStaff EngineerKubernetesTerraformAWSAzureGCPMLOpsCI/CDZero Trust
About the role
Role overview
Staff DevOps Engineer (full remote) at webAI. You will architect, build, and scale secure infrastructure for deploying AI workloads across cloud and edge environments. This is a staff-level individual contributor role acting as a subject matter expert in cloud architecture, security best practices, and platform reliability.
Responsibilities
- Architect, design, and operate secure infrastructure for deploying AI workloads across cloud and edge
- Design and run production Kubernetes clusters optimized for AI/ML workloads (GPU support)
- Implement container security, multi-tenancy, and resource optimization
- Lead MLOps infrastructure initiatives
- Model deployment pipelines, model/versioning and lifecycle management
- Feature stores, experiment tracking, and monitoring for performance and drift
- Drive platform reliability and infrastructure strategy through technical initiatives
- Build CI/CD pipelines with integrated security controls
- Apply security-by-design patterns (e.g., encryption, secrets management, IAM, network security)
- Use GitOps and declarative infrastructure management
Requirements
- Strong scripting/programming skills: Python (preferred) plus Bash or Go for automation
- 5+ years implementing Infrastructure as Code with Terraform, Ansible, or Pulumi, managing 50+ cloud resources
- Production observability/monitoring experience: Prometheus, Grafana, ELK, CloudWatch, Datadog (or similar)
- Deep cloud experience: AWS, Azure, or GCP (compute, networking, storage, managed services)
- MLOps workflows experience: deployment automation, versioning, lifecycle management
- Security best practices: encryption, secrets management, IAM, and network security
- Expert-level Docker and Kubernetes (CKA/CKAD preferred)
- GitOps experience and declarative infrastructure management
- 7+ years hands-on experience in DevOps/SRE/Infrastructure engineering with production system ownership
- Excellent communication for technical documentation and cross-functional collaboration
Nice-to-haves / Additional skills
- Experience with GitHub Actions, GitLab CI, Jenkins, ArgoCD (secure CI/CD)
- Multi-cloud/hybrid cloud architecture with portability and interoperability
- Experience deploying large language models (LLMs)/transformer models at scale
- Zero Trust architecture and modern security patterns
- Service mesh experience: Istio or Linkerd
- AI/ML infrastructure: feature stores, model registries, A/B testing infrastructure, monitoring
- Edge computing and distributed systems architectures
- Cost optimization / FinOps (rightsizing and cost management)
- Mentoring or leading technical initiatives
- Certifications: CKA/CKAD, Terraform Associate, AWS Solutions Architect, Azure Administrator, or GCP Professional Cloud Architect
About webAI
webAI builds an AI platform and supporting infrastructure for deploying AI workloads. The role focuses on secure, scalable operations across cloud and edge environments, including Kubernetes and MLOps capabilities.
Scraped 5/13/2026