xelys jobs xelys jobs

Staff Site Reliability Engineer

ServiceNow

Full remote - Dublin, IE Today via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

About the role

Join our team as a Staff Site Reliability Engineer, where you will design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness. You will work closely with engineering teams to improve platform reliability, release quality, and cloud-native adoption. This role requires strong software engineering skills, experience with Kubernetes, and a passion for continuous learning and automation. Key missions: Concevoir, construire et exploiter des plateformes d'ingénierie cloud-native pour la validation des logiciels, la validation des versions et la préparation à la production.. Développer des solutions d'automatisation qui améliorent la productivité des ingénieurs, rationalisent les opérations et réduisent le travail manuel.. Construire et intégrer des pipelines de test automatisés, des signaux d'observabilité, d'intelligence de déploiement et de qualité dans les workflows CI/CD. Profile: - Strong software engineering skills with hands-on experience designing, developing, testing, and debugging applications using Python, Go, Java, or Ruby - Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation - Demonstrated ability to solve complex technical problems, drive projects independently, and collaborate effectively across engineering teams - Experience leveraging AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, or operational automation is a plus - Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification - Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows - Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems - Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry - Thrives in fast-paced, ambiguous environments with a strong ownership mindset, bias for action, and a passion for continuous learning and automation - Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments - Experience building and operating cloud-native platforms supporting scalable, highly available services - 8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience - Experience designing and implementing automation to improve developer productivity, release quality, and operational efficiency - Low ego, intellectually curious, and an effective collaborator who enjoys partnering with globally distributed teams to deliver reliable engineering solutions - Experience with observability and monitoring platforms for applications, services, and distributed systems at scale - Experience with DevOps automation, CI/CD pipelines, GitOps, and Agile development practices using tools such as GitLab CI/CD, Argo CD, or Flux - Experience building and maintaining enterprise-scale test automation frameworks using technologies such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent - Experience with test orchestration, intelligent regression testing, test impact analysis, flaky test detection, parallel execution, and test data management - Experience with service virtualization, contract testing, synthetic testing, and building developer self-service engineering platforms - Experience with Infrastructure as Code and configuration management tools such as Ansible, Terraform, or equivalent - Experience with the Kubernetes ecosystem, including Helm, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtime technologies - Experience operating Kubernetes platforms across public cloud providers, including AWS (EKS), Azure (AKS), and Google Cloud (GKE) - Familiarity with AI-assisted engineering, intelligent testing, operational automation, or cloud-native engineering platforms - Experience implementing progressive delivery practices, including canary deployments, feature flags, deployment verification, and automated rollback

Scraped 9/25/2026