xelys jobs xelys jobs

Site Reliability Engineer (w/m/d)

IONOS SE

hybridmidpermanentdevopsbackend Karlsruhe 33 days ago via Arbeitnow

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Site Reliability EngineeringKubernetesTerraformGitLab CI/CDArgo CDHelmPrometheusGrafanaELK StackInfrastructure as Code

About the role

Role overview

You will work as a Site Reliability Engineer (SRE) in the Application Hosting team, forming the technical backbone of IONOS’s product platform for Managed Nextcloud, Nextcloud Workspace, IONOS GPT, and other web services running on a Kubernetes platform.

Responsibilities

  • Evolve the infrastructure/platform for the products and integrate new products/web services into the Kubernetes and cloud infrastructure
  • Ensure stable and secure operation of the product platform
  • Perform deep analysis and optimize primarily containerized, Kubernetes-based application infrastructure
  • Drive automation and manage infrastructure declaratively and reproducibly
    • Use Terraform, GitLab CI/CD, and Argo CD to provision and manage infrastructure
  • Analyze and resolve complex distributed-system issues and continuously improve the platform
  • Build and maintain Monitoring/Logging/Alerting for proactive detection of bottlenecks and failures (e.g., Prometheus, Grafana, ELK stack)

Requirements

  • Multi-year experience as an SRE or in a related role (Linux System Administrator, Platform Engineer, DevOps Engineer, Full Stack Developer) in Linux and Kubernetes environments
  • Strong, long-term hands-on experience with Linux, container technologies, and Kubernetes
  • Experience with Infrastructure as Code (preferably Terraform)
  • Experience with CI/CD pipelines (e.g., GitLab CI/CD or GitHub Actions)
  • Experience using Helm charts
  • Proficiency in at least one programming/scripting language for automation and monitoring tasks (e.g., Go, Python, Bash), and ideally some experience building operators
  • Experience operating and troubleshooting high-availability, distributed production systems, including monitoring, alerting, and log analysis
    • Examples: Prometheus, Grafana, FluentD, ELK, VictoriaMetrics, icinga
  • Proactive, solution-oriented and independent work style; ability to systematically analyze complex technical problems
  • German and English communication skills

Benefits / work model

  • Hybrid work model and flexible hours (trust-based working time)
  • Additional perks such as subsidized canteen (at some sites), modern office locations, employee discounts, events, workshops, training, and health offerings

About IONOS SE

IONOS SE is the leading European digitalization partner, providing cloud infrastructure, cloud services, and hosting for small and mid-sized businesses. It operates a global platform across multiple European markets and teams work on reliable, high-performance product services such as Managed Nextcloud and other web services.

Scraped 8/21/2026