xelys jobs xelys jobs

Database Reliability Engineer

ClickHouse

full-remoteseniorpermanentbackend Full remote 73 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Database Reliability EngineeringClickHouseDistributed SystemsSQLProduction DebuggingIncident ResponsePost-mortemsPythonShellAWS

About the role

Role overview

Join ClickHouse as a Database Reliability Engineer to improve the reliability, availability, scalability, and performance of ClickHouse. You’ll help shape ClickHouse Cloud operations through incident management, investigations, and continuous improvement.

Key missions

  • Collaborate with multiple teams to implement ClickHouse for customers and provide best-practice guidance.
  • Own engineering escalation management and incident response (triage, investigations, and follow-ups).
  • Drive post-mortem analysis and continuous improvement of reliability and performance.
  • Enhance incident response processes for ClickHouse core.

Requirements

  • Strong problem-solving skills and solid production debugging ability.
  • Strong understanding of distributed database internals and SQL (ClickHouse experience is a major plus).
  • Experience operating ClickHouse or other SQL databases in production.
  • High ownership and accountability; comfortable working in a fast-paced global environment.
  • Degree in Computer Science or a related field.
  • Scripting skills with Shell or Python, and ability to read/understand C++ code.
  • At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering.
  • Excellent communication skills.

Nice to have

  • Knowledge of cloud computing platforms such as AWS, Azure, or GCP.

About ClickHouse

ClickHouse is a leading company in the database industry, known for its high-performance analytical database technology. The role is focused on improving the reliability and operational excellence of ClickHouse and ClickHouse Cloud through cross-team collaboration and incident-driven continuous improvement.

Scraped 5/13/2026