xelys jobs xelys jobs

Database Reliability Engineer

ClickHouse

full-remoteseniorpermanentbackenddevops Full remote 73 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

Database Reliability EngineeringClickHouseSQLDistributed SystemsAWSAzureGCPIncident ResponsePost-MortemsProduction Debugging

About the role

Role overview

As a Database Reliability Engineer at ClickHouse, you ensure the reliability, availability, scalability, and performance of ClickHouse (including ClickHouse Cloud). You’ll work with multiple teams, lead on incident escalation and investigations, and drive continuous reliability improvements.

Responsibilities

  • Collaborate with different teams to implement ClickHouse in the best way for customers
  • Own engineering escalation management, response, investigations, and post-mortem analysis
  • Continuously improve reliability and performance of ClickHouse core
  • Enhance incident response processes and post-mortem practices

Requirements

  • Knowledge of cloud platforms such as AWS, Azure, or GCP
  • Bachelor’s or Master’s degree in Computer Science or related field
  • Strong problem-solving skills and solid production debugging experience
  • Ability to thrive in a fast-paced, global environment; communicate effectively
  • Demonstrated ownership and accountability
  • Previous experience operating ClickHouse or other SQL databases in production
  • Excellent understanding of distributed database internals and SQL (ClickHouse is a major plus)
  • Scripting experience with Shell or Python and ability to read/understand C++ code
  • At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering

Nice to have

  • Previous hands-on operations experience specifically with ClickHouse

About ClickHouse

ClickHouse is a company focused on database technology, known for the ClickHouse analytics database and ClickHouse Cloud. They help customers run high-performance, scalable data workloads and rely on strong engineering to ensure reliability and performance.

Scraped 5/13/2026