xelys jobs xelys jobs

Director of Data & Storage Reliability Engineering

ServiceNow

Full remote - Santa Clara, CA, US Today via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

About the role

Join ServiceNow as the Director of Data & Storage Reliability Engineering. In this role, you will lead the organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure. You will build and develop high-performing engineering teams, establish a strong engineering-first culture, and drive engineering improvements that eliminate issues before they impact customers. You will also partner closely with various teams to ensure reliability and performance considerations are incorporated throughout the software development lifecycle. Key missions: Lead the Data & Storage Reliability Engineering organization to improve reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure.. Build, develop, and scale high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering.. Establish a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, customer experience, and systemic risk reduction. Profile: - 8+ years of engineering leadership experience, including leading managers and globally distributed teams - Experience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams to drive systemic improvements based on operational learnings - Extensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizations - Experience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes - 15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environments - Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements - Strong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomes - Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry - Experience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomes - Strong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineering - Experience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments and long-term roadmaps - Experience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability, performance, scalability, and resilience objectives - Deep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architectures - Experience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scale - Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience - Exceptional communication, stakeholder management, and leadership skills - Experience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authority - Experience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomes - Previous Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environment - Experience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizations - Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads - Experience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizations - Experience with observability platforms, telemetry systems, diagnostics frameworks, and production analytics - Experience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformations - Experience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivity - Strong understanding of distributed systems architecture, cloud platform operations, and hyperscale environments - Experience developing executive-facing reliability scorecards, engineering metrics, and business impact reporting - Experience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmaps - Experience with Linux-based production environments and large-scale cloud infrastructure - Experience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms - Familiarity with ServiceNow platform architecture and large-scale SaaS operations

Scraped 9/25/2026