Data Scientist, AI/ML
Gremlin
seniorpermanentdata Marion County, IN 71 days ago via LinkedIn
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
Machine LearningAI/MLCausal InferenceGraph Machine LearningTime-Series ModelingReinforcement LearningFeature StoresData PipelinesMLOpsChaos Engineering
About the role
Role: Data Scientist, AI/ML
You will help improve the reliability of internet-scale systems by turning millions of chaos engineering experiments into automated failure analysis and remediation.
What you’ll do
- Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to find failure patterns, root causes, and resilience signals across distributed systems.
- Pretrain and fine-tune ML models to detect, classify, and explain failures from chaos experiments.
- Build systems that recommend (and eventually orchestrate) automated remediation by learning from historical experiment outcomes and system behavior.
- Develop scalable data pipelines and feature stores to process, enrich, and serve experiment data for both model training and real-time inference.
- Collaborate with platform engineers and SREs to integrate AI-driven failure analysis and remediation into Gremlin’s core product.
- Apply advanced ML techniques such as causal inference, graph ML, time-series modeling, and reinforcement learning to improve accuracy and actionability.
- Translate insights into AI-powered features for customers to understand blast radius, pinpoint root causes, and accelerate recovery.
- Research and productionize novel ML approaches (including causal AI and agentic systems) to create reliable remediation strategies.
What we’re looking for
- 5+ years professional experience building and productionizing machine learning (ideally for distributed systems, infrastructure, or DevOps/SRE use cases).
- Hands-on experience with causal inference, graph ML, time-series modeling, or reinforcement learning.
- Experience building data pipelines and feature stores supporting both offline training and real-time inference.
- Experience working in Agile environments.
- Strong experience with rigorous experimentation, model evaluation, and engineering best practices.
- Ability to partner with platform engineers and SREs to ship research into product features.
Bonus experience
- Experience with chaos engineering, site reliability engineering, or distributed systems.
- Background in agentic AI systems or large-scale causal inference in production.
- Experience standing up MLOps tooling (e.g., model serving, monitoring, or feature store infrastructure).
About Gremlin
Gremlin is a software company focused on reliability engineering, including chaos engineering and reliability testing. Its Reliability Platform helps software teams proactively monitor and test systems for reliability risks, enforce standards, and automate reliability practices across organizations.
Scraped 7/15/2026