Data Scientist
Graphik Dimensions
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role overview
Data Scientist to join ScienceLogic’s Data Science team and improve a production suite of small, locally-hosted language models (not a single hosted frontier API). This role emphasizes LLM system evaluation and reliability—defining what “good” means for non-deterministic, resource-constrained models, building evaluation infrastructure, and turning interaction data into guidance for engineering and product.
Responsibilities
Evaluation & response quality
- Design and own evaluation harnesses for LLM and agent outputs (golden sets, regression suites, rubric-based scoring).
- Build and calibrate LLM-as-judge pipelines; validate judges against human labels and control for bias/variance.
- Define and track response-quality metrics, including:
- Faithfulness/groundedness
- Hallucination rate
- Answer relevance and completeness
- Instruction-following
- Persona adherence
- Curate, version, and expand evaluation datasets as the product evolves.
- Benchmark models in the suite against each other; measure the quality cost of smaller local models vs larger alternatives.
Adversarial & robustness testing
- Conduct red-team testing (prompt injection, jailbreaks, tool misuse, and edge cases).
- Design chaos/stress tests to probe reliability under degraded or hostile conditions.
- Characterize failure modes and feed results into guardrails and regression coverage.
Retrieval & agentic trajectory analysis
- Evaluate retrieval quality over the corpus (e.g., recall@k, MRR/nDCG, context precision/recall).
- Run experiments on chunking, indexing, and hybrid retrieval strategies.
- Analyze multi-step agent trajectories:
- tool-call correctness
- trajectory efficiency
- replayable-state inspection
- guardrail-breach behavior
- Assess intent classification and routing as measurable components.
Behavioral regression & drift
- Create standing evaluations to catch quality/behavioral regressions when models are swapped, upgraded, re-quantized, or prompts/pipelines change.
- Monitor production output distribution and quality drift; distinguish true regressions from stochastic noise.
- Recommend and validate fixes using the appropriate levers.
Requirements
- Strong thinking in evaluation suites, failure modes, and groundedness for LLM systems.
- Ability to improve reliable, high-quality behavior from small local models under real resource budgets.
Nice-to-haves
- Experience with LLM-as-judge evaluation, retrieval metrics, adversarial/red-team testing, and monitoring for regression/drift in non-deterministic systems.
About Graphik Dimensions
ScienceLogic is an enterprise IT operations and AIOps company that helps organizations achieve Autonomic IT. Its AI/analytics platform delivers unified visibility across hybrid and multi-cloud environments, automates workflows, and improves performance at scale. The company focuses on automation, AI, and analytics with real enterprise security and compliance constraints.
Scraped 7/29/2026