xelys jobs xelys jobs

Data Scientist

Graphik Dimensions

middata United States Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

LLM EvaluationLLM-as-JudgePrompt InjectionRed TeamingRetrieval EvaluationAnomaly DetectionForecastingRobustness TestingAgentic SystemsGroundedness

About the role

Role overview

Data Scientist to join ScienceLogic’s Data Science team and improve a production suite of small, locally-hosted language models (not a single hosted frontier API). This role emphasizes LLM system evaluation and reliability—defining what “good” means for non-deterministic, resource-constrained models, building evaluation infrastructure, and turning interaction data into guidance for engineering and product.

Responsibilities

Evaluation & response quality

  • Design and own evaluation harnesses for LLM and agent outputs (golden sets, regression suites, rubric-based scoring).
  • Build and calibrate LLM-as-judge pipelines; validate judges against human labels and control for bias/variance.
  • Define and track response-quality metrics, including:
    • Faithfulness/groundedness
    • Hallucination rate
    • Answer relevance and completeness
    • Instruction-following
    • Persona adherence
  • Curate, version, and expand evaluation datasets as the product evolves.
  • Benchmark models in the suite against each other; measure the quality cost of smaller local models vs larger alternatives.

Adversarial & robustness testing

  • Conduct red-team testing (prompt injection, jailbreaks, tool misuse, and edge cases).
  • Design chaos/stress tests to probe reliability under degraded or hostile conditions.
  • Characterize failure modes and feed results into guardrails and regression coverage.

Retrieval & agentic trajectory analysis

  • Evaluate retrieval quality over the corpus (e.g., recall@k, MRR/nDCG, context precision/recall).
  • Run experiments on chunking, indexing, and hybrid retrieval strategies.
  • Analyze multi-step agent trajectories:
    • tool-call correctness
    • trajectory efficiency
    • replayable-state inspection
    • guardrail-breach behavior
  • Assess intent classification and routing as measurable components.

Behavioral regression & drift

  • Create standing evaluations to catch quality/behavioral regressions when models are swapped, upgraded, re-quantized, or prompts/pipelines change.
  • Monitor production output distribution and quality drift; distinguish true regressions from stochastic noise.
  • Recommend and validate fixes using the appropriate levers.

Requirements

  • Strong thinking in evaluation suites, failure modes, and groundedness for LLM systems.
  • Ability to improve reliable, high-quality behavior from small local models under real resource budgets.

Nice-to-haves

  • Experience with LLM-as-judge evaluation, retrieval metrics, adversarial/red-team testing, and monitoring for regression/drift in non-deterministic systems.

About Graphik Dimensions

ScienceLogic is an enterprise IT operations and AIOps company that helps organizations achieve Autonomic IT. Its AI/analytics platform delivers unified visibility across hybrid and multi-cloud environments, automates workflows, and improves performance at scale. The company focuses on automation, AI, and analytics with real enterprise security and compliance constraints.

Scraped 7/29/2026