xelys jobs xelys jobs

Data Scientist

ScienceLogic

middata United States Yesterday via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

LLM EvaluationLLM-as-judgeRetrieval-Augmented Generation (RAG)Adversarial TestingPrompt InjectionAgentic SystemsAnomaly DetectionTime Series ForecastingGroundednessRegression Testing

About the role

Role: Data Scientist (LLM Evaluation & Operational Intelligence)

You’ll join ScienceLogic’s Data Science team to improve the quality and reliability of a suite of small, locally-hosted language models in production. Instead of classical predictive modeling, you’ll measure and optimize the behavior of the LLM system itself (answers, retrieval, multi-step agent behavior, and robustness under adversarial conditions).

Responsibilities

Evaluation & Response Quality

  • Design and own evaluation harnesses for LLM and agent outputs (golden sets, regression suites, rubric-based scoring).
  • Build and calibrate LLM-as-judge pipelines; validate judges against human labels and manage judge bias/variance.
  • Define and track response-quality metrics, including:
    • Faithfulness/groundedness and hallucination rate
    • Answer relevance and completeness
    • Instruction following and persona adherence
  • Curate, version, and expand evaluation datasets as product surfaces evolve.
  • Benchmark models against each other to route tasks to the best model; quantify the quality cost of small local models vs larger alternatives.

Adversarial & Robustness Testing

  • Run red-team testing: prompt injection, jailbreaks, tool misuse, and edge-case discovery.
  • Create chaos/stress tests to probe reliability under degraded or hostile conditions.
  • Characterize failure modes and feed them into guardrails and regression coverage.

Retrieval & Agentic Trajectory Analysis

  • Evaluate retrieval quality over the document corpus (e.g., recall@k, MRR/nDCG, context precision/recall).
  • Run experiments on chunking, indexing, and hybrid retrieval strategies.
  • Analyze multi-step agent trajectories:
    • Tool-call correctness
    • Trajectory efficiency
    • Replayable-state inspection
    • Guardrail-breach behavior
  • Assess intent classification and routing quality as measurable components.

Behavioral Regression & Drift

  • Build standing evaluations to detect quality and behavioral regressions when models/pipelines change (swap, upgrade, re-quantize, prompt changes).
  • Monitor output-distribution and quality drift in production and distinguish true regressions from stochastic noise.
  • Recommend and validate fixes using appropriate levers in the system.

About ScienceLogic

ScienceLogic builds an AIOps platform for modern enterprises, enabling Autonomic IT where systems are self-healing, self-optimizing, and aligned with business outcomes. The platform provides unified visibility across hybrid and multi-cloud environments, automates workflows, and uses AI and analytics to improve performance at scale.

Scraped 7/29/2026