xelys jobs xelys jobs

QA/Test Engineer

NielsenIQ

full-remoteseniorpermanentqa American Canyon, CA 2 days ago via LinkedIn

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

QA TestingTest EngineeringPythonGitAI Model EvaluationDebuggingTest Case DesignQuality AssuranceEdge Case TestingBenchmarking

About the role

Role overview

You will serve as a QA/Test Engineer, acting as the quality backbone for a lab building agentic evaluation benchmarks. Each benchmark task is complex and multi-step, requiring rigorous test design and review to ensure tasks are unambiguous, correctly graded, and robust to shortcuts.

This is a full-time W-2 role (via Cincinnatus LLC) with the opportunity to be placed with a leading AI lab as part of their extended workforce. The role is fully remote within the United States (~35 hours/week).

Responsibilities

  • Design checks: Create test cases that verify each task works as intended, including difficult edge cases.
  • Review tasks: Carefully review tasks and reference solutions to catch ambiguity and gaps before finalization.
  • Debug: Use Python to debug tasks and quality checks when behavior doesn’t match expectations.
  • Shape the process: Help build repeatable quality checklists and provide actionable feedback to task authors.
  • Protect results: Identify and prevent shortcuts and grading gaps in AI agent runs so benchmark scores remain trustworthy.
  • Collaborate tightly: Work in a fast feedback loop with researchers and task authors.

Requirements

  • Education/experience: MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.
  • Experience: 5+ years in test engineering, QA, or a research/software engineering role with strong quality ownership.
  • Testing & review: Demonstrated ability to design test cases and quality-review processes.
  • Debugging: Ability to debug complex systems end-to-end.
  • Tools: Proficiency in Python and Git; comfort navigating unfamiliar codebases/environments.
  • Documentation: Exceptional attention to detail and strong written documentation habits.
  • Work style: Independent, able to tackle ambiguous/open-ended problems.
  • Availability: Reliable to work ~35 hours/week.

Nice to have

  • Experience in AI training, model evaluation, or quality review of AI-generated work.

About NielsenIQ

The posting references NielsenIQ, a company operating in data analytics and measurement, supporting businesses with insights across marketing and consumer-related intelligence. The role is positioned within an AI lab context focused on building agentic evaluation benchmarks for frontier models.

Scraped 7/26/2026