QA/Test Engineer
NielsenIQ
See how well this job matches your profile
Sign up to get an AI match score and generate a tailored application in seconds.
Get your match scoreTags
About the role
Role overview
You will serve as a QA/Test Engineer, acting as the quality backbone for a lab building agentic evaluation benchmarks. Each benchmark task is complex and multi-step, requiring rigorous test design and review to ensure tasks are unambiguous, correctly graded, and robust to shortcuts.
This is a full-time W-2 role (via Cincinnatus LLC) with the opportunity to be placed with a leading AI lab as part of their extended workforce. The role is fully remote within the United States (~35 hours/week).
Responsibilities
- Design checks: Create test cases that verify each task works as intended, including difficult edge cases.
- Review tasks: Carefully review tasks and reference solutions to catch ambiguity and gaps before finalization.
- Debug: Use Python to debug tasks and quality checks when behavior doesn’t match expectations.
- Shape the process: Help build repeatable quality checklists and provide actionable feedback to task authors.
- Protect results: Identify and prevent shortcuts and grading gaps in AI agent runs so benchmark scores remain trustworthy.
- Collaborate tightly: Work in a fast feedback loop with researchers and task authors.
Requirements
- Education/experience: MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.
- Experience: 5+ years in test engineering, QA, or a research/software engineering role with strong quality ownership.
- Testing & review: Demonstrated ability to design test cases and quality-review processes.
- Debugging: Ability to debug complex systems end-to-end.
- Tools: Proficiency in Python and Git; comfort navigating unfamiliar codebases/environments.
- Documentation: Exceptional attention to detail and strong written documentation habits.
- Work style: Independent, able to tackle ambiguous/open-ended problems.
- Availability: Reliable to work ~35 hours/week.
Nice to have
- Experience in AI training, model evaluation, or quality review of AI-generated work.
About NielsenIQ
The posting references NielsenIQ, a company operating in data analytics and measurement, supporting businesses with insights across marketing and consumer-related intelligence. The role is positioned within an AI lab context focused on building agentic evaluation benchmarks for frontier models.
Scraped 7/26/2026