AI Benchmark Quality Reviewer
Turing · AR / Costa Rica, BR
Compensation
$23/hr
Description
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.
What you’ll do
• Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
• Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
• Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
• Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.
• Write concise, evidence-backed findings and verify that revisions address the issues found.
What we’re looking for
• At least five years of relevant technical or analytical experience.
• Ability to read Python, SQL, shell scripts, structured data and execution logs to understand task setup and grading behavior.
• Strong written English, analytical judgment and attention to detail.
• Ability to give specific, reproducible feedback and explain uncertainty clearly.
• Experience in AI evaluation, technical QA, data analysis or benchmark development is helpful, but not required. Familiarity with Harbor task setup is a plus.
Engagement
• Remote contractor assignment for eight weeks.
• 40 hours per week, including eight hours of daily overlap with Pacific Time.
• Shortlisted applicants may be asked to complete an interest form before final review.
Compensation
• $23 per hour.
- Contract
- Short
- Engagement
- Freelance
Skills & categories
- Posted
- Sep 23, 2026
- Slots remaining
- 25
- First seen
- Sep 23, 2026
- Last seen
- Sep 24, 2026