Kairos
Back to gigs

AI Benchmark Quality Reviewer

Remote

Turing · AR / Costa Rica, BR

Fresh
Gig
Braintrust

Compensation

$23/hr

Apply on Braintrust

Description

Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.


What you’ll do

• Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.

• Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.

• Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.

• Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues.

• Write concise, evidence-backed findings and verify that revisions address the issues found.


What we’re looking for

• At least five years of relevant technical or analytical experience.

• Ability to read Python, SQL, shell scripts, structured data and execution logs to understand task setup and grading behavior.

• Strong written English, analytical judgment and attention to detail.

• Ability to give specific, reproducible feedback and explain uncertainty clearly.

• Experience in AI evaluation, technical QA, data analysis or benchmark development is helpful, but not required. Familiarity with Harbor task setup is a plus.


Engagement

• Remote contractor assignment for eight weeks.

• 40 hours per week, including eight hours of daily overlap with Pacific Time.

• Shortlisted applicants may be asked to complete an interest form before final review.


Compensation

• $23 per hour.

Contract
Short
Engagement
Freelance

Skills & categories

PythonSQL
Posted
Sep 23, 2026
Slots remaining
25
First seen
Sep 23, 2026
Last seen
Sep 24, 2026