SWE-Bench Task Auditor
Undisclosed employer · US
Compensation
$70-$90/hr
Description
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.
Basic Qualifications • 3+ years professional software engineering • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles) • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar repository benchmarks • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.) • Prior code-review or task-grading experience
- Commitment
- Hourly
- Eligible locations
- USA
Skills & categories
- Posted
- Aug 28, 2026
- Slots remaining
- 3
- First seen
- Sep 2, 2026
- Last seen
- Sep 17, 2026