Generalist Annotator — Health AI Conversation Quality Evaluation
Undisclosed employer · — worldwide
Compensation
$20-$160/hr
Description
About the Role
Mercor is hiring generalist annotators to evaluate conversations between users and a Health & Fitness Assistant AI. You will judge how well the AI communicates: whether it answered the actual question, whether a non-expert could follow it, and whether it was honest even when honesty was the harder answer.
This is a non-clinical role. You do not need medical training, and you should not try to supply it — clinicians assess accuracy and safety on the same conversations, as a separate track. Your value is precisely that you are not one of them: you can say how a real person would experience the response, which a clinician reading the same text cannot.
What you will do
- Relevance and clarity — did the AI answer what was asked, in language a family member with no medical background could act on?
- Sycophancy detection — the core of this work. Spotting where the AI told someone what they wanted to hear instead of what they needed to hear: backing down when a user pushed back, hiding behind "you know your body best", or framing a risky decision as empowering rather than naming the risk. Automated evaluators reliably score this as helpful, which is why it needs a person.
- Tone and outcome — whether the conversation left the user informed and calmer, or anxious and no better off, and whether they walked away with anything concrete to do.
- Written justification — every rating below the top option requires a specific, quotable explanation of what went wrong and what a better response would have said. The ratings tell us something is wrong; only your writing tells us what to fix.
Conversations run from 1 to 7 turns; many are a single turn. Expect roughly 2–13 minutes each depending on length. Full written instructions and worked examples are provided before you start.
Required qualifications
- Strong written English, and the ability to explain a judgment in two to four specific sentences rather than a one-line verdict
- Careful reading — the work rewards noticing what a response quietly left out, not just what it got wrong
- Comfort applying a detailed rubric consistently across many items
- Reliable availability, with the ability to concentrate hours when a batch is time-boxed
Preferred qualifications
- Prior annotation, human-feedback, evaluation, or content-review experience
- A background where you have had to explain something technical to a non-expert audience — teaching, editing, writing, customer support, patient advocacy
- Experience with rubric-based grading or quality assurance work
Important
Using an AI tool to write your comments is prohibited and will end your work on this project. This project exists to capture human judgment that AI systems lack; AI-written feedback corrupts the dataset we use to check those systems. It is checked for.
Why this work
Most health-AI failures are not exotic. They are ordinary answers that sound responsible, read as warm and careful, and still leave someone worse off than if they had never asked. Automated evaluators are good at spotting responses that look careful, which is exactly why they miss these. Catching them takes a person who read the whole conversation and noticed where it quietly went wrong.
- Commitment
- Task-based
- Posted
- Sep 11, 2026
- Slots remaining
- 38
- First seen
- Sep 18, 2026
- Last seen
- Sep 19, 2026