
Compensation
Salary undisclosedDescription
Role Overview
We are seeking detail-oriented evaluators to conduct human quality evaluations for an enterprise AI customer support product on Instagram, WhatsApp, and Messenger. You will evaluate and benchmark AI model responses against complex evaluation rubrics using provided business knowledge bases.
What You’ll Own:
- Model Evaluation: Review and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions (e.g., Action Fidelity, Faithfulness, Hallucination, Compliance, Tone, and Handoff).
- Intent & Fact Verification: Benchmark both informational (R1) and transactional (R2) customer queries against authoritative business sources (FAQs, product catalogs, SOPs) within the task UI.
- Quality Assurance: Participate in dual-review processes and daily calibration audits to ensure inter-rater agreement and establish ground-truth performance targets.
- Performance Targets: Deliver precise evaluation
You’ll Thrive in This Role If You Have:
- Customer Service Background: Prior experience in customer service, call centers, retail, or handling customer communications via email, chat, or phone (highly prioritized).
- English Proficiency: Exceptional written English skills with a strong command of tone, brand voice, grammar, and nuance.
- Vietnamese Proficiency: Exceptional written Vietnamese skills with a strong command of tone, brand voice, grammar, and nuance.
- Analytical Precision: Ability to strictly follow multi-tier evaluation guidelines, complex logic trees, and technical rubrics without deviation.
- Tech Adaptability: Comfort using dedicated web-based tools and labeling interfaces.
Work Environment & Schedule:
● Duration: Full-Time, 2-Month
The expected hourly salary range for this position is up to $20 p/hour, based on experience, skills, and qualifications.
- Posted
- Aug 26, 2026
- Last seen
- Aug 27, 2026
- First seen
- Aug 27, 2026
