Kairos
Back to jobs

Member of Technical Staff, Agents

Remote
CleraPalo Alto, CA, US22 hours agoWebsite
Fresh
Full-time
Engineering

Compensation

Salary undisclosed
Apply
Share

Description

About the Role

This is a hands-on engineering role focused on building and improving autonomous AI agents that handle real executive work end to end, from scheduling and emails to phone calls and relationship management. You will sit at the core of a small, senior team and own the full loop: shipping agents, measuring their quality, and making improvement a repeatable science. The work matters because the agents run in production for real users who depend on them without checking.

What You'll Do

  • Build agents that complete delegated executive tasks end to end, without human follow-up.

  • Design and maintain evaluation frameworks and metrics so the team knows, with confidence, how often agents get things right.

  • Expand the set of tasks agents can handle in tools executives already use, with evidence of readiness before launch.

  • Turn production failures into proven, permanent fixes quickly.

  • Own detection at scale using observability and eval tooling to surface known failure patterns and the long tail the team has not seen yet.

  • Build self-recovery so agents stay on task when a tool or external service fails.

  • Make agent improvement measurable: change one thing, run realistic end-to-end simulations, keep what works.

What We're Looking For

  • 5 or more years in data science, machine learning, or analytics, with a focus on evaluation systems and quality measurement for production AI.

  • Demonstrated experience designing evaluation methodologies: success criteria, dataset construction, metric selection, and distinguishing real signal from noise.

  • Production-quality Python and SQL; comfortable in Django, React, and TypeScript when work crosses into those areas.

  • Strong statistical and experimental design skills including sampling, variance, bias detection, and significance testing for non-deterministic systems.

  • Experience with LLM-as-a-judge systems, model-based graders, and grader calibration.

  • Ability to debug analytically across prompts, traces, model outputs, code, databases, and APIs to find the real cause of a failure.

  • Experience building and running agentic systems with multi-step execution and tool use in production.

  • Ground-truth data development experience including labeling guidelines, annotation quality control, and dataset maintenance.

  • Strong CS or engineering fundamentals from a rigorous degree program.

  • Track record of owning something end to end, shipping it to real users, and showing what changed because of it.

Compensation & Benefits

Salary range: $200,000 to $300,000 USD annually. Visa sponsorship is not available for this role.

Location

On-site in Palo Alto, California, United States.

Stack

LLMsPythonData ScienceTypeScriptSQLReactAgentic AIMachine LearningDjango
Posted
Oct 1, 2026
Last seen
Oct 1, 2026
First seen
Oct 1, 2026

Similar roles

Browse more AI jobs