Kairos
Back to jobs

Senior/Staff FDE - CUA

Hybrid
Snorkel AISan Francisco, CA, US / New York City, NY, US16 hours agoWebsite
Fresh
Staff / Principal
310 - DaaS FDE

Compensation

$180,000-$320,000/yr
Apply
Share

Description

About the Role

Snorkel AI is hiring a Forward Deployed Engineer focused on Computer Use Agents to partner with leading AI labs and enterprises on their most critical agentic-AI initiatives.

In this role, you will lead the technical execution of complex customer engagements involving agents that operate computers, browsers, and software environments to complete realistic, multi-step tasks. You will translate ambiguous product and model challenges into robust task environments, datasets, evaluators, and delivery plans that improve agent reliability and downstream performance.

You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.

Main Responsibilities

Computer Use Agents, Data, and Evaluation

  • Design and build task environments, datasets, and evaluation workflows for computer-using agents operating across browsers, desktop applications, terminals, and other software interfaces
  • Translate customer goals, agent failure modes, and real-world workflows into representative, multi-step tasks with clear success criteria
  • Develop data-generation, validation, and quality-assurance pipelines for multimodal and agentic training and evaluation data
  • Build automated evaluators, checks, and measurement frameworks to assess task completion, correctness, robustness, efficiency, and adherence to requirements
  • Diagnose agent failures across planning, tool use, perception, state management, and interaction with user interfaces; turn findings into improved tasks, data, and evaluations
  • Design and run experiments to measure how data, task design, and evaluation changes affect downstream agent performance
  • Deliver reusable, production-grade task suites, datasets, and evaluation assets that help customers train, benchmark, and improve computer-use agents

Forward Deployed Engineering & Customer Partnership

  • Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions
  • Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value
  • Rapidly prototype and productionize solutions across models, agent frameworks, APIs, browser or desktop environments, and custom applications
  • Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders
  • Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment

Technical Leadership & Scale

  • Identify recurring patterns across customer engagements and turn successful solutions into reusable task frameworks, evaluators, tooling, and best practices
  • Define and improve technical standards for agent task design, environment reliability, evaluation, and delivery
  • Partner with DaaS Engineering, Research, and Product teams to influence platform and product capabilities based on real-world customer needs
  • Lead technical design reviews, share expertise, and provide guidance to other engineers
  • Stay current with emerging agentic-AI, computer-use, evaluation, and data-curation techniques and assess their applicability to customer problems

What We're Looking For

  • 5+ years of experience in machine learning engineering, software engineering, applied AI, forward deployed engineering, solutions engineering, or a similar technical role
  • Strong Python skills and experience building reliable production software, data, or ML systems
  • Hands-on experience building, evaluating, or deploying LLM-based or agentic systems, including computer-use agents (CUA)
  • Strong understanding of experimentation and evaluation, including LLM-as-a-judge / model-based evaluation, defining metrics, and using empirical results to guide technical decisions
  • Experience designing task environments, datasets, and verifiers for agents, including reward & verifier design (RL with verifiable rewards)
  • Experience building or working with APIs, automation, web applications, browser-based systems, desktop applications, or developer tools
  • Experience manipulating, analyzing, and validating large or complex datasets using Python and the modern GenAI/LLM stack
  • Demonstrated ability to take ambiguous technical problems from problem definition through implementation and delivery
  • Strong technical communication skills and experience working directly with customers or cross-functional stakeholders
  • Demonstrated experience setting technical direction, creating reusable approaches across projects, and influencing broader engineering or product decisions

Preferred Qualifications

  • Experience developing agent benchmarks, task suites, or simulators
  • Experience with multimodal models, visual grounding, or evaluating agents that interact with graphical user interfaces
  • Experience with repo-scale agentic coding tasks, agent tool protocols & interop, or productized browser/computer-use agents
  • Experience building data pipelines for fine-tuning, reinforcement learning, preference optimization, benchmarking, or model evaluation
  • Experience working in fast-paced, customer-facing environments where requirements and technical approaches evolve quickly

Compensation

The base salary range for this position is $180,000–$320,000, with an additional variable compensation opportunity. The exact mix of base salary and variable compensation will depend on the role level and work location. Final compensation will be determined based on job-related skills, experience, relevant education or training, interview performance, and other business considerations.

All offers also include equity in the form of employee stock options, as well as benefits.

Stack

LLMsGenerative AIPythonAgentic AIMachine LearningFine-tuningReinforcement LearningData EngineeringMultimodal
Posted
Aug 27, 2026
Last seen
Aug 27, 2026
First seen
Aug 27, 2026

Similar roles

Browse more AI jobs