Staff Product Manager - Data Products
Hybrid
Fresh
Staff / Principal
314 - Product
Compensation
$260,000-$300,000/yrDescription
The Role
We're looking for our founding AI Data Product Manager to own Snorkel's Agentic Data and RL Environments roadmap. In this role, you'll lead the product strategy for a variety of data types (e.g. Agentic Coding, Computer Use). You will shape the roadmap for the datasets Snorkel invests in by understanding the market, incorporating frontier lab needs and collaborating with researchers at Snorkel and our academic partners.
This role is highly cross-functional, sitting between Research, GTM and Operations. As a founding member for this role, you will be in charge of setting up the frameworks to build the roadmap, gather data from relevant sources, and share the roadmap with both internal and external stakeholders.
What You'll Do
- Own the "data as a product" roadmap for Snorkel's Agentic and RL Environment focus areas, working x-functionally with research, academic partners, and GTM to define the skills and capabilities for our datasets
- Shape new "data" product areas and work with academic partners and research leaders to build Snorkel's competitive edge in the market
- Collaborate cross-functionally to help shape the roadmap and data strategy and influence business strategy
Minimum Qualifications
- 4-6 years of experience shaping technical roadmaps and working with researchers as stakeholders
- Comfort with ambiguity and working with multiple technical and non-technical stakeholders
- Experience working in fast-paced environments, setting up 0→1 products
- AI and ML fluency, especially related to Frontier Agentic Workflows and RL Environments
- 8+ years in product management, including 3+ years at senior/staff level owning a roadmap end-to-end (or 6+ years with a PhD/research background in ML)
- Demonstrated ownership of a technical product where data itself was the deliverable — datasets, benchmarks, evals, annotation pipelines, or labeled corpora sold or shipped to external consumers
- Working fluency in modern LLM post-training: SFT, preference data (RLHF/RLAIF), RLVR, reward modeling, and how data composition affects model capability. Must be able to hold a substantive conversation with a research scientist without an interpreter
- Familiarity with agentic systems and the current agentic eval landscape (e.g. SWE-bench-style coding evals, terminal/computer-use benchmarks, tool-use and long-horizon task evaluation) and an informed view on where they fall short
- Track record building product frameworks from zero — prioritization models, roadmap artifacts, intake processes — in an environment with no existing playbook
- Experience operating across research, GTM, and operations simultaneously, with evidence of driving decisions through influence rather than authority
- Direct customer-facing experience with highly technical buyers; ability to run a discovery conversation with an ML researcher or post-training lead and convert it into a roadmap commitment
- Quantitative rigor: can size a market, model unit economics of a data program (cost per trajectory/task/environment), and defend prioritization with numbers
- Strong technical foundation — comfortable reading research papers, discussing training dynamics with scientists, and reasoning about data pipelines end-to-end.
- Ability to write clearly for two audiences at once — internal research/ops and external frontier lab stakeholders
- Excellent analytical instincts — able to define success metrics for products that live close to research, where outcomes are often indirect.
Preferred Qualifications
- Prior experience at an AI data/environments company or inside a frontier lab's data, post-training, or evals org
- Has built or specified RL environments — sandboxed/containerized task environments, verifiable reward design, task generation, environment scaling and reproducibility
- Direct experience with agentic coding or computer-use data specifically: trajectory collection, rubric design, verifier construction, failure-mode taxonomy
- Hands-on technical ability — can write Python, query data, run a model, and prototype an eval without engineering support
- Experience selling or delivering into frontier labs, with an existing network among post-training, evals, or data acquisition leads
- Experience structuring academic or research partnerships, including co-development of datasets or benchmarks
- Published research, open-source datasets/benchmarks, or public writing that establishes credibility with the research community
- Experience with pricing and packaging for bespoke or semi-standardized data contracts, and the tension between custom deals and repeatable product
- Competitive intelligence muscle — has run structured win/loss or market mapping in a fast-moving, opaque market
- Prior founding-PM or 0→1 experience at a company between Series B and IPO
- Domain depth in one or more target verticals for agentic data (software engineering, enterprise workflows/CRM-ERP automation, finance, healthcare)
- Experience with human-in-the-loop data pipelines, annotation quality systems, or synthetic data generation at scale.
- Track record of leading large, cross-team initiatives without formal authority.
Stack
LLMsPythonAgentic AIMachine LearningReinforcement LearningData Engineering
- Posted
- Sep 14, 2026
- Last seen
- Sep 15, 2026
- First seen
- Sep 15, 2026


