Ooak Data turns company data into training data for AI agents.
Frontier labs can train models to reason. They cannot train them to work: navigating a real company's Slack threads, half-finished Notion docs, contradictory Jira tickets, and permission boundaries. That requires real enterprise data, and you cannot synthesize it. You have to source it.
We plug into enterprise tools, anonymize everything into a structurally identical digital twin, and generate reinforcement-learning environments with expert-level tasks calibrated against frontier models.
We are a Y Combinator company with 7 figures signed contracts with three frontier AI labs, and we are scaling delivery aggressively over the next twelve months. Three founders, full-time since December 2025: Pierre-Louis (CEO, ex-COO in edtech, sold data to frontier labs), Grégoire (CPO, first PM & head of Ops at Epsor through Series B), Thomas (CTO, ex-Head of Data at PayLead, ex-Samsung AI lab).
Our engineering team ships faster than founders can specify. We have four product surfaces, most of them internal, all of them technical, and every one of them decides whether an environment ships this week or next.
We are hiring our first Product Manager to own the product: what gets built, in what order, and whether it actually made us faster.
The full chain, from the moment a CEO agrees to share their data to the moment a frontier lab receives an environment.
Partner experience
The data pipeline
Pharos, our delivery platform
RL environments
You will run two kinds of work at once: short projects that need tight coordination across Tech, Ops and GTM this week, and long ones where nothing exists yet and you get to decide what should.
The extra that make the difference
Ooak Data (YC S26) builds the data infrastructure that frontier AI labs use to train and evaluate their agent models. They collect enterprise data from real company tools — Google Drive, Slack, Gmail, Notion, Jira, SharePoint, Teams, emails, videos, images — anonymize it, and turn it into RL environments and digital twins for agent training. Their core product, Alexandria, is designed to be the world's largest library of real-world business workflow datasets. They're also growing into "enrichment" — creating task ecosystems that let labs train models directly on real-world multi-step workflows, not just receive raw data. Signed large purchase orders with frontier AI labs (not SaaS), now in delivery mode. Data anonymization pipeline has been reduced from ~1 month to ~2 weeks per client.
Salary
$70,000 - $90,000
Equity
0.2% - 0.5%
Location
Paris, Paris
Experience
6+ years
Last stage
Seed
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.