First full-time data engineer hire (CDI) at Ooak Data (YC S26). The core challenge is two-fold: (1) scaling the data anonymization pipeline (currently ~2 weeks per client, needs to get faster as POs come in for delivery), and (2) building the fine-tuning/adaptation platform — a local-model training infrastructure with automated data selection. They collect enterprise data across 50+ source types and need serious data engineering to make it usable for frontier AI lab training.
💰 Compensation: €70K–€100K base (CDI) + up to 0.5% equity 📍 Location: Paris, France — Hybrid 🕐 Experience: 5+ years in data engineering
Ooak Data (YC S26) builds the data infrastructure that frontier AI labs use to train and evaluate their agent models. They collect enterprise data from real company tools — Google Drive, Slack, Gmail, Notion, Jira, SharePoint, Teams, emails, videos, images — anonymize it, and turn it into RL environments and digital twins for agent training. Their core product, Alexandria, is designed to be the world's largest library of real-world business workflow datasets. They're also growing into "enrichment" — creating task ecosystems that let labs train models directly on real-world multi-step workflows, not just receive raw data. Signed large purchase orders with frontier AI labs (not SaaS), now in delivery mode. Data anonymization pipeline has been reduced from ~1 month to ~2 weeks per client.
Salary
$70,000 - $100,000
Equity
0.1% - 0.5%
Location
Paris, France
Experience
5+ years
Last stage
Seed
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.