About the role
We’re hiring a Founding ML Research Engineer to work work through the full stack across data, pre-training, post-training & evals for training Generalist Audio Models. You’ll work through the entire stack with small team, tons of compute, high autonomy, and see your research ideas making it to production within a week(s).
What you’ll do
- Research better multi-modal architectures & codecs that are efficient across both spoken speech & general audio.
- Post-train audio models to have LLM like instruction following & in-context learning but over both text and audio.
- Build large-scale speech model pre-training and post-training (SFT/RLHF-style, distillation, preference optimization, etc.).
- Build scalable data + compute pipelines: dataset curation, filtering, mixing, tokenization/feature pipelines, evaluation harnesses.
- Look at lots of data & hear lots of audio.
What we’re looking for
- Industry/Academia experience pre-training / post-training large neural networks; speech/audio is a plus but not required, language/vision experience is also relevant.
- Strong ML systems and engineering depth (distributed training, performance, reliability).
- Comfort operating in ambiguity: you can spec, build, debug, and ship.
- A hunger to always ask - what would the next frontier look like?
To apply
Share arxiv links to papers that: you co-authored, you enjoyed reading, you found surprising.
About Kalpa Labs
Scaling Generalist Speech models
Required skills
Job details
Salary
$180,000 - $250,000
Equity
0.5% - 1%
Location
San Francisco, CA
Experience
3+ years
Funding
Total raised
$500K
Last stage
Seed
Investors
What happens next.
No applications, no recruiter spam. Just the intro.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.
