At Miso Labs, our goal is to pass the voice Turing test. In order to do so, we must move beyond the current STT->LLM->TTS paradigm, and train full-duplex speech-to-speech models. This requires exploring new architectural ideas, and scaling them rapidly.
We’ve raised a large seed round from top investors and are hiring our founding research team. We’re a small team, so you can expect to work on every part of the model training stack from small-scale architecture experiments, to large pre-training runs on up to 1000 H100s. We are an in-person culture in San Francisco, and you will be expected to relocate to the SF area (we will cover the moving costs).
It is important to us that team members get credit for their work. Outside of our core IP, researchers are encouraged to publish their work at Miso Labs in conferences and on open source, and we are happy to sponsor travel to top conferences in AI/ML.
Preferred Qualifications:
If you’re interested in this role, please apply even if you don’t fit all the qualifications. We’re more interested in evidence of exceptional ability than prior experience in deep learning/AI.
This role comes with full benefits including:
Emotive foundation voice models
Salary
$400,000 - $500,000
Equity
2% - 3%
Location
San Francisco, CA
Experience
0+ years
Last stage
Seed
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.