Lemma reads all of your agent conversations, finds the failures nobody knew to look for, and opens the PR that fixes them.
We sell reliability tooling. There is no version of this company where our own pipeline is the flaky part.
Right now it works because it's small and we're watching it. This role is about making it work when it's large and nobody is watching it.
The interesting problems here aren't modeling problems. Agent runs go for hours, call thousands of tools, and produce deeply nested spans that look completely different at every customer. We read all of production, not a sample, and we surface the anomalous fraction of a percent without anyone telling us what anomalous means. Doing that accurately is hard. Doing it at a cost per event that doesn't eat the business is the actual job.
Then the part that decides whether we have a company: reproducing a failure we saw once, proving it's real and not variance, and being right enough that a team lets us open PRs against their repo. Being confidently wrong once costs more trust than being right fifty times earns.
We move fast on hiring. Target is an offer within two weeks of first contact.
Onsite in San Francisco. We sponsor visas.
Production Monitoring for AI agents
Salary
$150,000 - $200,000
Equity
0.5% - 1%
Location
San Francisco, CA
Experience
0+ years
Total raised
$575K
Last stage
Pre-seed
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.