About the role
We would love to meet you if you:
- Philosophy: You are your own worst critic. You have a high bar for quality and don’t rest until the job is done right—no settling for 90%. We want someone who ships fast, with high agency, and who doesn't just voice problems but actively jumps in to fix them.
- Experience: You have deep expertise in Python and PyTorch, with a strong foundation in low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale. You're experienced with modern inference systems like TGI, vLLM, TensorRT-LLM, and Optimum, and comfortable creating custom tooling for testing and optimization.
- Approach: You combine technical expertise with practical problem-solving. You're methodical in debugging complex systems and can rapidly prototype and validate solutions.
The core work will include:
- Architecting and implementing robust, scalable inference systems for serving state-of-the-art AI models
- Optimizing model serving infrastructure for high throughput and low latency at scale
- Developing and integrating advanced inference optimization techniques
- Working closely with our research team to bring cutting-edge capabilities into production
- Building developer tools and infrastructure to support rapid experimentation and deployment.
Bonus points if you:
- Have experience with low-level systems programming (CUDA, Triton) and compiler optimization
- Are passionate about open-source contributions and staying current with ML infrastructure developments
- Bring practical experience with high-performance computing and distributed systems
- Have worked in early-stage environments where you helped shape technical direction
- Are energized by solving complex technical challenges in a collaborative environment
This is an in person role at our office in SF. We’re an early stage company which means that the role requires working hard and moving quickly. Please only apply if that excites you.
About Reducto
Reducto is a Series B team in San Francisco, founded in 2023 by Adit Abraham (CEO) and Raunak Chowdhuri (CTO), who met at MIT. Chowdhuri had published computer vision papers with over 100 citations before finishing high school. The company went through YC Winter 2024 and has raised $108M from a16z, Benchmark, First Round, BoxGroup, and YC.
The machine learning team is eight people and reports to the CTO. It runs a mix of in-house models and larger open-source ones, and an ML tech lead role is expected to be filled around the new year. Engineering is in the office five days a week near Union Square and Market, which is intentional at this stage.
Culture signals that hold across every role here: high agency and self-direction, no settling for 90%, and attention to detail that matters concretely. In document parsing a comma versus a period turns $20,000 into 20.000, so sloppy work is a real red flag. Engineers are expected to hold technical conversations with enterprise customers, not just stay in the code.
Other roles at Reducto
Machine Learning Engineer
San Francisco, United StatesFull-timeTechnical Recruiter
San Francisco, United StatesFull-timeLead Software Engineer, Platform
San Francisco, United StatesFull-timeML Infrastructure Engineer
San Francisco, United StatesFull-timeForward Deployed Engineer, Infra
San Francisco, United StatesFull-time
Job details
Salary
$200,000 - $300,000
Equity
0.1% - 1%
Location
San Francisco, CA
Experience
3+ years
Funding
Total raised
$108M
Last stage
Series B
Investors
What happens next.
No applications, no recruiter spam. Just the intro.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.
