About the role
You'll diagnose and resolve performance problems across Decart's ML systems, spanning research, training, and production inference. The largest share of the work is writing and optimizing kernels for TPU and Trainium. You'll also advise researchers on the performance cost of proposed model changes. We're looking for engineers with a demonstrated record in large-scale systems engineering and low-level optimization.
Minimum requirements
What we're looking for
Projects you might work on
Decart builds real-time generative AI infrastructure and models that enable ultra-low-latency training and inference for large-scale multimodal systems. Leveraging our systems-level optimization stack, we develop foundational interactive video and simulation models capable of generating and transforming dynamic environments in real time, making generative AI experiences live, persistent, and accessible at scale.
Salary
$200,000 - $240,000
Location
San Francisco, United States
Experience
2+ years
Total raised
$453.0M
Last stage
Series C
Investors
Moshe Shalev
Co-Founder & CPO
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.