What we're looking for:
We need someone with 5+ years of infrastructure engineering experience who has deep hands-on expertise operating production Kubernetes at scale and working with LLM inference serving systems. You should be comfortable debugging NVIDIA GPU systems end-to-end (drivers, CUDA, NCCL, network fabric) and have a track record of building highly available, multi-cloud infrastructure for latency-sensitive AI workloads. Bonus points if you've deployed across heterogeneous accelerator types (NVIDIA GPUs, AWS Trainium, Google TPUs) or contributed to open-source inference frameworks like vLLM or SGLang.
What you'll do:
Decart builds real-time generative AI infrastructure and models that enable ultra-low-latency training and inference for large-scale multimodal systems. Leveraging our systems-level optimization stack, we develop foundational interactive video and simulation models capable of generating and transforming dynamic environments in real time, making generative AI experiences live, persistent, and accessible at scale.
Salary
$200,000 - $260,000
Location
San Francisco, United States
Experience
5+ years
Total raised
$453.0M
Last stage
Series C
Investors
Moshe Shalev
Co-Founder & CPO
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.