About the role
Qualifications
- CUDA + GPU inference optimization
- vLLM, SGLang, or TensorRT-LLM experience
- KV caching, paged attention, batching, token streaming, etc.
- Distributed compute (with GPUs is a super plus)
- No degree required
Company
Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.
Role
Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.
Day to day responsibilities:
- Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
- Conducting model performance reviews
- Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
- Sometimes write kernels and, yes, occasional tasteful shitposting
About Luminal
Making AI run fast on any hardware.
Required skills
Torch/PyTorch
CUDA
Other roles at Luminal
Job details
Salary
$150,000 - $250,000
Equity
0.15% - 0.75%
Location
San Francisco, CA
Experience
0+ years
Funding
Total raised
$5.3M
Last stage
Seed
Investors
Founders
What happens next.
No applications, no recruiter spam. Just the intro.
01
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
02
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
03
A meeting lands on your calendar
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.
