Cloud Inference Engineer

San Francisco, CA
Full-time
Visa Sponsorship

About the role

Qualifications

  • CUDA + GPU inference optimization
  • vLLM, SGLang, or TensorRT-LLM experience
  • KV caching, paged attention, batching, token streaming, etc.
  • Distributed compute (with GPUs is a super plus)
  • No degree required

Company

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.

Role

Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

Day to day responsibilities:

  • Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • Conducting model performance reviews
  • Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • Sometimes write kernels and, yes, occasional tasteful shitposting

About Luminal

Making AI run fast on any hardware.

Required skills

Torch/PyTorch
CUDA

Other roles at Luminal

Interested?

Let me introduce you to the founders.

Skip the process

Job details

Salary

$150,000 - $250,000

Equity

0.15% - 0.75%

Location

San Francisco, CA

Experience

0+ years

Company

NameLuminal
IndustryInfrastructure
Team Size7

Funding

Total raised

$5.3M

Last stage

Seed

Investors

Y Combinator
Felicis Ventures

Founders

J

Jake

Co-founder

J

Joe

Co-founder

M

Matthew

Co-founder

Joe Fioti

Joe Fioti

LinkedIn
Jake Stevens

Jake Stevens

LinkedIn
Matthew Gunton

Matthew Gunton

LinkedIn

What happens next.

No applications, no recruiter spam. Just the intro.

01

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

02

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

03

A meeting lands on your calendar

If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.