The role
We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform.
You’ll be building and operating the core systems that power agentic AI at scale. Your mission: keep our ultra-low-latency, stateful, serverless compute engine rock-solid as we serve billions of agent requests for the most sophisticated AI teams in the world.
This role is highly technical and execution-heavy. You’ll own our reliability posture end-to-end—observability, performance tuning, incident ops, infrastructure health, and the automation systems that keep everything running smoothly. We want you to design new reliability systems, push the boundaries of automation, and continuously evolve the platform to meet the demands of next-generation AI workloads. If you're a builder who thrives on owning critical infrastructure at scale, this role is for you.
What you'll do
Collaborating closely with the founders, the infra team, and the dev team—and leveraging AI wherever it creates leverage—you will architect and operate the systems that keep Blaxel fast, resilient, and secure.
Who you are
Required skills
Preferred
Bonus
Experience with any of the following is a plus (not required):
About Blaxel
Blaxel is AWS for AI agents. We’re a new kind of cloud computing infrastructure optimized for the unique demands of agentic AI, leveraging a purpose-built 25ms cold-start serverless compute engine.
Now processing billions of agent requests, we power the coding agents and background AI tasks infrastructure for top AI startups. Founders choose us when they hit the limits of general-purpose clouds. We solve the hard infrastructure problems—statefulness, ultra-low latency, and secure sandboxed code execution—so they can focus on building their core AI products.
We raised a $7.3M seed round led by First Round Capital.
Blaxel is a cloud infrastructure company purpose-built for AI agents, often described as 'AWS for AI Agents.' The platform provides a serverless, perpetual sandbox environment enabling developers to build, deploy, and scale autonomous AI agents without managing underlying infrastructure. Its core differentiator is a sub-25ms sandbox resume time powered by microVM technology, co-located agent hosting, a unified AI gateway, and GitHub-native developer workflows. Founded in 2024 by six co-founders who previously built and sold ForePaaS to OVHcloud, Blaxel is backed by First Round Capital and Y Combinator.
Salary
$175,000 - $250,000
Location
San Francisco
Experience
3+ years
Total raised
$7.3M
Last stage
Seed
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.