Team leadership: Lead a team of experienced SREs to ensure uptime, resiliency and fault tolerance of AI model training and inference systems. Observability: Design and help maintain monitoring, alerting, and logging systems to provide real-time visibility into model serving pipelines and infra. Automation & Tooling: Lead building of automation for deployments, incident response, scaling, and failover in hybrid cloud/on-prem CPU+GPU environments. Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements. Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments. Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows. Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with Site Reliability Engineering, DevOps, or Infrastructure Engineering Leadership roles AND 8+ years experience with Kubernetes, Docker, and container orchestration, AND 6+ years experience with programming/scripting skills not limited to Python, Go, or Bash Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience AND 10+ years experience with Kubernetes, Docker, and container orchestration, AND 10+ years' experience with public cloud platforms like Azure/AWS/GCP and infrastructure-as-code OR equivalent experience 6+ years people management experience. 8+ years experience in monitoring & observability tools (Grafana, Datadog, OpenTelemetry, etc.). Knowledge of CI/CD pipelines for Inference and ML model deployment. Solid knowledge of distributed systems, networking, and storage. Experience running large-scale GPU clusters for ML/AI workloads (preferred). Familiarity with ML training/inference pipelines. Experience with high-performance computing (HPC) and workload schedulers ( Kubernetes operators). Background in capacity planning & cost optimization for GPU-heavy environments
Salary
$165,600 - $296,400
Location
United States
Total raised
$142.0M
Last stage
Series E
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.