Design, implement, test, and optimize distributed training infrastructure in Python and C++ for large-scale GPU clusters. Build and evolve telemetry systems to provide visibility into infrastructure & ML model performance, utilization, and cost related metrics Profile, benchmark, and debug performance bottlenecks across compute, memory, networking, and storage subsystems Drive architectural improvements across various ML services which deliver measurable efficiency improvements Build and evolve tools to automatically provide insights and recommendations to improve fleet-wide efficiency Optimize collective communication libraries (e.g., NCCL) for emerging NVLink and InfiniBand topologies Partner with ML researchers and infrastructure engineers to understand their plans and future needs and develop plans to balance growth with efficiency Collaborate with hardware teams to optimize for next-generation accelerators (NVIDIA, MAIA, and beyond) Embody our Culture and Values. Bachelor's Degree in Computer Science, or related technical discipline AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Bachelor's Degree in Computer Science or related technical field AND 10+ years technical engineering experience with coding in languages including, but not limited to, C++ or Python OR Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C++ or Python OR equivalent experience Deep understanding of the fundamentals of GPU architectures and DL/LLM architectures Deep experience in profiling and analyzing performance in large-scale distributed computing systems. Deep experience in profiling and analyzing performance in ML models especially GenAI models Experience with low-level GPU programming (CUDA, Triton, NCCL) and frameworks such as PyTorch or JAX. Experience in leading technical projects and supporting architectural decisions with data. Experience building infrastructure for large-scale machine learning or generative AI workloads. Experience in networking (InfiniBand, NVLink), storage systems, or distributed training parallelisms. Track record of contributing to high-performance computing or large-scale AI infrastructure projects.
Salary
$119,800 - $234,700
Location
United States
Total raised
$142.0M
Last stage
Series E
Investors
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.