Member of Technical Staff, ML Inference Engineering

Palo Alto, CA

About the role

Member of Technical Staff, ML Inference Engineering Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more. Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language. Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house. Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication. If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you.

About

the Role Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable. We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do Performance Optimization Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference Analyze and improve latency, throughput, memory usage, and compute efficiency Profile system performance to detect and resolve GPU- and kernel-level bottlenecks Implement low-level optimizations using CUDA, Triton, and other performance tooling Improve support for mixed precision, quantization, and model graph optimization Build and maintain performance benchmarking and monitoring infrastructure Scale inference and training systems across multi-GPU, multi-node environments Inference Systems & Reliability Own and evolve our inference engine, enabling reliability and performance at scale Develop and optimize runtime inference services for large-scale AI applications Implement robust, fault-tolerant systems for data ingestion and processing

Requirements Must-have: 5+ years of experience writing high-quality, high-performance code Familiarity with NVIDIA GPU architecture and CUDA Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to A record of shipping research or systems that other people build on, whether in a lab or in industry Nice-to-have: Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes Experience developing large-scale, high-load production systems Experience maintaining or contributing to open-source ML projects Experience managing machine learning workloads on Kubernetes clusters Experience with InfiniBand or RoCE networking Experience with bare-metal provisioning and lifecycle management Experience operating large-scale AI training or inference clusters Experience with hardware health monitoring and predictive failure detection Experience with distributed storage systems Apply now Engineering Palo Alto, CA Share on: Apply now Terms of service Privacy Cookies Powered by Rippling

About Sanas

Sanas is a real-time speech-understanding platform that modulates accents while preserving voices and emotions for natural interactions. The platform assists multilingual speakers with accent correction, allowing them to communicate effectively without language barriers. Sanas's technology is utilized by contact centers and enterprises to enhance communication, improve customer satisfaction, and reduce miscommunication.

Other roles at Sanas

Interested?

Let me introduce you to the founders.

Skip the process

Job details

Location

Palo Alto, CA

Company

NameSanas
IndustryApplication Software
Team Size4

Funding

Total raised

$121M

Last stage

Series B

Investors

Quadrille Capital
Insight Partners
General Catalyst
GV (Google Ventures)
Human Capital

Founders

Shawn Zhang

Shawn Zhang

Co-Founder & CTO

LinkedIn
Maxim Serebryakov

Maxim Serebryakov

Co-Founder & CEO

LinkedIn
Sharath Keshava Narayana

Sharath Keshava Narayana

Co-Founder & President

LinkedIn

What happens next.

No applications, no recruiter spam. Just the intro.

01

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

02

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

03

A meeting lands on your calendar

If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.