Software Engineer (Product, Infrastructure and Platform Reliability)

Tokyo, Japan

About the role

Software Engineer (Product, Infrastructure and Platform Reliability) Blog --> Career Opportunities --> ← All open positions Product Software Engineer (Product, Infrastructure and Platform Reliability) Join

the team behind Sakana Chat 🐟, Sakana Marlin 🐬, and Sakana Fugu 🐡 Infrastructure and Platform Reliability Engineer to keep our products fast, dependable, and cost-efficient as that moment turns into sustained scale Tokyo Full-time Japanese (Conversational) & English (Business) At Sakana AI, we are turning world-class AI research into products used by real users and enterprise customers, including Sakana Chat, Marlin, Fugu, and Namazu. These products are becoming core ways for our research to reach production environments at scale. In this role, you will join the PD Team and own the reliability of our products and the production inference platform behind them. You will work across cloud infrastructure, GPU resources, LLM serving, SLA management, monitoring, incident response, cost optimization, and enterprise non-functional

requirements. Key

Responsibilities Own reliability and SLA management across Sakana AI products, including Sakana Chat, Marlin, Fugu, and Namazu. Own production inference infrastructure for our LLM products, improving reliability, latency, throughput, GPU utilization, and cost efficiency. Operate LLM serving systems and support safe deployment and rollback workflows. Build monitoring, alerting, incident response, postmortems, and prevention practices, and participate in the on-call rotation. Translate enterprise security, availability, compliance, and SLA needs into infrastructure design. Support capacity planning and cost management, and provide technical

requirements and capacity forecasts to the business team responsible for GPU and cloud procurement. Collaborate within a global team that works in both Japanese and English. Required

Qualifications Experience designing and operating production infrastructure in a cloud environment such as AWS or GCP. Experience operating containerized services, APIs, batch jobs, or model inference workloads in production. Experience managing and automating infrastructure with IaC or comparable tooling. Hands-on experience with monitoring, alerting, incident response, and postmortems. Ability to reason about GPU, cloud, and serving resource usage from cost, performance, and availability perspectives. Experience with LLM serving frameworks such as vLLM, TensorRT-LLM, or SGLang. English communication skills sufficient to drive technical projects with internal and external stakeholders, including discussions around

requirements, cost, capacity, and priorities. Japanese language ability (if you are a native speaker or have passed JLPT N1/N2, please mention this in your application). Preferred

Qualifications Experience designing or operating production inference infrastructure for LLMs or machine learning models. Experience with Kubernetes, GKE, Vertex AI, Terraform, or comparable infrastructure tooling. Experience building or operating GPU clusters, especially NVIDIA H100/B200-class environments. Expertise in inference optimization techniques such as quantization, speculative decoding, and PD disaggregation. Knowledge of SRE practices, SLO design, capacity planning, or enterprise infrastructure.

Who We Are Looking For You can translate research and product needs into reliable production infrastructure. You optimize for latency, reliability, cost, operability, and product value. You are comfortable reading backend or inference pipeline code when needed. You can work effectively with business stakeholders on cost, capacity, and procurement decisions. Other Important Information Please submit your CV and cover letter in English. As mentioned, when applying for this role, please select the Software Engineer (Product, Infrastructure and Platform Reliability) option in Google Forms, and make sure that in your cover letter, explicitly write that you are applying for the “ Software Engineer (Product, Infrastructure and Platform Reliability) ” role to be considered for this position. How to Apply 1 Find a role Review the open positions and identify the one that fits your background. 2 Email [email protected] Send a brief introduction and indicate which role you are applying for. 3 Submit the form Complete the Google Form with your CV and cover letter (English). ← Back to all open positions © Sakana AI 株式会社

About Sakana AI

Sakana AI develops research-oriented artificial intelligence systems that draw inspiration from natural processes like evolution and collective intelligence. The company designs methods such as evolutionary model merging, multi-agent orchestration, and autonomous research agents to create AI models capable of adaptation and collaboration across tasks. Its work includes systems that automate elements of scientific research and frameworks that combine multiple models to generate new, task-specialized AI learners. Sakana AI engages in partnerships with enterprise and research organizations to apply its technology in domain-specific workflows and innovation projects. The company contributes tools and models to the broader AI community and explores applications that extend beyond conventional AI pipelines.

Other roles at Sakana AI

Interested?

Let me introduce you to the founders.

Skip the process

Job details

Location

Tokyo, Japan

Company

NameSakana AI
IndustryArtificial Intelligence (AI)
Team Size2

Funding

Total raised

$379M

Last stage

Series B

Investors

Lux Capital
Khosla Ventures
New Enterprise Associates
NVIDIA
MMitsubishi UFJ Financial Group

Founders

DH

David Ha

Co-Founder & CEO

LinkedIn
Llion Jones

Llion Jones

Co-Founder & CTO

LinkedIn
RI

Ren Ito

Co-Founder & COO

LinkedIn

What happens next.

No applications, no recruiter spam. Just the intro.

01

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

02

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

03

A meeting lands on your calendar

If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.