About the role
Software Engineer (Product, Infrastructure and Platform Reliability) Blog --> Career Opportunities --> ← All open positions Product Software Engineer (Product, Infrastructure and Platform Reliability) Join
the team behind Sakana Chat 🐟, Sakana Marlin 🐬, and Sakana Fugu 🐡 Infrastructure and Platform Reliability Engineer to keep our products fast, dependable, and cost-efficient as that moment turns into sustained scale Tokyo Full-time Japanese (Conversational) & English (Business) At Sakana AI, we are turning world-class AI research into products used by real users and enterprise customers, including Sakana Chat, Marlin, Fugu, and Namazu. These products are becoming core ways for our research to reach production environments at scale. In this role, you will join the PD Team and own the reliability of our products and the production inference platform behind them. You will work across cloud infrastructure, GPU resources, LLM serving, SLA management, monitoring, incident response, cost optimization, and enterprise non-functional
requirements. Key
Responsibilities Own reliability and SLA management across Sakana AI products, including Sakana Chat, Marlin, Fugu, and Namazu. Own production inference infrastructure for our LLM products, improving reliability, latency, throughput, GPU utilization, and cost efficiency. Operate LLM serving systems and support safe deployment and rollback workflows. Build monitoring, alerting, incident response, postmortems, and prevention practices, and participate in the on-call rotation. Translate enterprise security, availability, compliance, and SLA needs into infrastructure design. Support capacity planning and cost management, and provide technical
requirements and capacity forecasts to the business team responsible for GPU and cloud procurement. Collaborate within a global team that works in both Japanese and English. Required
Qualifications Experience designing and operating production infrastructure in a cloud environment such as AWS or GCP. Experience operating containerized services, APIs, batch jobs, or model inference workloads in production. Experience managing and automating infrastructure with IaC or comparable tooling. Hands-on experience with monitoring, alerting, incident response, and postmortems. Ability to reason about GPU, cloud, and serving resource usage from cost, performance, and availability perspectives. Experience with LLM serving frameworks such as vLLM, TensorRT-LLM, or SGLang. English communication skills sufficient to drive technical projects with internal and external stakeholders, including discussions around
requirements, cost, capacity, and priorities. Japanese language ability (if you are a native speaker or have passed JLPT N1/N2, please mention this in your application). Preferred
Qualifications Experience designing or operating production inference infrastructure for LLMs or machine learning models. Experience with Kubernetes, GKE, Vertex AI, Terraform, or comparable infrastructure tooling. Experience building or operating GPU clusters, especially NVIDIA H100/B200-class environments. Expertise in inference optimization techniques such as quantization, speculative decoding, and PD disaggregation. Knowledge of SRE practices, SLO design, capacity planning, or enterprise infrastructure.
Who We Are Looking For You can translate research and product needs into reliable production infrastructure. You optimize for latency, reliability, cost, operability, and product value. You are comfortable reading backend or inference pipeline code when needed. You can work effectively with business stakeholders on cost, capacity, and procurement decisions. Other Important Information Please submit your CV and cover letter in English. As mentioned, when applying for this role, please select the Software Engineer (Product, Infrastructure and Platform Reliability) option in Google Forms, and make sure that in your cover letter, explicitly write that you are applying for the “ Software Engineer (Product, Infrastructure and Platform Reliability) ” role to be considered for this position. How to Apply 1 Find a role Review the open positions and identify the one that fits your background. 2 Email [email protected] Send a brief introduction and indicate which role you are applying for. 3 Submit the form Complete the Google Form with your CV and cover letter (English). ← Back to all open positions © Sakana AI 株式会社
About Sakana AI
Sakana AI develops research-oriented artificial intelligence systems that draw inspiration from natural processes like evolution and collective intelligence. The company designs methods such as evolutionary model merging, multi-agent orchestration, and autonomous research agents to create AI models capable of adaptation and collaboration across tasks. Its work includes systems that automate elements of scientific research and frameworks that combine multiple models to generate new, task-specialized AI learners. Sakana AI engages in partnerships with enterprise and research organizations to apply its technology in domain-specific workflows and innovation projects. The company contributes tools and models to the broader AI community and explores applications that extend beyond conventional AI pipelines.
Other roles at Sakana AI
Job details
Location
Tokyo, Japan
Funding
Total raised
$379M
Last stage
Series B
Investors
Founders
What happens next.
No applications, no recruiter spam. Just the intro.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.
