Principal, System Reliability Engineer

San Jose, CA

About the role

Position: Principal Engineer, System Reliability Location: San Jose, CA Job Id: 650 # of Openings: 0 Principle Systems Reliability Engineer Location: San Jose , CA (On-site) Ayar Labs is shattering AI data bottlenecks by moving data at the speed of light. As pioneers of co-packaged optics (CPO), we are using light instead of electricity to move data faster, further, and with a fraction of the energy needed to fuel the explosive growth of AI models. Backed by industry giants like NVIDIA, AMD, Mediatek and Intel and manufactured in partnership with the world’s leading semiconductor ecosystem, Ayar Labs’ co-packaged optics solution is key to unleashing next-generation AI scale-up architectures.

About

the Role Ayar Labs builds optical I/O technology for hyperscale AI infrastructure. This role owns reliability qualification at the system and fleet level: building the test infrastructure and evidence base that proves our products are ready for large-scale deployment, and then partnering directly with customers and integration partners to carry that evidence through joint qualification programs and pilot deployments. You will be responsible for both the engineering (test infrastructure, statistical reliability planning, standards compliance) and the relationship (translating technical evidence into the specific commitments our partners need to move forward).

What You'll Own Test infrastructure and evidence generation: Design, build, and operate custom test infrastructure that emulates real system-level electrical, thermal, and workload conditions ahead of, or independent of, any single customer's specific hardware. This includes sourcing or generating representative workload and power profiles and using them to drive continuous, long-duration test campaigns. Reliability statistics and demonstration planning: Define statistical sample sizes and test durations needed to demonstrate specific reliability and confidence targets, and apply the right methodology to each failure population in the system rather than a single blanket target. Standards and compliance: Ensure that test infrastructure, interfaces, and telemetry are built against relevant industry interconnect and management standards, so evidence generated internally holds up when reviewed by, or transferred to, an external partner's platform. Staged qualification execution: Execute staged qualification gates, from lab-level interoperability and functional testing through environmental and stress testing to live pilot deployment, applying a reliability / availability / serviceability lens throughout rather than treating any one dimension in isolation. Fault injection and lifecycle monitoring: Design and run fault-injection and failure-mode testing, including scenarios that exercise field-serviceable components under representative operating conditions, in partnership with customer and integration-partner operations teams. Economic modeling and fleet reporting: Provide reliability and performance data into cost-of-ownership and total-cost modeling shared with customers, and own ongoing fleet reliability reporting and incident response once products reach production. Cross-functional partnership: Serve as the primary technical point of contact with Tier-1 integration partners and hyperscale customer engineering teams across the qualification lifecycle, from early technical evidence through pilot sign-off and into steady-state operations. Basic

Qualifications 5+ years in fleet/systems reliability engineering for large-scale infrastructure, with a BS in Electrical Engineering, Computer Engineering, or a related field. You've built or operated test infrastructure that validates a component or subsystem's behavior before it is deployed into a customer's actual system, using representative rather than production hardware. You've defined and executed statistical reliability demonstration plans (sample sizes, test durations, confidence and reliability targets) for hardware components, and can explain the methodology behind them, not just apply a lookup table. You've worked with die-to-die, chip-to-chip, or board-level interconnect standards and telemetry or management interfaces relevant to high-speed data center hardware. You've partnered directly with external OEM or hyperscale customer engineering teams on a joint qualification or certification program, and are comfortable owning that relationship technically. Comfort operating a live, continuously running test or fleet environment: on-call coverage, incident response, and building the telemetry pipeline that turns raw sensor data into fleet-level reliability statistics. Strong written and verbal communication skills; you will regularly translate internal engineering data into evidence and documentation for external partners. Preferred

Qualifications MS in Electrical Engineering, Computer Engineering, or a related field. Experience with FPGA-based test or signal-generation systems. Experience characterizing or emulating real compute or network workload behavior for use in a test or validation environment. Background in data-center fleet operations, site-reliability engineering, or customer-facing qualification and certification processes. Familiarity with co-packaged optics, silicon photonics, or other emerging interconnect packaging technologies. Salary Range: $185,000 - $290,000 NOTE TO RECRUITERS: Principals only. We are not accepting resumes from recruiters for this position. Remuneration for recruiting activities is only applicable subject to a signed and executed agreement between the parties. Please don’t send candidates to Ayar Labs, and do not contact our managers. Ayar Labs is an Equal Opportunity Employer and is strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, sex, national origin, race, color, ethnicity, creed, religion, gender identity, sexual orientation, disability, veteran status, or any other characteristic protected by law. It is the policy of Ayar Labs to provide reasonable accommodation when requested by a qualified applicant or employee with a disability, unless such accommodation would cause an undue hardship. Veterans are more than welcome and encouraged to apply. ​ Apply for this Position Go back to the job list

About Ayar Labs

Ayar Labs uses silicon photonics and microring-based architectures to increase bandwidth, reduce latency, and improve power efficiency in accelerator and XPU environments. The company’s TeraPHY optical engine integrates UCIe-based optical chiplets to enable high-throughput, low-power data transfer across extended distances within and between compute systems. Its SuperNova external light source provides multi-wavelength laser output designed for scalable, standards-based deployment. Ayar Labs supports disaggregated memory and scale-up architectures, enabling AI hardware systems to operate with higher interconnect density and accelerator utilization.

Other roles at Ayar Labs

Interested?

Let me introduce you to the founders.

Skip the process

Job details

Salary

$185,000 - $290,000

Location

San Jose, CA

Company

NameAyar Labs
IndustryAI Infrastructure
Team Size3

Funding

Total raised

$870M

Last stage

Series E

Investors

NNeuberger Berman
AAdvent International
LLight Street Capital
PPlayground Global
Insight Partners

Founders

Mark Wade

Mark Wade

Co-Founder & CEO

LinkedIn
Chen Sun

Chen Sun

Co-Founder & VP of Silicon Engineering / Chief Scientist

LinkedIn
Vladimir Stojanovic

Vladimir Stojanovic

Co-Founder & CTO

LinkedIn

What happens next.

No applications, no recruiter spam. Just the intro.

01

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

02

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

03

A meeting lands on your calendar

If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.