Evaluating robots in the real world.
Robocurve builds open-source tools and independent benchmarks to measure how well robots can do real-world jobs. Instead of relying on unverified demo videos from frontier labs, we score their models on reproducible benchmarks that anyone can trust.
Today there are no well-run, standardized robotics benchmarks. Labs evaluate in-house, and no independent group has stepped in to run continuous benchmarking as a service. The result is that no one actually knows how good anyone else is, or where the real frontier sits.
Good benchmarks require operating and maintaining physical hardware and real-world setups. Simulation only goes so far, since models that look strong in sim can show large performance gaps once deployed in the real world. Building real-world benchmarks means coordinating job-domain experts, evals engineering, and hands-on robotics all at once.
Frontier labs are targeting general-purpose robotics by 2028, yet the field of robotics evals barely exists. Whoever builds the trusted measure of robot capability becomes the reference everyone relies on. We combine backgrounds in AI evals and robotics to build and grow this field as fast as possible. We already shipped v1 of our open-source framework, Inspect Robots, and ran our first pilot scoring a frontier model on a real robot.
Last stage
Seed
Investors
Aris Zhu
Founder and CTO of Robocurve. Studied CS & Physics at Harvard. Aris builds robotics systems and AI agents across research and production. At Amazon AGI Labs, her test-time scaling research improved the NovaAct browser agent on WebVoyager. At Yondu Robotics (YC W24), she architected the navigation stack and fleet management system for humanoid robots. At Amazon Robotics she deployed a vision model to edge devices in production. She co-authored a paper in IEEE Robotics and Automation Letters.
LinkedInJay Chooi
MA Statistics and BA CS/Math from Harvard. Previously Research Fellow at MATS, researcher at the UK AI Security Institute, and top contributor to Inspect Evals, the UK government's AI eval framework. Rhodes Scholar and gold medalist at the International Olympiad on Astronomy and Astrophysics.
LinkedInNo applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.