Evaluating models for behavior and performance through simulated games
Evaluating models for social behavior, safety, and performance through multi-agent simulated games. We use popular games like Catan, Risk, or Poker, and make humans come play them against talking AI opponents, but on the backend we're actually running a multi-agent environment used to evaluate the models for skill and behavior over large sample sizes.
You can play in Social Arena and view our preliminary benchmarks today at https://olamlabs.ai/
Last stage
Seed
Investors
shreshth sharma
currently building simulated envs to create better evals @ olam labs, previously: was doing a cs and business double degree at uwaterloo (dropped out for yc) building reinforcement learning systems and doing data science @ RBC i was briefly top 50 on fifa mobile
No applications, no recruiter spam. Just the intro.
A few questions to make sure this role is the right shape for you. Two minutes.
I write the intro, send it to the founder, and handle the back-and-forth.
If they’re a yes, I book the chat. You show up — that’s the whole job-hunt.