You'll work on the stack that gets a policy from "recorded a human doing it" to "running on a line." Depending on your strengths and what's on fire, that could mean:
- Data collection tooling. Teleop capture, episode management, labeling pipelines. The difference between 8 minutes of usable demonstrations and 8 minutes of garbage is almost entirely tooling.
- Training and evaluation infrastructure. Reproducible runs, checkpoint management, and — the part everyone skips — real evaluation harnesses. If we can't measure a policy's success rate with enough trials to trust the number, we don't know anything.
- Policy work. Imitation learning (ACT, diffusion policies), visual servoing, observation-space design. Our experience matches the published literature here: what you condition the policy on matters more than which backbone you pick.
- Perception glue. Segmentation and tracking models feeding structured observations to controllers. Lots of SAM2, DINOv2, and small task-specific networks.
- Deployment. Getting the above onto real hardware, at real control frequencies, without it falling over on hour four.
You will also do unglamorous things: fix the data loader, chase a 20 Hz timing mismatch, re-cut a fixture, sit next to the cell and watch it fail 40 times. This is the job.
What we're looking for
Required
- Strong Python. You've written code other people had to maintain.
- Comfort with PyTorch — you can read a training loop and know where to put a breakpoint.
- Enough Linux and git to be self-sufficient.
- Empirical instincts. When something doesn't work, your first move is to isolate a variable, not to change three things and re-run.
- Currently enrolled in CS, EE, ME, robotics, or equivalent — or you can demonstrate the skills some other way. We care about the second clause.