The new standard of Physical AI eval
We measure whether your model is getting better: your policy on our robots. Video and a score for every run, back the same day.
- Independent
- Continuous
- Commercial tasks
- At scale
- Public and private
Positronic lets us evaluate checkpoints continuously as we train, so we can course-correct our research quickly. Day-to-day evaluation doesn’t compete with our robot fleet or operators for time, which keeps them focused on data collection and real-world deployment.
Every training run ends with the same question
The best model we have tested does 64 picks an hour. A person doing the same job by hand does over 1,300.
Is this checkpoint better than the last one? On real robots that question has no cheap answer: somebody resets the scene after every attempt, which buys tens of rollouts a day, not thousands.
Binary success rate throws away most of what the rollout showed. The difference you care about ends up smaller than the noise.
Nothing stays still. Lighting, placement, wear, the operator. Two checkpoints run a week apart were never compared under the same conditions.
How teams solve it with us
Checkpoint in, rollouts out. The lab, the operators and the resets are ours. An infrastructure problem becomes an API call.
Many embodiments and simulators, same tasks, scoring fixed before anyone sees a result. Both checkpoints meet the same conditions, and nobody picks the metric that wins.
More signal from each rollout. Time-to-milestone scoring, so a call that needed hundreds of runs takes tens: the ~30x trial reduction in the PhAIL paper.
Have a checkpoint you cannot score? Tell us what you are training, or take half an hour.
What you send, what comes back
A served endpoint, or the weights. Keep them on your own servers if you prefer. Plenty of teams do.
Your policy, blind, against a maintained baseline or your own previous checkpoint. Same tasks, same scenes, scoring fixed up front.
Video for every run, the time-to-milestone scores, and the comparison itself, back the same day.
Bring us the checkpoint you cannot score
Leave an address and a line about what you are training, or take half an hour now. Either way we work out what is worth measuring on your setup, and what a first round would tell you.