Two arms, identical per-task fine-tunes, 127 unpaired rollouts per cell on the physical four-arm SO-ARM101 bench. The only difference between the arms is whether the Stera-10M pre-training stage happened.
The gain itself, with its own 95% range. A task counts when its whole range sits above zero. Pour-over coffee is the one that does not, and it is also the task with almost no matching footage in the corpus.
Stera holds 181 hours of labelled activity. Only a fraction survives to become motion a fixed base arm could execute, and knowing which fraction is what makes the corpus size an honest number rather than a headline.
The operator reports how far each attempt got on the task's declared ladder, not whether it worked. Most of the ego arm's gain is attempts that used to die at approach now reaching grasp.
The table the charts are drawn from. Every rate carries its count and its interval, so a rate from a handful of trials cannot be mistaken for a rate from many.
| Task | Stera h | base | ego | lift |
|---|---|---|---|---|
| Pooled | 37.7 | 21.9% | 36.9% | +15.0 [+10.0 to +19.8] |
One pass of the ego stage at the budget this report is costed on. Loss is shown because it gets asked for, not because it predicts the bench. This project has already published the case where it anti-predicted capability.