Gantry

Evaluation report, Stera-10M

Two arms, identical per-task fine-tunes, 127 unpaired rollouts per cell on the physical four-arm SO-ARM101 bench. The only difference between the arms is whether the Stera-10M pre-training stage happened.

Baseline
21.9%
How often the arm finished the task with no human video training.
CI 18.8 to 25.3 · n=635
With ego stage
36.9%
How often it finished after learning from human video first.
CI 33.2 to 40.7 · n=635
Pooled lift
+15.0 pp
How much better the trained version did, counting all five tasks together.
95% range +10.0 to +19.8
Dose response
r² 0.80
How closely each task's gain tracks how much matching video it had.
1.0 would be perfect
Bench rollouts
1,270
How many attempts on the physical arm the report is built from.
5 tasks × 2 arms × 127

Success rate by task

base, task fine-tune only ego, human video then task fine-tune

How much each task gained

The gain itself, with its own 95% range. A task counts when its whole range sits above zero. Pour-over coffee is the one that does not, and it is also the task with almost no matching footage in the corpus.

Where the corpus went

Stera holds 181 hours of labelled activity. Only a fraction survives to become motion a fixed base arm could execute, and knowing which fraction is what makes the corpus size an honest number rather than a headline.

Where attempts stopped

The operator reports how far each attempt got on the task's declared ladder, not whether it worked. Most of the ego arm's gain is attempts that used to die at approach now reaching grasp.

Per task detail

The table the charts are drawn from. Every rate carries its count and its interval, so a rate from a handful of trials cannot be mistaken for a rate from many.

TaskStera hbase egolift
Pooled37.7 21.9%36.9%+15.0 [+10.0 to +19.8]

Ego stage training

One pass of the ego stage at the budget this report is costed on. Loss is shown because it gets asked for, not because it predicts the bench. This project has already published the case where it anti-predicted capability.

corpus 143 sessions · 6,143 segments · 14.55 h · rev 548a1f267416 protocol unpaired, n=127 per cell · seedable=false · Wilson 95%