Gantry

Evaluation report, SABER-10K

SABER-10K is DreamVu's public slice of a retail manipulation corpus, 10,000 episodes in LeRobot v2.0 layout. This report reads what the corpus already proved about itself, shows where that proof runs out, and states the corpus edit that targets the part it does not reach.

Episodes
10,000
The public release, split across three recording streams.
2,854,545 frames · 24.1 GB
Baseline
13.4%
How often the robot finished the task with no SABER training.
CI 11.4 to 15.7 · n=1,000
With SABER
29.3%
How often it finished after the corpus was added to post-training.
CI 26.6 to 32.2 · n=1,000
Carried by two
60%
Share of the whole gain that comes from the two fridge tasks alone.
96 of 159 points
Did not move
3 of 10
Tasks whose range still contains zero after the corpus was added.
two of them by 0.0 points

What SABER's own evaluation reports

Ten retail manipulation tasks, 100 rollouts each, GR00T N1.6 before and after SABER post-training. These are DreamVu's published numbers, not a Gantry measurement. The intervals are ours, added because a rate from 100 trials needs one before it can be compared with another. Evaluation was in simulation, which the paper states as its own limitation.

baseline, fine-tune only with SABER post-training

The gain is not evenly earned

Each task's gain with its own 95% range. Seven clear zero. Three do not, and two of those did not move at all. A mean of 29.3% is a real number and also a crowded one, because it averages a task that went to 100% with two that stayed exactly where they started.

Where the headroom is

Arithmetic on the table above, not a rollout. If the three tasks that did not move had instead moved like the median task that did, the corpus mean would read 33.0% rather than 29.3%. That 3.7 points is the part of the corpus the published result does not reach, and it is what a corpus edit is aimed at.

What the release is made of

Three streams share one corpus name and are not the same kind of data. Half the episodes carry a four-dimensional latent action, which is a learned code rather than a joint command. Two percent carry the whole-body stream, which the paper itself flags as undersized. A grade that reports one number for SABER-10K would be averaging across that.

StreamEpisodesAction dim What the action isVideo
stream1, egocentric5,0004 latent code learned from video, no physical units640×480 h264
stream3, dexterous hands4,80036 hand pose retargeted into robot joint space640×360 mp4v
stream2, humanoid20072 whole body retargeted, the stream the paper calls small640×360 mp4v
Total10,000mixed 2,854,545 frames at 29.97 fps24.1 GB

How an episode gets scored

A probe that never saw the episode predicts its actions from its own frames. The same probe is fitted again on deranged copies of the same episodes, where every clip keeps its frames and is handed somebody else's actions. The margin is how much better the real pairing predicts than the shuffled one. An episode near zero is one whose pictures do not tell you what its actions were, which is the definition of a clip that cannot teach.

Frozen for this study. Five folds, five derangements per fold, the gate's shipped grid of eight, images only, seeds from gates.signal.SEED. Nothing about the corpus is used to choose them, so the cut is a function two people would independently reproduce.

The scorer, running on a public corpus

Every episode of RoboMimic square multi-human, scored the way SABER-10K would be. Real output, 300 episodes, five folds. The line is where the bottom 25% falls. Fifteen episodes score at or below zero, meaning their own frames predicted a stranger's actions at least as well as their own.

The result we publish against ourselves

On that same corpus the margin is flat across the human operator tiers it was recorded with. Medians of 16.3, 15.8 and 15.7 for better, okay and worse, Kruskal Wallis p=0.955. The cut takes 19, 30 and 27 episodes from the three tiers, close to what a coin takes. The score is not measuring operator skill, and we say so before anyone asks.

The edit, and the four arms it is tested with

Score every episode, drop the lowest-margin quarter, retrain. The comparison that decides the question is CUT against RANDOM, never CUT against FULL. A cut corpus is smaller, and at a fixed gradient budget a smaller corpus is seen more often, so beating FULL can happen for reasons that have nothing to do with which episodes were dropped. In a null fixture that comparison came back p=0.041 from noise alone.

Full
the corpus as released
The reference point. Not the comparator, because it differs from every other arm in size as well as content.
Cut
worst quarter removed
The prescription. This is the dataset a customer would get back from the report.
Random
same size, different quarter
The comparator. Isolates which episodes were dropped from the fact that some were.
Anti
keeps the worst quarter
The sign check. If ANTI matches CUT, the score carries no quality information whatever CUT versus RANDOM says.

The gate before any GPU is bought

Plant a known defect in a quarter of the episodes using the product's own derangement, score the corpus again, and count how many planted episodes land in the worst quarter. Chance is 12.5% of the corpus. A score that cannot find a maximal, deliberately planted defect will not find a subtle natural one, and no arm gets trained. This runs on CPU and costs nothing.

What the report can conclude

Three verdicts, fixed before the first rollout. The middle one exists because an underpowered test is not evidence of absence, and it is reported with the smallest effect the design could have caught.

Improved

CUT beats RANDOM. Acting on the report produced a better policy than dropping the same amount at random.

Unresolved

Reported with its interval and its minimum detectable effect. Never written as "no effect".

Refuted

CUT does not beat RANDOM. The prescription is decoration, and that is publishable.

Sources

Corpus and stream statistics from the dataset card, DreamVu/SABER-10K on Hugging Face, CC BY-NC 4.0, LeRobot v2.0.
Per-task success rates, evaluation protocol and the simulation limitation from arXiv 2605.09613, SABER, a scalable action-based embodied dataset for real-world VLA adaptation. The evaluated corpus there is the full 44.8K-sample SABER, of which this release is a 10K subset.
Margin distribution and tier result computed by Gantry on RoboMimic square multi-human, 300 episodes, and committed at experiments/curation_efficacy/results/.

SABER-10K · 10,000 episodes · 2,854,545 frames · 3 streams · CC BY-NC 4.0 published eval n=100 per task · Wilson 95% · cut = lowest-margin 25%