Gantry trains a policy on your dataset and tests it on real hardware. You get a report on whether the data is worth training on, and what to fix if it is not.
See what it has graded
Submission to verdict, no configuration. Gates run in order of cost, and the cheap ones refuse a broken upload before anything expensive starts.
Counts, frame rates, licences, frozen streams, dropped frames. The bench inspects with the pipeline's own connectors, so what it sees is what training would see.
Identical frames, identical actions, identical training budget. The only change is that each episode's actions get handed to a different episode's video.
Your data has to beat the shuffle. Whatever fine-tuning-in-general buys, the control arm buys too, so only correspondence can explain a win.
Pass with the margin and its interval, refuse with the failing gate and a prescription, or abstain with the reason shown. Never a bare score.
Fifty intact datasets across three task families, every one passed. A screen that cries wolf on good data is worse than no screen, so we measured that first.
Delay every action by up to 20 frames and no dependence score flinches. So a separate gate hunts the offset directly and recovers it to the frame, in 12 seconds.
The probe lands on the same verdicts as a method that burns 17 GPU hours on VAEs and kNN estimators. Ours runs on a laptop while you wait.
Including the ones we damaged on purpose to see what escapes the grade.