Figure Trains Helix 2.5 to Clean 30 Unseen Homes
Figure has released its Helix 2.5 humanoid policy, which performed household chores across 30 unfamiliar homes to show how video pretraining dramatically improves real-world generalization.

Figure has announced Helix 2.5, a humanoid robot policy that successfully completed household tasks in 30 rented Bay Area homes that were entirely excluded from its training data. Operating under a zero-shot framework with no fine-tuning or local adaptation, the robot performed long-horizon tasks including tidying living rooms, folding towels, and making beds. The system treats these varied domestic environments as whole-body control problems, coordinating locomotion, active perception, and bimanual manipulation to recover from physical errors.
To isolate the impact of pretraining, Figure compared a model trained from scratch against one initialized with Index, its massive library of human-behavior video. In blind evaluations, the policy trained from scratch achieved a mere 9% success rate, whereas the Index-pretrained Helix 2.5 reached a 56% success rate. This represents a 47-percentage-point increase, or roughly a 6.2-fold improvement. Furthermore, Helix 2.5 matched the performance of the older Helix 02 model while requiring only half of the task-specific adaptation data, despite being tested across 30 times as many environments.
The company also demonstrated what it calls the first human-to-humanoid scaling law. By training four models across an eightfold range of Index data, Figure mapped a highly predictable decline in action-prediction loss. This relationship allowed engineers to forecast the final test loss of the largest training run to four decimal places, with a forecast error of just 0.54% of the measured loss variation. To support this trajectory, Figure has committed $3.5 billion in compute to the Helix project, with its Index dataset currently expanding by approximately 35 minutes of human video every second.
For robotics practitioners, these findings offer a concrete blueprint for deploying humanoid systems in unstructured environments. The results suggest that pretraining on diverse human video, followed by targeted adaptation on smaller robot datasets, is a highly viable path to achieving zero-shot generalization. While Helix 2.5 still failed 44% of its pooled trials and has not yet been released publicly, the predictable scaling of its action-prediction loss provides a reliable framework for teams to estimate performance gains before investing in expensive training runs.
This is our own summary of reporting by AlphaSignal



