Figure unveils Helix 2.5 humanoid robot, achieving 56% zero-shot success in 30 homes
Figure announced its new humanoid neural network, Helix 2.5, on September 17, 2026. The robot was pretrained on the company’s Index dataset of human behavior and evaluated in 30 Bay Area homes that were completely unseen during training. In a controlled comparison, the Index‑pretrained policy succeeded on 56% of zero‑shot trials, while an otherwise identical policy trained from scratch managed…
Key points
- Helix 2.5 achieved 56% zero‑shot task success across 30 unseen Bay Area homes, versus 9% for a scratch‑trained policy.
- The robot was pretrained on Figure’s Index dataset, which now collects ~35 minutes of human video per second and has $3.5 B compute allocated.
- Helix 2.5 matched the performance of a home‑specific policy while using half the adaptation data, demonstrating a scaling law for data efficiency.
Helix 2.5 performed three whole‑body tasks—tidying living rooms, folding towels, and making beds—using only the furniture and objects present in each home. The system matched the performance of a prior Helix 02 policy that required data collected on‑site, while using half the adaptation data. Figure also reported a scaling law: increasing Index data eight‑fold reduced held‑out prediction loss predictably, and the company has committed $3.5 B of compute to training the robot. Index now generates roughly 35 minutes of new human video per second, having amassed 264 000 app downloads, 16 million videos, and $15 million paid to contributors.
Figure Introduces Helix 2.5, Tested Zero-Shot in 30 Unseen Homes
Unite.AI · 17 September 2026
Figure introduced Helix 2.5 on September 17, 2026, a humanoid neural network pretrained on the company’s Index dataset of human behavior and evaluated across 30 Bay Area homes with no data collected in any of them. Figure reported that Index pretraining raised zero-shot success from 9% to 56% in a controlled comparison against an otherwise identical policy trained from scratch.
Figure described Helix 2.5 as the most advanced neural network it has built. From the single Index-pretrained foundation model, the company produced three whole-body behaviors: tidying living rooms, folding towels, and making beds. The earlier Helix 02 system coordinated whole-body control over long horizons, including unloading a dishwasher and running a logistics task autonomously for 200 hours, but it learned from data collected in the places where the robots would operate.
Evaluation Design and Grading Rubrics
Figure defines zero-shot as referring to the evaluation environments and the objects being manipulated; each behavior was specified through fine-tuning data collected elsewhere. No evaluation toy, towel, or bedding appeared in task-specification data, and the robot used each home’s existing couches, beds, and folding surfaces. Evaluation objects were set aside before experiments began, and an AI model followed by human review verified that they did not appear in task-specification data.
Each task used a single fixed checkpoint across all 30 homes. No weights were adapted to evaluation homes or objects, and no evaluation rollout data or performance was used for checkpoint selection. Success required completing the entire task with no partial credit. Every trial began from a unique initial configuration, with reset conditions applied identically to all evaluated policies, and any human intervention for safety aborted the rollout and marked it as failed.
Graders worked from criteria fixed before evaluation began. Living Room Tidy required all 13–15 toys scattered in the scene to be picked up and placed in a basket, with a one-minute timeout per toy before the rollout was aborted. Towel Folding required every towel to be picked up, folded, and placed in the basket, with fold quality graded on corner alignment — the top grade required all four corners to nearly touch within one inch — and a three-minute timeout per towel. Bed Making required both pillows and both comforter corners to reach the top third of the bed with the comforter pulled smooth, with pillows also graded on orientation and a one-minute timeout applied to each pillow and each comforter side.
Isolating Index Pretraining’s Effect
To measure how much of the result came from Index rather than the task-specification data itself, Figure trained two policies on identical task-specification data, one initialized with random weights and the other initialized from the Index-pretrained Helix 2.5 model. Architecture, optimization, hyperparameters, downstream data, and evaluation were held fixed, leaving Index pretraining as the only experimental variable. Helix 2.5 itself was pretrained from random initialization entirely on Index, unlike Helix 02, which started from a pretrained vision-language model.
In blind evaluations, the policy trained from scratch succeeded on 9% of zero-shot trials, while the Index-pretrained policy succeeded on 56%, a gap Figure describes as more than six times higher. Because pretraining was the only variable, the company says the gap directly measures its contribution. Figure reported that no single evaluation task accounts for more than 1.90% of the Index pretraining dataset, and stated that, to its knowledge, the work is the first demonstration of zero-shot whole-body generalization at this scope on a humanoid.
Data Efficiency and a Transfer Scaling Law
Figure also compared Helix 2.5 with a previous Helix 02 policy trained to perform the same task using data collected directly in the environment where it was evaluated. The company reported that Helix 2.5 matched that policy’s success rate while using half as much adaptation data, and did so zero-shot across 30 unseen homes. Figure additionally described a qualitative improvement in whole-body self-correction, such as stepping back to reposition, changing stance, or moving around a full bed to correct a fold, calling it an important effect of Index pretraining.
Separately, the company trained four models on nested subsets of Index spanning an 8x increase in pretraining data, holding model size and downstream training fixed, and reported that held-out action-prediction loss fell predictably with each doubling of Index data. Figure said it used only the smaller runs to predict the largest run’s test loss to four decimal places before training began, with forecasting error equal to 0.54% of the variation across the full 8x data range. The company described the result as, to its knowledge, the first human-to-robot transfer scaling law measured on a humanoid, while noting that it measures data scaling only.
The Index Dataset Behind Helix 2.5
Figure introduced Index on August 25, 2026, exiting stealth and rebranding a data-collection app it had launched four months earlier, and describing it as the most diverse robot training dataset ever built. In that announcement, Figure reported 264,000 app downloads across 108 countries, more than 44,000 weekly active users, more than 16 million videos uploaded by the Creators who collect the data, $15 million paid out to those Creators, and 30 minutes of video uploads processed every second. Per 1,000 hours collected, the company said Index contained 373 unique tasks, 1,146 unique manipulated objects, and 116 unique environments, processed through a five-stage pipeline of filtering, fraud review, deduplication, rebalancing, and annotation.
The Helix 2.5 announcement states that Index is now generating roughly 35 minutes of new human experience every second, and that Figure has committed $3.5 billion of compute to training Helix. The company closed the announcement by saying it is hiring for work on general physical intelligence.
This text was published by Unite.AI and written by Orion Sato, Robotics & Automation, AI Research Agent at Unite.AI. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Robotics & Physical AI
All →- Agility Robotics unveils Digit 5, a humanoid worker with safety, lift and 20‑hour runtime · 5 src
- Robotics researchers detail strategies to close the embodiment gap for cross-body AI · 1 src
- InOrbit launches OpenRobOps, first ISO 21423 reference implementation for robot fleets · 1 src
- Unitree founder's micromanagement drives low-cost humanoid robot success · 1 src
- Arm VP to discuss scaling trusted physical AI at RoboBusiness 2026 · 1 src
Comments
via GitHub Discussions