On September 17, Figure AI put its newest humanoid policy, Helix 2.5, into 30 homes across the Bay Area that no Figure robot or dataset had ever touched, and asked it to do three ordinary chores: tidy a living room, fold towels into a basket, and make a bed -- pillows placed, comforter corners squared. No home got environment-specific fine-tuning, and no data was collected there beforehand. Figure calls the result the first demonstration, to its knowledge, of zero-shot whole-body generalization at this scope on a humanoid.
"To our knowledge, [this is] the first demonstration of zero-shot whole-body generalization at this scope on a humanoid."
The company's own comparison isolates what pretraining actually bought. Holding task-specific data and network architecture fixed, a policy pretrained on Figure's Index dataset succeeded at the entire task -- no partial credit given -- in 56% of attempts, against 9% for an otherwise identical policy trained from scratch. A human safety intervention during a run counted the attempt as a failure, not a pass. That's Figure's evidence its approach scales; it is also, so far, evidence only Figure has measured.
What Index pretraining changed
- Whole-task zero-shot success, 30 unseen homes
- Task-specific adaptation data needed to match Helix 02's success rate
- Environment-specific fine-tuning used
Three tasks and 30 homes is a real test, not an exhaustive one. Tidying toys, folding towels and making a bed all involve soft, forgiving objects and generous success windows -- one minute per toy-tidy attempt, three minutes for towels, a minute per side of the bed. Figure hasn't published a zero-shot result yet on harder categories: cooking, anything involving liquids, or tasks with a real safety cost to a mistake. The 9%-to-56% jump is evidence of a real capability gain on the tasks tested, not proof the same multiple holds everywhere a home might need a robot.
The dataset behind that gain is growing fast: Figure says Index now generates roughly 35 minutes of new human-behavior data every second, and the company has committed $3.5 billion in compute, arranged with cloud provider Nscale, to keep training Helix on it. No single one of the three evaluation tasks makes up more than 1.9% of that pretraining set -- Figure's own answer to the obvious follow-up question, which is whether the model simply memorized the test.
Helix's model lineage
- Feb 2025 — Figure introduces the original Helix VLA model, controlling a humanoid's upper body from one network.
- Jan 27, 2026 — Helix 02 extends control to the full body, replacing what Figure says was 109,504 lines of hand-engineered control code.
- May 2026 — Figure reports Helix 02 units completing full 8-hour autonomous shifts.
- Sept 17, 2026 — Helix 2.5 is evaluated zero-shot across 30 unseen homes.
The bet is well-capitalized. Figure raised more than $1 billion in a September 2025 Series C at a $39 billion post-money valuation, with Nvidia, Microsoft, Brookfield Asset Management and the OpenAI Startup Fund among the backers -- money that funds exactly the kind of large-scale data collection and compute commitment Index and the Nscale deal represent. A demo this well-funded still has to answer to the same standard as a scrappier one: does the number replicate outside the company that measured it.
Every number above is Figure's own measurement, on its own benchmark, reported in its own announcement. That doesn't make it false, but it puts it in the same category as most humanoid-robotics claims made this year: a company's word about its own demo, ahead of any outside replication.
- Index pretraining raised zero-shot success from 9% to 56% across 30 unseen homes
- Figure's software-first bet will out-compete Tesla's hardware-first bet in the same category
Figure's bet is that a smarter model, trained on more of the right data, gets a humanoid further than more hardware does. Tesla is running a different experiment at the same time: on September 17, the same day as Figure's announcement, Tesla placed an order for roughly 5,000 Optimus units with its supply chain and reiterated a target of about 50,000 units rolled out by the end of 2026, moving the program from lab demos toward mass-production audits at its factories. Tesla hasn't published a zero-shot generalization test comparable to Figure's; Figure hasn't disclosed a comparable production or unit-shipment number. Whether a model-first or a hardware-first approach reaches a paying household first is exactly the comparison neither company's own numbers can currently answer.
What Figure's test does establish, if it holds up under outside scrutiny, is a specific mechanism: a humanoid that gets better at homes it's never seen mainly by watching more human behavior, not by being shown that exact home first. If that scales the way Figure's pretraining curve already suggests, the bottleneck for a robot competent in your specific kitchen stops being your kitchen.
- Figure tested Helix 2.5 zero-shot in 30 Bay Area homes it had never seen before.
- Index pretraining raised whole-task success from 9% to 56%, Figure's own comparison found.
- Figure has committed $3.5 billion in compute with Nscale to keep training on Index data.
- Tesla, the same week, ordered about 5,000 Optimus units toward a 50,000-unit 2026 goal.
- Caveat: every Helix 2.5 number is Figure's own measurement -- no outside lab has replicated it.