The Bear Case
Why this could be another robotics winter, and the milestones that would prove it wrong.
- Last reviewed
- Reading time
- 9 min
- Citations
- 20
Every other module in this domain documents genuine progress, and the progress is real. This module gives the opposite reading the same evidentiary standard. The bear case says robotics is repeating the pattern that produced earlier AI winters: capital arrives on demonstrations, deployment evidence arrives late or never, and the funding window closes before the science closes the reliability gap. What follows is that case stated by its strongest named proponents. Every item is labeled evidence (a published position, a filing, a documented event) or speculation (a scenario about what happens next), and the module ends with the eight milestones whose outcomes would settle the argument either way.
Here is the scoreboard the argument turns on. Eight milestones, each chosen so its outcome moves the bear case whoever is right about everything else. Switch the status filter to Met and the board empties; switch it to Partial and four rows survive, each with the published evidence behind the call.
8 milestones: 4 not met, 4 partial, 0 met
| Milestone | Why it matters | Status |
|---|---|---|
| The generalization test: a single policy succeeding across many homes it has never seen, with no per-site data, is the clearest signal that robot intelligence is actually general. | partial | |
| Commercial viability at scale: five figures of humanoids doing documented productive work would end the pilot-program era. | not met | |
| The evaluation crisis: until labs publish comparable numbers on shared tasks, every demo is its own benchmark and progress claims cannot be arbitrated. | partial | |
| The reliability-gap test: reinforcement learning closes the gap to 99%+ one task at a time today; whether it does so across many tasks at once decides if the gap is engineering or science. | partial | |
| The direct test of Brooks's argument that vision-only training cannot produce dexterity: if touch is the missing channel, integrating it should move contact-rich performance. | not met | |
| If simulation fidelity reaches contact-rich manipulation, the data bottleneck breaks: training data becomes cheap, unlimited, and parallel. | not met | |
| The form-factor question in one number: if a humanoid's total cost per task matches a purpose-built system in the same application, the general-purpose thesis survives contact with accounting. | not met | |
| The scaling-law question: the bull case rests on robot performance scaling with data the way language did; a plateau would vindicate the structure-and-priors camp. | partial |
One policy, >90% in unseen homes
partialWhy it matters
The generalization test: a single policy succeeding across many homes it has never seen, with no per-site data, is the clearest signal that robot intelligence is actually general.
Current status
π0.5 cleaned kitchens and bedrooms in homes that were entirely absent from its training data, but the published evaluation covered a small number of homes; no lab has run a systematic multi-home evaluation with standardized tasks. Black 2025
How we’d know
A published, reproducible evaluation of one policy across more than ten unseen homes with standardized task definitions and success above 90%.
The skeptic record
Evidence: Rodney Brooks, co-founder of iRobot and Rethink Robotics and now running Robust.AI, is the field's most prominent named skeptic, and his claim is specific. His September 2025 essay calls the belief that current methods will produce practical humanoid dexterity within decades "pure fantasy thinking," on tactile grounds: vision-only training gives the robot no force channel, and Johansson's anesthetized-finger experiments show what humans become without one. Dexterity covers the argument in full Brooks 2025. His January 2026 predictions scorecard holds the line: deployable dexterity "will remain pathetic compared to human hands" beyond 2036, and walking humanoids stay too dangerous near people without new mechanical systems Brooks 2026. These are dated, falsifiable predictions from someone who has shipped robots for four decades, which is what separates them from ambient doubt.
Evidence: Morgan Stanley's July 2026 note argues the industry has a "PR problem": humanoids are publicly framed as direct substitutes for workers rather than as tools for hazardous or labor-constrained tasks, so "the industry's social license to deploy may matter just as much as technical performance" Wilkins 2026. The same note says investors "are increasingly looking for tangible evidence of real-world return on investment" rather than polished videos, and flags the US ban on Chinese humanoid imports as an R&D cost increase. The bank kept its 50,000-unit China shipment estimate for 2026. This is not a short call; it is a bank telling clients the bottleneck has moved from capability to proof.
Unit economics
Evidence: Unitree, the volume leader, shipped roughly 5,500 humanoids in 2025, more than any Western competitor Noreika 2026, and then saw Q1 2026 adjusted net profit fall 52.55% year over year Ramsey 2026. Volume and profit are different variables, and the bear case starts there: the company winning the shipment race is not yet winning the income statement.
Evidence: Figure's private valuation of $39B (September 2025) exceeds Goldman Sachs' projection for the entire humanoid market in 2035, $38B Noreika 2026. One private company is priced above the projected size of its whole market nine years out, so the valuation is a bet on the market outgrowing the projection, not on current revenue.
Evidence: the inputs are expensive too. Bessemer estimates more than $3B of aggregate robot data costs over the next two years and frames that spend as the moat Levine 2026. A moat made of capital is also an exposure: the spending requirement that protects incumbents is the same burn rate the funding cycle has to keep feeding. What a shipped cell actually has to earn back, against integrator margins and jam-clearing labour rather than against research budgets, is priced in Industrial Deployment.
Evidence: at Computex 2026, a Qualcomm-powered humanoid collapsed face-first during the live keynote, a failure attributed to a communication glitch, and was covered and carried off stage Malayil 2026. One fall says nothing about any single program. As a public datum about the distance between stage demos and floor reliability, it is on the record.
The demo-to-deployment gap
The Reliability Gap documents the arithmetic: 80% per step is a demo, 99.9% is a product, and compounding makes the distance between them brutal at thirty steps. The bear case is that closing the gap takes longer than the capital lasts. Lisa Yan, quoted in Bessemer's own analysis, describes the climb from 99% to 99.9% as steep Levine 2026. Morgan Stanley's ROI warning is the market making the same point from the other side Wilkins 2026. And the strongest verified deployment records, Agility's 65,000+ operating hours and Figure's 1,250+ hours at BMW Spartanburg, are supervised programs in narrow applications, which is what a bridge looks like before it becomes autonomy Noreika 2026. The evaluation crisis compounds the problem: when every demo is its own benchmark, even genuine progress is hard to underwrite.
The capital cycle
Evidence: robotics startups raised more than $23B globally by early June 2026 by PitchBook's tally, closing in on the $26B raised in all of 2025 Savage 2026. Crunchbase's narrower venture count puts the figure at $18.8B, already past the $14.1B peak set in 2021 Azevedo 2026. The two tallies disagree about the number and agree about the shape: capital is arriving faster than deployment evidence.
Whether that inflow constitutes a bubble is itself contested, and both sides are named. Bessemer Venture Partners argues there is "no robotics bubble" and that the sector is "structurally underinvested," while conceding that "not every company being funded will succeed" and that "some valuations are stretched" Levine 2026. The skeptic reading, Brooks's timeline predictions and Morgan Stanley's demand for ROI evidence, takes the same inflow as too much capital chasing demonstrations Brooks 2026 Wilkins 2026. This module does not resolve the disagreement; Competing Theses maps the technical version of it. The winter scenario needs only the disagreement itself: the market has not converged on how to value these companies, and unconverged valuations reprice violently.
How a winter would happen
The mechanisms below are scenarios, not predictions. Each rests on the evidence above, and each is labeled accordingly.
Speculation: the reliability gap outlasts investor patience. If 99.9% takes five more years of per-task engineering, the 2026 fund vintage marks to market against pilots rather than products, and the pullback slows every program, including the ones whose science is working.
Speculation: the data thesis plateaus. If 10x more data stops buying success-rate gains, the fix is new science (tactile integration, contact-faithful simulation, world models), and science does not arrive on a fund's schedule. Ken Goldberg's 100,000-year data gap is the strongest statement of this premise Goldberg 2025.
Speculation: the form factor is wrong. If task-specific systems keep winning cost per task, the humanoid thesis that most of the capital is priced on collapses and takes the funding cycle with it, even though warehouse automation keeps working. Brooks's "pure fantasy thinking" line is the named version of this scenario Brooks 2025.
The watchlist
Back to the board at the top of this module. Each row carries the question it settles, a current-status call with the published evidence behind it, and the observation that would flip it to met. As of August 2026 it reads four not met, four partial, zero met.
Several rows are closer than the labels suggest. π0.5 has already cleaned kitchens in homes it never trained on Black 2025, RL-100 and π*0.6 hold the per-task reliability records Lei 2025 Amin 2025, and RoboArena, RoboChallenge, and ManipulationNet are converging on shared evaluation infrastructure Atreya 2025 Yakefu 2025 Chen 2026. Others are early. Tactile foundation models such as Sparsh-X and TouchWorld exist but sit outside every major VLA pipeline Higuera 2025 Zhou 2026, contact-rich sim-to-real remains the unsolved case in reality-gap surveys Aljalbout 2025, and EgoScale's scaling law is measured on validation loss rather than real-world success rate Zheng 2026.
What would prove the bears wrong
The bear case is falsifiable, which is what makes it worth taking seriously. A verified 10,000-unit deployment breaks the commercial-viability premise. One policy above 90% across more than ten unseen homes breaks the generalization premise, and Generalization tracks how close the field is. A published cost-per-task parity analysis breaks the form-factor premise. A real-world success-rate scaling law breaks the plateau premise. Until those arrive, the honest summary is the board itself: zero of eight milestones met, four moving, and more than $23B raised by early June against that scoreboard.
See also
- Competing Theses
End-to-end scaling versus hierarchy versus world models versus RL fine-tuning, with falsification criteria.
- The Reliability Gap
80% is a demo, 99.9% is a product: what deployment numbers actually show.
- Generalization
What the pi0.5 and pi0.7 results demonstrate, and what they do not: the open-world gap.
- The Evaluation Crisis
Why N-of-10 trials and unreported variance mislead: 95% per-step success is unusable at 30 steps.
Linked from
- Competing Theses
End-to-end scaling versus hierarchy versus world models versus RL fine-tuning, with falsification criteria.
References
Rodney Brooks, 2025.
https://rodneybrooks.com/why-todays-humanoids-wont-learn-dexterity/
Rodney Brooks, 2026.
https://rodneybrooks.com/predictions-scorecard-2026-january-01/
Joseph Wilkins, CNBC, 2026.
https://www.cnbc.com/2026/07/29/morgan-stanley-humanoid-robots-pr-problem.html
Alius Noreika, 2026.
https://www.technology.org/2026/07/18/humanoid-robots-in-2026-what-is-actually-deployed/
Mireya Ramsey, TechTimes, 2026.
https://www.techtimes.com/articles/320197/20260711/robot-boom-meets-earnings-reality-unitree-profits-halved-optimus-not-sale.htm
Jeremy Levine, Talia Goldberg, Janelle Teng Wade, Alexandra Sukin, Bhavik Nagda, Jason Scheller, Christine Deakers, 2026.
https://www.bvp.com/atlas/bessemer-predicts-robotics-and-physical-ai
Jijo Malayil, Interesting Engineering, 2026.
https://interestingengineering.com/ai-robotics/qualcomm-robot-unexpected-collapse
Andre Savage, Market Briefs, 2026.
https://www.briefs.co/news/robotics-startups-raised-23-billion-in-2026-closing-in-on-all-of-2025/
Mary Ann Azevedo, Crunchbase News, 2026.
https://news.crunchbase.com/robotics/startup-venture-funding-surges-2026-data/
Ken Goldberg, Science Robotics, 2025.
https://doi.org/10.1126/scirobotics.aea7390
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, and 28 more, 2025.
https://arxiv.org/html/2504.16054v1
Kun Lei, Huanyu Li, Dongjie Yu, Zhenyu Wei, Lingxiao Guo, Zhennan Jiang, Ziyu Wang, Shiyu Liang, and 1 more, 2025.
https://arxiv.org/abs/2510.14830
Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Kevin Black, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, and 47 more, 2025.
https://www.pi.website/download/pistar06.pdf
Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, and 24 more, 2025.
https://arxiv.org/abs/2506.18123
Adina Yakefu, Bin Xie, Chongyang Xu, Enwen Zhang, Erjin Zhou, Fan Jia, Haitao Yang, Haoqiang Fan, and 29 more, 2025.
https://arxiv.org/abs/2510.17950
Yiting Chen, Kenneth Kimble, Edward H. Adelson, Tamim Asfour, Podshara Chanrungmaneekul, Sachin Chitta, Yash Chitambar, Ziyang Chen, and 15 more, 2026.
https://arxiv.org/abs/2603.04363
Carolina Higuera, Akash Sharma, Taosha Fan, Chaithanya Krishna Bodduluri, Byron Boots, Michael Kaess, Mike Lambeta, Tingfan Wu, and 3 more, 2025.
https://arxiv.org/abs/2506.14754
Jianyi Zhou, Feiyang Hong, Yunhao Li, Yicheng Zhao, Yongjue Cen, Zirui Liu, Jiakang Huang, Zirui Chen, and 4 more, 2026.
https://arxiv.org/abs/2607.07287
Elie Aljalbout, Jiaxu Xing, Angel Romero, Iretiayo Akinola, Caelan Reed Garrett, Eric Heiden, Abhishek Gupta, Tucker Hermans, and 4 more, Annual Review of Control, Robotics, and Autonomous Systems 2026 (accepted), 2025.
https://arxiv.org/abs/2510.20808
Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Castañeda, Fengyuan Hu, and 7 more, 2026.
https://arxiv.org/abs/2602.16710
Spot a factual error or missing qualification? Report a content correction.