The Reliability Gap
80% is a demo, 99.9% is a product: what deployment numbers actually show.
- Last reviewed
- Reading time
- 9 min
- Citations
- 8
A robot that finishes a task 80% of the time is a demo. A robot that finishes it 99.9% of the time is a product. The distance between those numbers is the reliability gap: the difference between the success rates a system demonstrates in controlled settings and the rates that deployment requires. Bessemer's 2026 robotics analysis states the framing directly: getting from 80% task success to 99.9% is not a linear problem, and the last stretch requires fundamentally different approaches than the first Levine 2026. This module covers why the gap is structural, what is closing it, and what the deployment record shows once claims are separated from verified numbers.
The arithmetic of compounding
A manipulation task is a chain of decisions, and the chain succeeds only when every link does. If each step succeeds independently with probability , an -step task completes with probability . The exponent is the whole story. A policy with 95% per-step success completes a 30-step task 21% of the time. At 99% per-step the same task completes 74% of the time. Even at 99.9% per-step, a figure almost no system publishes, a 30-step task completes 97% of the time and still fails once in thirty runs. The per-step number and the task number are different metrics, and a result reported without its horizon cannot be compared against either. Drag the per-step success slider below from 95% up toward 100 and watch how late the episode number recovers.
(0.950)^30 = 21.5% episode success
On the reliability-gap calculator a 95.0 percent per-step policy yields 21.5% episode success at 30 steps and only 0.6% at the 100-step far end, with the 50 percent crossing near step 14.
Sampled episode success by episode length
| episode length | episode success |
|---|---|
| 0 steps | 100.0% |
| 10 steps | 59.9% |
| 25 steps | 27.7% |
| 50 steps | 7.7% |
| 75 steps | 2.1% |
| 100 steps | 0.6% |
The calculator computes the product at any setting, and its sliders span the full range so the boundary cases stay honest: at 100% per-step the task always completes, at 0% it never does, and at a horizon of one step the readout equals the per-step value. Two refinements matter for reading real systems. First, the independence assumption is generous: a mistake carries the robot into a state outside its training distribution, where the next mistake is more likely, so compounding error makes the true curve worse than the calculator's. Second, the inverse reading sets the engineering bar. Holding 99% end-to-end success on a 50-step task requires about 99.98% per-step reliability, a figure no generalist policy publishes.
Why the last stretch is different
Moving per-step reliability from 95% to 99% is an engineering problem with known methods: more data, better models, targeted evaluation. Moving from 99% to 99.9% is a different kind of problem, because the remaining failures concentrate in the long tail: unusual object configurations, lighting, and contact dynamics that any finite training set underrepresents Levine 2026. Lisa Yan, co-founder and CEO of Argus Systems and previously at Waymo, describes the deployment experience in Bessemer's analysis: closing the gap between 99% and 99.9% reliability is "a steep hill climb that takes longer than most people realize" Levine 2026. Bessemer's own forecast is more optimistic about timing, arguing the field's ChatGPT moment is not years away, and the disagreement is about the slope of the improvement curve rather than the existence of the gap Levine 2026.
The classical safety toolbox does not close the gap either. Formal methods such as barrier functions and reachability analysis do not scale to high-dimensional manipulation with learned policies, and as of mid-2026 no deployed vision-language-action model publishes formal safety guarantees. Assurance in practice rests on empirical evaluation and runtime monitoring, which makes the reliability gap and the evaluation crisis the same problem seen at different altitudes: one is about making failure rare, the other about being able to measure how rare it is (see The Evaluation Crisis).
What is closing the gap
Three methods are doing measurable work, each with a known ceiling.
RL from experience. Physical Intelligence's π*0.6 with Recap trains on the robot's own deployment experience, cutting failure rates by 2x or more on hard tasks and reaching above 90% success on espresso, laundry, and box-assembly after on-robot RL, with more than 2x throughput on some of the hardest tasks Amin 2025. RL-100 goes further on its task set: 100% success across 1,000 episodes on 8 tasks, including 250 consecutive successes on one task, and a juicing robot that served customers in a shopping mall for about seven hours without failure Lei 2025. The ceiling is that both are per-task specialist results earned through on-robot training, not generalist reliability. RL-100's own zero-shot figure under environmental and dynamics shifts is about 90%, which is the gap reasserting itself one distribution over Lei 2025.
Parada 2026 Runtime monitors. Google's ASIMOV-Agentic benchmark evaluates the orchestration layer around the policy rather than the policy itself: whether an embodied reasoning agent refuses unsafe tool calls from the VLA, predicts whether a task is feasible, and requests human intervention when uncertain Google DeepMind 2026.
Parada 2026 DeepMind describes Gemini Robotics ER 2 as its safest robotics model to date on safety-constraint-following and human-proximity benchmarks, with the ability to detect nearby humans and trigger safe stops. Monitoring catches failures instead of preventing them, so it raises effective reliability without changing the policy's own success rate.
Human in the loop. π0.7 accepts step-by-step language coaching that steers it through tasks it has not seen, recovering from failures that would otherwise end the episode Ai 2026. Supervised operation bridges the gap for paying customers while the autonomy catches up, but a bridged gap is not a closed one: the reliability number that matters is the one with the coach removed.
What is actually deployed
The gap is easiest to see in the deployment record, where claimed and verified numbers diverge. Circulating figures put Tesla's cumulative Optimus builds in the tens of thousands and Figure's deployments above ten thousand; neither figure comes from the companies they describe, and Tesla has never published a production count Noreika 2026. The verified record is narrower. Agility's Digit has accumulated more than 65,000 operating hours across nine customer facilities. Figure's eleven-month pilot at BMW Spartanburg logged more than 1,250 hours on a live assembly line. Unitree shipped roughly 5,500 humanoid units in 2025, with its G1 starting near $16,000, though volume has not produced profit: Q1 2026 adjusted net profit fell 52.55% year over year Noreika 2026.
6 of 6 rows
| Program | Value | Status | Source |
|---|---|---|---|
| Agility Digit Across nine customer facilities; named customers include GXO, Schaeffler, Toyota Motor Manufacturing Canada, and Mercado Libre. | 65,000+Operating hours | verified | Technology OrgJul 2026 |
| Figure 02 at BMW Spartanburg Eleven-month pilot on a live assembly line: 90,000+ parts loaded, above 99% placement accuracy per shift, 84-second cycle time. | 1,250+Operating hours | verified | Technology OrgNov 2025 |
| Unitree humanoid line Shipped across the G1/H1/H2 line in 2025, more than any Western competitor; 10,000 to 20,000 units targeted for 2026. | ~5,500Units shipped | verified | Technology OrgJul 2026 |
| Tesla Optimus Fremont production had not begun as of mid-July 2026, and Tesla has never published a production count; Q4 2025 earnings call described units as for learning, not productive tasks. | Not startedProduction status | verified | Technology OrgJul 2026 |
| Tesla Optimus (circulating figure) A widely circulated claim with no company source; Tesla has never published a production count, audited or otherwise. | 50,000+Units built | claimed | Technology OrgJul 2026 |
| Figure Helix 02 May 13, 2026 vendor livestream of package sorting; a real broadcast, but a single task, one site, and no independent audit of the success rate. | 8 hoursAutonomous shift | claimed | TechTimesMay 2026 |
The dashboard keeps the two classes apart. A verified row is documented against company statements, filings, or named customers; a claimed row is a circulating figure without a company source, or a vendor-run demonstration without an independent audit. Figure's eight-hour autonomous shift of May 2026 sits in the claimed rows deliberately: the livestream was real, but a single task at a single site with no published success rate is a demonstration, not a deployment metric Belmonte 2026. The valuation context belongs here as well. Figure's private valuation of $39B (September 2025) exceeds Goldman Sachs' projection for the entire humanoid market in 2035, $38B, a distance that only closes if reliability does Noreika 2026.
What solved would look like
A closed reliability gap has a measurable shape: a deployed system with a documented mean time between failures above 1,000 hours on real tasks, formal safety guarantees for human-robot interaction, and a published protocol for the failure modes that remain. No current system publishes any of the three. Until one does, the honest unit of progress is not the demo but the verified operating hour, and on that metric the field measures tens of thousands of hours against the millions that general-purpose deployment implies.
The second of those three is the one with an established answer everywhere except robot learning. Safety and assurance covers the standards stack industrial robotics already has, why a learned policy cannot be assigned a rating under any of it, and the wrapper architecture teams ship in place of the guarantee.
Self-check
Read the reasoning
- Yes, roughly: catching failures before they compound means most detected episodes still complete, so the effective rate climbs toward what the customer needsThis overstates what a monitor does. ASIMOV-Agentic and Gemini Robotics ER 2 raise effective reliability by refusing unsafe calls and triggering safe stops, which turns silent failures into visible ones; a task the system safely aborted still did not complete, and the monitored number is not the policy’s own success rate.
- No, only on-robot experience moves the policy’s own number: a monitor converts failures into safe stops and human escalationsThis is the module’s own division of labor. Monitoring catches failures instead of preventing them, and the methods that actually move per-step reliability are RL from experience, where π*0.6 cuts failure rates by 2x or more on hard tasks and RL-100 reaches 100% across 1,000 episodes, but as per-task specialists with known ceilings.
- No, so put a human coach on every episode insteadCoaching works, and π0.7 accepts step-by-step language correction to recover failures it has not seen. But a bridged gap is not a closed one: the reliability number that matters is the one with the coach removed, and supervised operation is a deployment tactic, not a path to the customer’s 99%.
Monitors make failures visible and coaches rescue episodes, but only on-robot experience moves the policy’s own per-step number: detection is not prevention.
See also
- The Evaluation Crisis
Why N-of-10 trials and unreported variance mislead: 95% per-step success is unusable at 30 steps.
- RL Fine-Tuning of Policies
DPPO, ConRFT, Recap, pi_RL, residual RL, and HIL-SERL: closing the reliability gap with on-policy experience.
- Humanoid Whole-Body Control
Motion tracking from PHC to ASAP and GMT, and the three decompositions of 2026.
- Dexterity
Contact-rich manipulation, the tactile sensing gap, in-hand reorientation, and deformables.
Linked from
- The Evaluation Crisis
Why N-of-10 trials and unreported variance mislead: 95% per-step success is unusable at 30 steps.
- Industrial Deployment
The installed base robot learning is trying to enter, and the jam-rate arithmetic that decides whether a 99 percent cell ships.
- Dexterity
Contact-rich manipulation, the tactile sensing gap, in-hand reorientation, and deformables.
- Generalization
What the pi0.5 and pi0.7 results demonstrate, and what they do not: the open-world gap.
- Competing Theses
End-to-end scaling versus hierarchy versus world models versus RL fine-tuning, with falsification criteria.
- The Bear Case
Why this could be another robotics winter, and the milestones that would prove it wrong.
- Safety and Assurance
Industrial robotics can certify a control system but not a learned policy, so what ships is a verifiable safety layer wrapped around an unverifiable one.
- Autonomous Vehicles
The AV stack as a robotics problem: perception, prediction, planning, and the long tail.
- Surgical Robotics
Intuitive, CMR, and Moon Surgical: the precision and reliability bar for certified robots.
- Space Robotics
NASA/JPL systems, orbital servicing, and ISRU: robotics where repair is impossible.
References
Jeremy Levine, Talia Goldberg, Janelle Teng Wade, Alexandra Sukin, Bhavik Nagda, Jason Scheller, Christine Deakers, 2026.
https://www.bvp.com/atlas/bessemer-predicts-robotics-and-physical-ai
Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Kevin Black, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, and 47 more, 2025.
https://www.pi.website/blog/pistar06
Kun Lei, Huanyu Li, Dongjie Yu, Zhenyu Wei, Lingxiao Guo, Zhennan Jiang, Ziyu Wang, Shiyu Liang, and 1 more, 2025.
https://arxiv.org/abs/2510.14830
Google DeepMind, 2026.
https://huggingface.co/datasets/google/asimov_agentic
Carolina Parada, 2026.
https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/
Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, and 79 more, 2026.
https://www.pi.website/blog/pi07
Alius Noreika, 2026.
https://www.technology.org/2026/07/18/humanoid-robots-in-2026-what-is-actually-deployed/
Kyle Belmonte, 2026.
https://www.techtimes.com/articles/316632/20260514/figure-ais-helix-02-robots
Spot a factual error or missing qualification? Report a content correction.