Teleoperation Rigs
ALOHA, GELLO, UMI, and VR teleop: cost, data quality, throughput, and the embodiment gap.
- Last reviewed
- Reading time
- 9 min
- Citations
- 11
Nearly every demonstration in a robot learning dataset was produced by a human moving hardware. The human drives the robot itself, moves a scaled twin of it, holds a gripper proxy in hand, or works inside a headset while the robot mirrors them. Four rig families cover most of that collection in 2026: ALOHA-class bimanual workstations, GELLO kinematically matched exoskeletons, UMI handheld grippers, and VR teleoperation. Each family is a different bet on four axes: what the rig costs, how faithfully its demonstrations map onto robot execution, how cheaply its demonstration volume scales, and how large the gap is between the recorded motion and the robot's. Cost figures do not all describe comparable purchases: GELLO's sub-$300 paper BOM buys the controller, not its robot arm Wu 2023.
alpibrusl 2026 A community issue researched in June 2026 estimates the combined “ALOHA / ALOHA 2” category at about $17k-32k, but does not specify the currency code, configurations, or itemized inclusions and exclusions. That estimate is not a current vendor quote or a verified ALOHA 2 system total.
The comparison matrix
4 of 4 rigs
| Sources | |||||
|---|---|---|---|---|---|
GELLOKinematically matched exoskeletonGELLO for Franka, UR5, and xArm | $300Parts BOM under $300; excludes the target robot arm | highA scaled kinematic twin of the target arm; joint readings map one-to-one onto robot commands | mediumUnder $300 and assembles from printed and catalog parts, but each target arm model needs its own build | lowThe replica shares the arm joint structure, so the operator feels the arm constraints directly | GELLO paperProject site |
UMIHandheld gripper, no robot at collectionUMI gripper with GoPro wrist camera | $371$73 printed gripper + $298 GoPro and accessories, per gripper | medium155-degree fisheye wrist camera with SLAM-recovered gripper pose; no force channel | high111/h on the cup-arrangement benchmark, over 3x faster than spacemouse teleoperation | mediumA handheld gripper stands in for the robot gripper; latency matching and relative-trajectory actions close the gap at deployment | UMI paperProject site |
ALOHA-class workstationBimanual leader-follower workstationALOHA 2 (Stanford), Trossen AI Stationary and Mobile AI | not disclosedCommunity issue, researched Jun 2026: ALOHA / ALOHA 2 ~ $17k-32k; currency code, configurations and inclusions/exclusions not itemized; not a vendor quote | highLeader and follower arms share kinematics; demonstrations land directly in the robot joint space at 500 Hz | lowOne operator per fixed workstation; every additional collector costs another full rig | lowThe operator drives a kinematically identical arm, so recorded motion is the robot motion | ACT paperTrossen AICommunity estimate (Jun 2026) |
VR teleoperationHeadset plus retargeted human motionDROID Quest 2 rig, Open-TeleVision, Bunny-VisionPro | not disclosed | mediumImmersive stereo view and mirrored motion; actions are retargeted from the operator, not recorded from robot joints | highDROID: 50 operators across 13 institutions collected 76,000 trajectories (350 hours) in 12 months | highHuman hand and controller poses must be remapped through IK onto robot joints; calibration and avoidance layers exist because the mapping drifts | DROIDOpen-TeleVisionBunny-VisionPro |
Three reading rules for the matrix. First, ratings run low, medium, high, and the note in each cell states what the rating means for that rig and dimension, because "high throughput" is good news and "high embodiment gap" is not. Second, the highlight buttons swap the legend and print the per-rig detail for one dimension at a time. Third, every column sorts, cells with no published value are marked "not disclosed", and those cells sort last in both directions. Two cost cells remain “not disclosed”: the VR family has no published system total, and the cited ALOHA/ALOHA 2 community estimate does not establish a configuration-specific USD total. The estimate remains visible as source context, not a sortable price.
Bimanual workstations: fidelity first
The ALOHA design seats a human at a leader arm that is a twin of the follower arm the policy runs on. ACT, the policy built for that rig, learned six difficult bimanual tasks to 80-90% success, each from about 10 minutes of demonstrations, precisely because the demonstrations arrive already written in the robot's joint space Zhao 2023. Mobile ALOHA bolted a mobile base onto the same idea for whole-body tasks: with 50 demonstrations per task and co-training on existing static ALOHA data, success rates rose by up to 90% on tasks like sauteing shrimp and opening a two-door wall cabinet Fu 2024.
The hardware line behind these papers rebranded as Trossen AI in 2025-2026: $23,995.95 for the bimanual Stationary AI, $33,695.95 for Mobile AI on a base, $4,545.95 for the single-arm WidowX AI entry point, all running 500 Hz CAN FD control on the iNerve board with LeRobot and OpenPI integration Trossen Robotics 2026.
alpibrusl 2026 The community issue in alpibrusl/lex-robot, self-described as researched in June 2026, lists “ALOHA / ALOHA 2” as “research bimanual” at approximately $17k-32k. It warns generally that DIY BOMs, assembled products, and tariffs produce wide spreads, but does not tie the two endpoints to named configurations. The currency code, included hardware, and exclusions are not itemized; this is a secondary estimate, not current vendor pricing.
The trade is throughput. Each additional collector needs another full workstation and the space to put it, so fleet scale here is bought with capital, not with cheap duplication. No fleet-scale collection numbers are published for ALOHA-class rigs; the family earns its keep on data quality instead.
GELLO: the kinematic twin
GELLO takes the matched-kinematics idea and shrinks it. The device is a scaled replica of the target arm built from 3D-printed links and off-the-shelf motors, so the operator moves it like the arm itself and joint readings map one-to-one onto robot commands. Parts cost under $300, the paper calls the assembly process straightforward, says it needs minimal technical expertise, and quotes no build time; the published designs cover Franka, UR5, and xArm Wu 2023. In the paper's user study, 12 participants ran five tasks on a bimanual pair of UR5 arms using GELLO, VR controllers, and a 3D spacemouse; GELLO came out more reliable and faster than both Wu 2023.
Two limits define the family. The $300 buys a controller only; the robot arm it drives is a separate cost. And kinematic matching is per-model: every new arm needs its own GELLO design, which is the price of the one-to-one joint mapping.
UMI: collection without a robot
UMI removes the robot from collection entirely. The operator holds a 3D-printed parallel-jaw gripper fitted with a wrist-mounted GoPro and a 155-degree fisheye lens, performs the task wherever it naturally happens, and SLAM recovers the gripper pose from the video Chi 2024. The paper bills the printed gripper at $73 and the GoPro plus accessories at $298, for a $371 collection rig Chi 2024. The paper measures throughput in 15-minute collection windows and reports the gripper more than 3x faster than spacemouse teleoperation at 48% of bare-hand speed; the project site's rates are 111 demonstrations per hour on the cup arrangement task, against 35 per hour for teleoperation and 231 per hour for the unassisted human hand Chi 2024.
The price of robot-free collection is an embodiment gap the interface has to close at deployment time. UMI does it with inference-time latency matching and a camera-relative action representation, and the resulting policies deploy zero-shot onto any arm with a compatible gripper and camera setup; the paper demonstrates UR5 and Franka, and its own printed gripper has an 80 mm finger stroke Chi 2024. What the data lacks is a force channel, which matters for contact-rich tasks. Where robot-free collection fits in the wider data economics is the subject of The Data Bottleneck.
VR teleoperation: scale through headsets
VR teleoperation keeps the robot fixed and lets the operator move from inside a headset, with the robot mirroring the operator's hand or controller poses through a retargeting layer. DROID is the fleet-scale proof of the family: 50 operators across 13 institutions teleoperated Franka Panda rigs with Meta Quest 2 headsets and collected 76,000 trajectories totaling 350 hours over 12 months Khazatsky 2024. Diffusion policies trained on that pool beat policies trained on the next-best dataset by 22% in-distribution in controlled comparisons Khazatsky 2024, and the project maintains calibration protocols because multi-site retargeting drifts without them.
The newer systems push on immersion and fidelity. Open-TeleVision renders the robot's surroundings stereoscopically and mirrors the operator's arm and hand motion, validated on long-horizon tasks on two humanoid platforms Cheng 2024. Bunny-VisionPro runs on an Apple Vision Pro (launch price $3,499 Apple 2024), adds low-cost haptic feedback devices for the operator, and builds collision and singularity avoidance into the retargeting Ding 2024. The family's weakness is the embodiment gap itself: human hands share no kinematics with robot arms, so every action passes through IK and calibration, and that mapping is where demonstration quality is kept or lost.
Choosing by axis
No rig wins all four axes, which is why all four families are still in use. If the constraint is data fidelity on one platform, an ALOHA-class workstation or a GELLO built for the target arm records demonstrations with the smallest gap between collection and deployment. If the constraint is demonstration volume across many scenes, UMI collects without a robot in the room and VR scales through commodity headsets, as DROID showed across 13 institutions. If the constraint is capital, distinguish GELLO's under-$300 controller BOM and UMI's $371 collection rig from a complete workstation. The community ALOHA/ALOHA 2 estimate is not a verified minimum workstation budget. The embodiment gap is the bill each family pays differently: matched kinematics up top, interface design (UMI) and calibration protocols (DROID) in the middle, retargeted human motion at the VR end.
See also
- The Data Bottleneck
Robot-hours versus LLM tokens: the log-log reality of embodied data and teleop-farm economics.
- Hardware Taxonomy
Arms, humanoids, hands, sensors, and compute: a buyer's guide from SO-101 to Jetson Thor.
- Action Chunking (ACT and ALOHA)
Predicting action sequences instead of single steps: the CVAE structure, the chunk-size tradeoff, and temporal ensembling.
Linked from
- The Data Bottleneck
Robot-hours versus LLM tokens: the log-log reality of embodied data and teleop-farm economics.
- Hardware Taxonomy
Arms, humanoids, hands, sensors, and compute: a buyer's guide from SO-101 to Jetson Thor.
- Industrial Deployment
The installed base robot learning is trying to enter, and the jam-rate arithmetic that decides whether a 99 percent cell ships.
- Competing Theses
End-to-end scaling versus hierarchy versus world models versus RL fine-tuning, with falsification criteria.
- Safety and Assurance
Industrial robotics can certify a control system but not a learned policy, so what ships is a verifiable safety layer wrapped around an unverifiable one.
- Surgical Robotics
Intuitive, CMR, and Moon Surgical: the precision and reliability bar for certified robots.
- Space Robotics
NASA/JPL systems, orbital servicing, and ISRU: robotics where repair is impossible.
References
Tony Z. Zhao, Vikash Kumar, Sergey Levine, Chelsea Finn, RSS 2023.
https://arxiv.org/abs/2304.13705
Zipeng Fu, Tony Z. Zhao, Chelsea Finn, 2024.
https://arxiv.org/abs/2401.02117
alpibrusl, 2026.
https://github.com/alpibrusl/lex-robot/issues/3
Philipp Wu, Yide Shentu, Zhongke Yi, Xingyu Lin, Pieter Abbeel, 2023.
https://arxiv.org/abs/2309.13037
Cheng Chi, Zhenjia Xu, Chuer Pan, Eric Cousineau, Benjamin Burchfiel, Siyuan Feng, Russ Tedrake, Shuran Song, 2024.
https://arxiv.org/abs/2402.10329
Cheng Chi, Zhenjia Xu, Chuer Pan, Eric Cousineau, Benjamin Burchfiel, Siyuan Feng, Russ Tedrake, Shuran Song, 2024.
https://umi-gripper.github.io/
Alexander Khazatsky, Karl Pertsch, Suraj Nair, 2024.
https://arxiv.org/abs/2403.12945
Xuxin Cheng, Jialong Li, Shiqi Yang, Ge Yang, Xiaolong Wang, 2024.
https://arxiv.org/abs/2407.01512
Runyu Ding, Yuzhe Qin, Jiyue Zhu, Chengzhe Jia, Shiqi Yang, Ruihan Yang, Xiaojuan Qi, Xiaolong Wang, 2024.
https://arxiv.org/abs/2407.03162
Apple, 2024.
https://www.apple.com/newsroom/2024/01/apple-vision-pro-available-in-the-us-on-february-2/
Spot a factual error or missing qualification? Report a content correction.