Manipulation & Learned Policies
From behavior cloning to vision-language-action models: how modern robots learn to act.
This domain overview lists 15 articles, ordered as behavior cloning first, then the architectures that scaled it, with the cross-embodiment and real-time constraints those architectures ran into at the end.
Covariate shift and compounding error: why naive imitation breaks in closed loop, with DAgger as the standard fix.
Predicting action sequences instead of single steps: the CVAE structure, the chunk-size tradeoff, and temporal ensembling.
Visuomotor control as conditional denoising over action sequences, with receding-horizon execution.
RT-1, RT-2, RT-X, Octo, and OpenVLA: web-scale pretraining meets robot control, and the cost of discrete action tokens.
pi0 to pi0.7: flow-matching action experts, FAST tokenization, open-world generalization, and source-scoped checkpoint availability.
Gemini Robotics, GR00T, Helix, Skild, and GO-2: how to read closed-model vendor claims.
Every major policy across eight architectural axes: action representation, horizon, frequency, backbone, conditioning, cross-embodiment, hierarchy, openness.
SayCan, code-as-policies, and keypoint affordances; why separate planners gave way to internalized hierarchy.
DPPO, ConRFT, Recap, pi_RL, residual RL, and HIL-SERL: closing the reliability gap with on-policy experience.
Temporal ensembling and real-time chunking: the latency budgets that decide whether the control loop closes.
Padded action vectors, motion transfer, and shared relative end-effector frames; the live disagreement.
Training the VLM backbone on discrete tokens while a flow-matching expert learns actions behind a stop-gradient.
A dependency-aware route from supervised learning to real robot policies, with the minimum robotics stack each stage assumes.
Joint, Cartesian, torque, impedance, chunked and tokenized actions: what each representation gives the learner and pushes onto the controller.
What foundation means in robotics, how VLA, world-model and multimodal pretraining differ, and what adaptation still costs.