Domain
Manipulation & Learned Policies
From behavior cloning to vision-language-action models: how modern robots learn to act.
12 of 12 modules published
Covariate shift, compounding error, and why naive imitation breaks in closed loop; the DAgger fix.
Predicting action sequences instead of single steps: the CVAE structure, the chunk-size tradeoff, and temporal ensembling.
Visuomotor control as conditional denoising over action sequences, with receding-horizon execution.
RT-1, RT-2, RT-X, Octo, and OpenVLA: web-scale pretraining meets robot control, and the cost of discrete action tokens.
pi0 to pi0.7: flow-matching action experts, FAST tokenization, open-world generalization, and where open weights stop.
Gemini Robotics, GR00T, Helix, Skild, and GO-2: the closed-model landscape and how to read vendor claims.
Every major policy across eight architectural axes: action representation, horizon, frequency, backbone, conditioning, cross-embodiment, hierarchy, openness.
SayCan, code-as-policies, and keypoint affordances; why separate planners gave way to internalized hierarchy.
DPPO, ConRFT, Recap, pi_RL, residual RL, and HIL-SERL: closing the reliability gap with on-policy experience.
Temporal ensembling, real-time chunking, and the latency budgets that decide whether the control loop closes.
Padded action vectors, motion transfer, and shared relative end-effector frames; the live disagreement.
Training the VLM backbone on discrete tokens while a flow-matching expert learns actions behind a stop-gradient.