robot-atlas

Domain

Manipulation & Learned Policies

From behavior cloning to vision-language-action models: how modern robots learn to act.

12 of 12 modules published

  1. Covariate shift, compounding error, and why naive imitation breaks in closed loop; the DAgger fix.

  2. Predicting action sequences instead of single steps: the CVAE structure, the chunk-size tradeoff, and temporal ensembling.

  3. Visuomotor control as conditional denoising over action sequences, with receding-horizon execution.

  4. RT-1, RT-2, RT-X, Octo, and OpenVLA: web-scale pretraining meets robot control, and the cost of discrete action tokens.

  5. pi0 to pi0.7: flow-matching action experts, FAST tokenization, open-world generalization, and where open weights stop.

  6. Gemini Robotics, GR00T, Helix, Skild, and GO-2: the closed-model landscape and how to read vendor claims.

  7. Every major policy across eight architectural axes: action representation, horizon, frequency, backbone, conditioning, cross-embodiment, hierarchy, openness.

  8. SayCan, code-as-policies, and keypoint affordances; why separate planners gave way to internalized hierarchy.

  9. DPPO, ConRFT, Recap, pi_RL, residual RL, and HIL-SERL: closing the reliability gap with on-policy experience.

  10. Temporal ensembling, real-time chunking, and the latency budgets that decide whether the control loop closes.

  11. Padded action vectors, motion transfer, and shared relative end-effector frames; the live disagreement.

  12. Training the VLM backbone on discrete tokens while a flow-matching expert learns actions behind a stop-gradient.