Domain
RL, Sim-to-Real & Locomotion
Reinforcement learning at scale, massively parallel simulation, and the transfer problem.
6 of 6 modules published
The MDP simulability gap: contact-rich manipulation resists the simulation that made walking routine.
Isaac Lab, Newton, MJX, and Brax: GPU-parallel environments and the wall-clock economics of training.
Domain randomization, teacher-student distillation, system identification, and real-to-sim correction.
From ANYmal to Unitree and the MIT humanoid line: how learned gaits became the default.
Motion tracking from PHC to ASAP and GMT, and the three decompositions of 2026.
LLM-written rewards, curricula, and where classical trajectory optimization still wins.