Quadruped locomotion from scratch with PPO
Isaac Labrsl_rlPyTorchPPOPython
Description
I am training a Unitree Go2 quadruped to walk with PPO, using a custom manager-based RL environment in Isaac Lab.
Reward Tuning
Most of the work to get stable walking was tuning the rewards, penalties, and regularization terms. I wrote a custom height reward to stop the robot from crawling, and used curriculum learning to turn on the foot-lifting rewards only after the robot learned to balance. I also added squared penalties on terms like action_rate, which penalize large changes between actions and make the joint movement smoother.
Early policies learned to crawl. Crawling still follows the commanded velocity and gets a high reward, but it would not work on real terrain. I added a custom height reward that penalizes the body being too low, so crawling no longer scores well.
This caused a second problem. With the height penalty, the robot had to learn balance and foot clearance at the same time, and it failed to learn either. I used curriculum learning to turn on the foot-lifting rewards only after the policy could stay upright, so it learned one skill at a time.
Next Steps
Next steps are getup recovery policies for when the robot falls, and the domain randomization work needed for sim-to-real transfer.