All projects

Quadruped locomotion from scratch with PPO

Isaac Labrsl_rlPyTorchPPOPython

Description

I am training a Unitree Go2 quadruped to walk with PPO, using a custom manager-based RL environment in Isaac Lab.

Reward Tuning

Most of the work to get stable walking was tuning the rewards, penalties, and regularization terms. I wrote a custom height reward to stop the robot from crawling, and used curriculum learning to turn on the foot-lifting rewards only after the robot learned to balance. I also added squared penalties on terms like action_rate, which penalize large changes between actions and make the joint movement smoother.

The part that took the longest

Early policies learned to crawl. Crawling still follows the commanded velocity and gets a high reward, but it would not work on real terrain. I added a custom height reward that penalizes the body being too low, so crawling no longer scores well.

This caused a second problem. With the height penalty, the robot had to learn balance and foot clearance at the same time, and it failed to learn either. I used curriculum learning to turn on the foot-lifting rewards only after the policy could stay upright, so it learned one skill at a time.

Next Steps

Next steps are getup recovery policies for when the robot falls, and the domain randomization work needed for sim-to-real transfer.