All projects

WATonomous · Humanoid team

Fine-tuning π₀.5 for pick and place

A π₀.5 vision-language-action model, fine-tuned so the Pioneer humanoid's left arm picks up a cube and places it in a tray in Isaac Sim. I recorded the demonstrations with a leader arm, wrote the training and evaluation scripts, and trained only the action expert on top of a frozen vision-language backbone.

π₀.5 VLAFlow matchingLeRobotPyTorchIsaac SimLeader-arm teleopWeights & Biases

The policy after 6,000 weight updates, running in Isaac Sim, over several attempts in a row. The two smaller windows are what the policy sees: the left wrist camera and the base camera.

71teleoperated demonstrations, 5,654 frames at 25 fps
3camera views at 256 × 256: base and two wrists
6,000weight updates of 32 samples each
48 GBone cloud GPU, with the backbone frozen

Overview

The WATonomous humanoid team is building a humanoid called Pioneer, and I worked on its two-arm rig. My job was to get the left arm to do a simple manipulation task from camera images and a text prompt: pick up the cube and place it in the tray. To do that I recorded demonstrations in simulation, fine-tuned a vision-language-action model on them, and checked whether it had actually learned the task.

How π₀.5 is built

Training a policy from scratch takes far more data than one person can record, so I fine-tuned π₀.5, a vision-language-action (VLA) model from Physical Intelligence, starting from LeRobot's pi05_base weights. It has two parts:

  • PaliGemma, a vision-language model (VLM) built on the 2-billion-parameter Gemma language model. It reads the camera images, the text prompt and the arm's joint state.
  • The action expert, a much smaller action head with about 300 million parameters. It turns what the VLM understood into motion by flow matching: starting from random noise, it refines a chunk of 50 actions (two seconds at 25 Hz) in 10 steps.
3 camera images base and both wrists Task prompt "pick up the cube …" Joint state 8 values, as text tokens PaliGemma VLM SigLIP encoder + Gemma 2B frozen Action expert Gemma-style transformer trained · about 300M params Noisy action chunk refined in 10 steps Action chunk 50 steps 2 s at 25 Hz
frozen or fixed input trained in this project

Freezing the backbone

  • The PaliGemma backbone is already trained on internet-scale images and text, so it can already read a scene and an instruction. What it has never seen is this arm, these cameras and this task.
  • It has far too many parameters to fine-tune on 71 short demonstrations: that would need much more memory and risk damaging what it already knows.
  • So I froze it and fine-tuned only the action head, the part that has to adjust to my task.

Recording demonstrations

There was no dataset for this arm, so I recorded one in simulation:

  • The Pioneer arm in Isaac Sim is a digital twin of the real one.
  • I drive it with a 3D-printed leader arm: seven Feetech STS3215 servos used only as encoders, one per joint and one for the gripper.
  • Each leader joint maps directly to the matching simulated joint, so moving the leader by hand moves the robot the same way, with no inverse kinematics.
A white 3D-printed leader arm with a servo at each joint, clamped to a desk, with the Pioneer arm simulation on the monitor behind it
The leader arm. Seven servos used as joint encoders. On the monitor behind it is the scene it drives: the simulated arm, the table, the cube and the tray.
Recording a demonstration. I move the leader arm by hand, and the simulated arm follows it joint for joint.
  • My recorder saves four things 25 times a second as a LeRobot v3.0 dataset: the arm's joint positions (the state), the leader's joint targets (the action), three camera images and the text prompt.
  • I built the scene (a table, a cube and a tray) and matched the three cameras to the real robot's specs: a base camera modelled on an Intel RealSense D455 and one on each wrist. The closer the simulated images are to the real ones, the less the policy has to relearn on the real arm.
  • The finished dataset is 71 episodes of about 3 seconds each, 5,654 frames in total.
Two views from the simulation: a wrist camera close to the red cube between the gripper fingers, and the base camera looking down at the table, the tray, the cube and both grippers
What the policy sees. Two of the three views: the left wrist camera, close to the cube, and the base camera looking down at the table.

From a recording to model input

LeRobot's pre-processor and post-processor convert between the dataset and the model:

  • Images are resized to 224 × 224, and the prompt and joint state are turned into tokens.
  • Joint states and actions are normalized to roughly −1 to 1, using each dimension's 1st and 99th percentile. Without that, a wrist that moves through almost 2 radians would dominate the loss over a gripper finger that moves 0.05 metres.
  • The post-processor undoes the normalization, so the output is joint targets in radians again.
  • I trained on delta actions: each step the model predicts is an offset from where the arm was when it planned, while the gripper stays absolute. A policy that has not learned much then holds still instead of drifting toward an average pose.

The training script

I wrote the training loop by hand instead of calling LeRobot's training command. From top to bottom:

Setup
Fixes the random seeds, downloads the dataset from the Hugging Face Hub, and checks the Hub and Weights & Biases logins up front, so a missing token fails in seconds, not after hours of training.
Model
Loads pi05_base and matches its inputs and outputs to this dataset's cameras and joints. Weights stay in 32-bit floats while the forward pass runs in bfloat16, which is faster and smaller without rounding away small updates.
Train and validation split
Holds out one shuffled episode for validation and trains on the other 70. The split is by episode because neighbouring frames are nearly identical.
Data loading
Each sample is one frame plus the next 50 actions, shuffled so every update mixes frames from many episodes.
Processors
Builds the normalization and tokenization steps from the dataset's statistics, using the delta-action statistics for actions.
Freezing
Turns gradients on only for the action expert and its small input and output layers, and gives the optimizer only those parameters.
Learning rate
AdamW with a peak of 2.5e-5: linear warmup over the first 100 updates, a hold at the peak, then cosine decay to 10% of the peak over the last 1,000 updates.
Training loop
Runs 16 samples per forward pass and adds up the gradients from two passes, so every weight update is based on 32 samples. Gradients are clipped to a norm of 1.0.
Validation
Scores the held-out episode every 250 updates with the same random noise each time, so a change in validation loss comes from the weights.
Checkpoints and logs
Logs loss, learning rate and gradient norm to a CSV file and to Weights & Biases, and uploads the final 14.5 GB checkpoint to the Hub.

Making it fit

  • π₀.5 ran out of memory on my laptop's 12 GB GPU, so I moved training to a 48 GB cloud GPU.
  • Gradient accumulation keeps the update size independent of the hardware: a batch of 2 over 16 passes on a small GPU and a batch of 16 over 2 passes on the cloud GPU both give 32 samples per update.
  • Freezing the backbone saves memory too: it needs no gradients and no optimizer state.
  • The first updates of a fine-tune are the riskiest, because the gradients are largest then: the gradient norm starts near 3 and falls to about 0.4 within 60 updates. The warmup keeps those steps small so they cannot wreck the pretrained weights, and the cosine decay lets them settle at the end.
Three charts over the first 200 updates: the learning rate ramping up linearly to 2.5e-5 by update 100, the gradient norm falling from about 3 to about 0.4, and the training loss falling from about 0.8 to about 0.15
The first 200 updates. The learning rate ramps up while the gradient norm and the training loss fall quickly.

The first run was underfit

  • The first cloud run used 1,000 weight updates, about 5.7 passes over the roughly 5,600 training frames. The loss went down, but that says little about whether the arm would do the task.
  • So I wrote an offline evaluation. It gives the policy recorded camera frames and joint states and compares its predicted actions with what I actually did on the leader arm, next to a baseline that just holds the arm's current position.
  • The policy lost to that baseline. Over the first 0.4 seconds of each prediction, the part the arm executes before the policy plans again, its arm-joint error averaged about 10 degrees on episodes it had trained on, against about 5.5 degrees for holding still.
  • A policy that is worse than standing still on its own training data has not learned the demonstrations yet, so it needed more training.
Two charts of prediction error against how far ahead the policy predicts, for the arm joints and for the gripper fingers. Over the first 0.4 seconds, the policy's error on both training and held-out episodes is above the hold-still baseline.
Error against look-ahead, first run. In the shaded 0.4 s that gets executed, the policy (blue and orange) is less accurate than holding still (grey).

Training longer

  • I raised the run to 6,000 updates, about 34 passes over the data. The training loss fell from about 0.70 to about 0.013, and in simulation the arm now picks up the cube and places it in the tray, as in the video at the top.
  • Validation loss on the held-out episode was lowest between about 750 and 2,000 updates, around 0.07, then climbed to 0.144 while the training loss kept falling: the usual sign of overfitting.
  • One held-out episode is a noisy measure, but the trend is steady: the first run trained too little, this one probably too much, and the best checkpoint is likely somewhere in between.
Two charts over 6,000 updates. Validation loss dips to about 0.07 near 750 updates, stays low until about 2,000, then rises to 0.144. Training loss falls steadily from about 0.8 to near 0.01.
The 6,000-update run. Validation loss (left) bottoms out early and then rises. Training loss (right) keeps falling.
What I would change next
  • Save checkpoints more often, and pick one by validation loss and success in simulation.
  • Record more demonstrations, with more variety in where the cube and tray start.
  • Measure a success rate over many resets.
  • Try the policy on the real arm.

What I learned

  • A loss curve is not an evaluation. The first run had a falling loss and a policy that was worse than doing nothing. One chart against a hold-still baseline made that obvious.
  • Freeze what already works. 71 demonstrations cannot improve a backbone trained on internet-scale data, only damage it. Training the action head alone was enough for this task.
  • Memory is a design constraint. Batch size, gradient accumulation and precision all had to be chosen around the GPU.