RLWRLD · Apr – Aug 2026
Adapting Robot Foundation Models to a 16-DoF Hand
Bringing large pretrained policies — GR00T N1.6, π0.5, and a 3D geometry-aware policy — to DexJoCo, a public 11-task dexterous-manipulation benchmark.
- Industry
- Simulation (MuJoCo)
- VLA fine-tuning
- Evaluation
A geometry-aware policy on a dexterous hand
GAM, a recently published 3D geometry-aware policy, was designed for parallel-jaw grippers. I adapted it to a Franka arm with a 16-DoF Allegro hand — new data loader and policy server, delta actions, depth inputs, and a corrected training recipe. Along the way I found that its training step limit counts micro-batches rather than optimizer updates, so with our gradient-accumulation settings early runs had received only ~156 updates instead of the intended 10k. Water-plant success rose from 11% to about 55% (one training run; 6 evaluation seeds × 50 episodes, a plateau over 20k–70k updates).
Fine-tuning and benchmarking VLAs
I fine-tuned GR00T N1.6 on all 11 DexJoCo tasks (46.2% mean success, versus 40.3% reported for GR00T N1.5) and re-evaluated the released π0.5 checkpoints in the same harness, matching the reported 52.5% mean (individual tasks vary).
Evaluation you can trust
Re-running the same seed flipped 22.5% of episode outcomes. I traced it to unseeded flow-matching noise in the policy and nondeterministic GPU rendering; after fixing both, reruns were identical, so differences between models could finally be measured against a known noise floor. I also built the training/evaluation platform behind these runs across three compute clusters.
Credits
- Work done as a Research Engineer at RLWRLD.
- DexJoCo benchmark: Wang et al., arXiv 2605.16257. GAM: Han et al., arXiv 2606.17046. GR00T N1.6: NVIDIA. π0.5: Physical Intelligence.