← All projects

RLWRLD · Apr – Aug 2026

Adapting Robot Foundation Models to a 16-DoF Hand

Bringing large pretrained policies — GR00T N1.6, π0.5, and a 3D geometry-aware policy — to DexJoCo, a public 11-task dexterous-manipulation benchmark.

  • Industry
  • Simulation (MuJoCo)
  • VLA fine-tuning
  • Evaluation

A geometry-aware policy on a dexterous hand

GAM, a recently published 3D geometry-aware policy, was designed for parallel-jaw grippers. I adapted it to a Franka arm with a 16-DoF Allegro hand — new data loader and policy server, delta actions, depth inputs, and a corrected training recipe. Along the way I found that its training step limit counts micro-batches rather than optimizer updates, so with our gradient-accumulation settings early runs had received only ~156 updates instead of the intended 10k. Water-plant success rose from 11% to about 55% (one training run; 6 evaluation seeds × 50 episodes, a plateau over 20k–70k updates).

Probing what the policy expects to see: its predicted depth for the next step (middle) versus the depth that follows (right). The prediction is smoother and misses thin parts like the peg.

Fine-tuning and benchmarking VLAs

I fine-tuned GR00T N1.6 on all 11 DexJoCo tasks (46.2% mean success, versus 40.3% reported for GR00T N1.5) and re-evaluated the released π0.5 checkpoints in the same harness, matching the reported 52.5% mean (individual tasks vary).

Per-task success on the 11 DexJoCo tasks for GR00T N1.6 (fine-tuned) and π0.5.
Hammer Nail from the same starting scene: fine-tuned GR00T N1.6 drives the nail while π0.5 hovers until the time limit. Over all 50 episodes π0.5 is still ahead on this task (44 vs 36).

Evaluation you can trust

Re-running the same seed flipped 22.5% of episode outcomes. I traced it to unseeded flow-matching noise in the policy and nondeterministic GPU rendering; after fixing both, reruns were identical, so differences between models could finally be measured against a known noise floor. I also built the training/evaluation platform behind these runs across three compute clusters.

Before the fix: same checkpoint, same seed, run twice. Rendering nondeterminism alone sends the passes down different paths.

Credits

  • Work done as a Research Engineer at RLWRLD.
  • DexJoCo benchmark: Wang et al., arXiv 2605.16257. GAM: Han et al., arXiv 2606.17046. GR00T N1.6: NVIDIA. π0.5: Physical Intelligence.