← All projects

Georgia Tech · RL2 · Sep 2026 – present

What Makes Retargeted Demonstrations Learnable?

The same human demonstrations, retargeted two ways, train policies with very different success. Which properties of the retargeted actions explain it?

  • Ongoing
  • Real robot + simulation
  • Imitation learning
  • Humanoid
Frames from my egocentric Aria-glasses demonstration of a two-handed basket carry.

The question

Retargeting maps human motion onto a robot. In a follow-up to WARP, two retargeters applied to the same source demonstrations produced policies with very different success (14% vs 37% on a coffee-making task in simulation, each in its own scene layout). I study which properties of the resulting action data — consistency, smoothness, diversity, and how predictable actions are from observations — make demonstrations easier to learn by imitation.

What I did

  • Recorded egocentric human demonstrations with Aria glasses (45 of the 98 demos of a two-handed basket carry), which the WARP pipeline retargets to the Rainbow RBY1 humanoid, and ran the real-robot policy rollouts.
  • Characterized how a demonstrator's style survives retargeting: the two demonstrators' motions differ clearly in speed and timing.
  • Designed a controlled simulation study on the RBY1 (DexMimicGen coffee task): 7 training-data conditions, 14 policies, 1,400 closed-loop rollouts, testing whether data-quality metrics from the literature — action variance, state–action mutual information, spectral arc length, and signature-kernel diversity — explain the gap.
Per-demonstration duration and arm speed for two demonstrators.
Two people demonstrating the same task leave clearly different signatures in the retargeted data.

Findings so far

Closed-loop success as the share of WARP-retargeted demonstrations in training increases.
Same source demonstrations; only the retargeting mix changes. 100 simulated rollouts per point, 95% intervals.

In the WARP scene, success rises with the share of WARP-retargeted data — though the two retargetings also place the robot base about 17 cm apart, so scene layout and data are not yet separated. None of the dataset-level metrics tracks success consistently across tasks. Looking frame by frame is more telling: a policy errs more exactly where its nearest training demonstrations disagree about what to do next.

Policy error rises with how much the nearest training demonstrations disagree about the next action.
Same direction in all 6 datasets (3 tasks × 2 retargetings; Spearman ρ 0.25–0.50). Correlational, open-loop; policies are Zhenyang Chen's checkpoints, analysis mine.

Credits

  • Advisor: Prof. Danfei Xu.
  • Part of the WARP project in RL2, led by Chuye Zhang (WARP: arXiv 2606.29940).
  • DexMimicGen retargeted datasets, the baseline policy checkpoints analysed in the local-consistency figure, and the training/rollout code: Zhenyang Chen. I trained the 14 controlled-study policies with that setup and ran the analyses.