HIGenNTO

Scalable Humanoid Interaction Generation via Noise‑Space Trajectory Optimization

 

HIGenNTO synthesizes humanoid-scene interaction motion references by optimizing the initial noise of a pretrained motion prior under sparse contact and scene constraints, with no interaction motion capture.

Video

What HIGenNTO does

01

Interaction synthesis in noise space

A frozen, scene-unaware motion prior is made contact- and geometry-aware at optimization time: losses on sparse spatiotemporal contact and scene constraints are backpropagated through the whole denoising chain into the initial noise, with no retraining and no interaction motion capture.

02

References physical policies execute

Privileged teachers track the generated motion sequences in simulation; depth-conditioned students keep 61-92% task success from onboard sensing alone and deploy on a Unitree G1.

03

Task programs by a coding agent

Tasks can be specified by a coding agent, which compiles task descriptions into prompt, constraint, and scene programs. It can also identify behaviors from a motion dataset and author new tasks, scene geometry included.

Method

HIGenNTO system overview
A task is a text prompt, sparse spatiotemporal constraints, and a scene signed-distance field. HIGenNTO optimizes the initial noise of a frozen text-conditioned motion prior against the losses on these constraints; the generated references train a tracking teacher that is distilled into a depth-conditioned student.

Generated motion sequences

    HIGenNTO dataset

    1,251 motion sequences across 134 scenes and 10 tasks at 50 Hz, with scene geometry, object trajectories, and contact labels. Browse it in 3D in the dataset viewer.

    Dataset viewer

    Executing the references

    For each task, a privileged teacher tracks the generated references; it is distilled into a depth-conditioned student that observes only onboard depth and proprioception. SONIC, a general-purpose scene-blind motion tracker run zero-shot from released weights, is shown for context. Each video is a 3×3 grid of nine scenes. In the teacher and SONIC videos, the blue humanoid (and box) is the generated reference the policy is tracking.

        Results

        Task success per task and arm
        Tracking errors per task and arm

        Task success and tracking errors across the eight evaluated tasks.

        Real-world deployment

        Real-robot deployment film strips

          Depth-conditioned students deployed on a Unitree G1 from onboard depth and proprioception.

          BibTeX

          Citation will be added upon publication.