PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments

Yuxuan Ma1,*, Zicheng Zeng1,2,3,*, Chunlin Peng1,4,5,*, Zhoujian Li1,6, Zetong Zhao1,7, Zhikai Zhang1,7, Yunrui Lian1,7, Han Xue1,7, Sikai Liang1,7, Weiyi Zhu1, Mulin Chen1,7, Chenghuai Lin1, Jiayu Zeng1, Yanwei An1, Songan Zhang5, Jiayuan Gu3, Jilong Wang1, Jingbo Wang1, He Wang1,8, and Li Yi1,2,7,†

1Galbot ·  2Shanghai Qi Zhi Institute ·  3ShanghaiTech University ·  4Zhongguancun Academy
5Shanghai Jiao Tong University ·  6National University of Singapore ·  7Tsinghua University ·  8Peking University
*Equal contribution ·  Corresponding author

Main video poster Main video · 3:00

Abstract

We present PASSAGE, a perception-conditioned planner–tracker framework for humanoid traversal. Using virtual reality and inertial motion capture, we collect 100 h of scene-aligned human motion across 1,500 cluttered scenes. A conditional flow-matching planner generates short-horizon references from motion history, a local destination, and a robot-centric multi-layer elevation map, while a perceptive whole-body tracker executes them at 50 Hz with geometric feedback. Real-time chunking promotes inter-chunk consistency, and planner-side RL post-training under the frozen tracker further improves closed-loop performance. Without skill annotations or obstacle-specific policies, one planner–tracker pair selects and composes traversal behaviors across unseen geometries. In simulation, component ablations quantify the contribution of each stage. Across three independent training seeds, scaling captured data from 6 to 100 h increases mean contact-free success from 48.1% to 68.9% on held-out scenes, while the final model with validated scene augmentation reaches 70.3%. The fully onboard system integrates egocentric 3D LiDAR perception, online occupancy mapping, 6.25-Hz planning, and 50-Hz control on a Jetson AGX Orin; tests across 50 unseen physical layouts demonstrate traversal without prebuilt maps or offboard computation.

Long Horizon Office Rollout

Long-horizon humanoid rollout through a cluttered office Office rollout · 0:39

Method

PASSAGE system overview: motion-in-scene dataset, multilayer elevation map, flow-matching motion generation, motion tracking and real-world deployment
Figure 1 — System overview.

Quantitative Results

Quantitative comparison across held-out scenes; the final PASSAGE model reaches 98.67 percent success and 70.27 percent contact-free success
Real-world evaluation results across physical obstacle layouts
Physical-layout evaluation.
Contact-free success increases from 48.1 percent with 6 hours of captured data to 68.9 percent with 100 hours
Performance with increasing data scale.

Data Collection

Scene-aligned data collection shown across egocentric, simulation and real-world views
Scene-aligned motion capture 0:21

Evaluation

ClutterRollout

Humanoid traversing a densely cluttered office course
Multi-obstacle course 0:24
Humanoid traversing a single cluttered office layout
Single clutter layout 0:08

Single Skill

Humanoid ducking under an obstacle
Duck under 0:09
Humanoid stepping over an obstacle
Step over 0:09
Humanoid passing an obstacle sideways
Sideways passage 0:11
Humanoid combining duck-under and step-over skills, sequence one
Duck + step I 0:05
Humanoid combining duck-under and step-over skills, sequence two
Duck + step II 0:08
Humanoid combining duck-under and step-over skills, sequence three
Duck + step III 0:08

Simulation

Indoor Scene

Humanoid traversing a simulated indoor living room
Living room 1:11
Humanoid ducking through a simulated indoor stair passage
Stair duck-under 0:19

LEGO Rollout

Nine simultaneous humanoid rollouts across LEGO-style obstacle fields
3 × 3 overview 0:23
Humanoid traversing an easy LEGO-style obstacle field
Easy layout 0:16
Humanoid traversing a medium LEGO-style obstacle field
Medium layout 0:17
Humanoid traversing a hard LEGO-style obstacle field
Hard layout 0:18

Citation

@misc{ma2026passagescalingscenealignedmotion,
  title         = {PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments},
  author        = {Yuxuan Ma and Zicheng Zeng and Chunlin Peng and Zhoujian Li and Zetong Zhao and Zhikai Zhang and Yunrui Lian and Han Xue and Sikai Liang and Weiyi Zhu and Mulin Chen and Chenghuai Lin and Jiayu Zeng and Yanwei An and Songan Zhang and Jiayuan Gu and Jilong Wang and Jingbo Wang and He Wang and Li Yi},
  year          = {2026},
  eprint        = {2609.18732},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.18732}
}