Adding new environments and objects is cheap in human effort: only capture, object selection, and
the text edit involve the operator, while reconstruction, asset generation, physics estimation, grasp
synthesis, and trajectory generation run automatically. Objects are populated consistently with each scene's
real-world scenario—lamps and thermos in the bedroom, kettles and bowls in the kitchen, laptops and
plants in the office, bottles and fruit in the market.
All trajectory data—synchronized multi-view RGB sequences, proprioceptive states, and action
tokens—exports directly to the standardized LeRobot dataset format, immediately
consumable by modern VLA architectures on a Unitree G1 humanoid (29-DOF upper body, 12-DOF
dexterous hands) with head-stereo and egocentric wrist cameras.