Humans effortlessly react to the world; they adeptly return a tennis shot, catch a falling cup, step aside from a passing cart, or pass a box tossed during truck unloading. These everyday actions require anticipating a moving target and reflexively responding with the whole body within about a second. Such visual reactive whole-body control remains challenging for humanoids, which must (1) coordinate whole-body control, (2) react to object dynamics in real time, and (3) ground perception in egocentric vision. We study catching thrown boxes as a representative task that demands all three capabilities; existing systems cover only part of it, relying on end-effectors, motion capture, or marked objects. We present Reflex, a three-stage framework that learns these capabilities stage by stage while connecting them through a compact representation of object dynamics. First, a state-based policy learns whole-body catching with reinforcement learning. Second, the policy learns to infer the dynamics of a moving box from delayed and incomplete observations. Third, an egocentric RGB-D encoder learns to recover this dynamics representation directly from visual observations while keeping the whole-body controller fixed. This decomposition provides each capability with a direct training signal and makes visual learning substantially more scalable. In simulation, Reflex achieves 85.7% catch success from egocentric RGB-D, compared with 87.9% when given privileged box observations. Temporal visual history is critical: removing it reduces success by 35.6 and 39.8 points on centered and laterally offset throws. On a real Unitree G1, Reflex catches human-thrown boxes with 65% and 40% success at 2.5 and 5 m, reaching 76% and 59% of human catchers' success at the same distances.
Egocentric only
2.5 m throw
5 m throw
7 m throw
Inset: the robot's onboard camera view.
Egocentric only
Small parcel (22×21×15 cm)
Regular box (31×26×21 cm)
Large carton (46×41×32 cm)
Payloads of 0.5, 1 and 2 kg, each weighed before the throw
Egocentric only
Two off-target throws, seen from behind the robot
Egocentric only
Catch, carry, hand over, and return for the next throw
Simulated catch success versus lateral landing offset (from the paper; 640 throws per point, 95% intervals). The shaded bands mark the offsets seen in training.
Pick one factor at a time (box size, throw distance, payload, or lateral landing offset) while the others stay at their defaults, press Launch, and orbit the 3D scene to watch the whole-body catch. Precomputed Isaac Sim rollouts of the full visual pipeline (rendered RGB-D → visual encoder → policy), visual encoder trained with throws up to 7 m. Each Launch replays one of the recorded successful catches; overall success rates are reported in the paper.
drag to orbit · scroll to zoom · right-drag to pan · t = 0 at release
@misc{jia2026reflex,
title = {Reflex: Visually Grounded Reactive Humanoid Control},
author = {Jia, Taoyang and Huang, Weikai and Song, Linxin and Darlington, Jared and
Zhao, Jieyu and Wang, Yue and Duan, Jiafei and Ren, Zhongzheng and Krishna, Ranjay},
year = {2026}
}