Reflex

Visually Grounded Reactive Humanoid Control

Teaser

Human-thrown boxes. Whole-body catching.
Onboard RGB-D sensing. About one second to react.
1University of Washington, 2Stanford University, 3University of Southern California,
4National University of Singapore, 5University of North Carolina at Chapel Hill, 6Allen Institute for AI
* Equal contribution.

Abstract

Humans effortlessly react to the world; they adeptly return a tennis shot, catch a falling cup, step aside from a passing cart, or pass a box tossed during truck unloading. These everyday actions require anticipating a moving target and reflexively responding with the whole body within about a second. Such visual reactive whole-body control remains challenging for humanoids, which must (1) coordinate whole-body control, (2) react to object dynamics in real time, and (3) ground perception in egocentric vision. We study catching thrown boxes as a representative task that demands all three capabilities; existing systems cover only part of it, relying on end-effectors, motion capture, or marked objects. We present Reflex, a three-stage framework that learns these capabilities stage by stage while connecting them through a compact representation of object dynamics. First, a state-based policy learns whole-body catching with reinforcement learning. Second, the policy learns to infer the dynamics of a moving box from delayed and incomplete observations. Third, an egocentric RGB-D encoder learns to recover this dynamics representation directly from visual observations while keeping the whole-body controller fixed. This decomposition provides each capability with a direct training signal and makes visual learning substantially more scalable. In simulation, Reflex achieves 85.7% catch success from egocentric RGB-D, compared with 87.9% when given privileged box observations. Temporal visual history is critical: removing it reduces success by 35.6 and 39.8 points on centered and laterally offset throws. On a real Unitree G1, Reflex catches human-thrown boxes with 65% and 40% success at 2.5 and 5 m, reaching 76% and 59% of human catchers' success at the same distances.

Reflex catches boxes thrown from 2.5 to 7 meters

Egocentric only

2.5 m throw

5 m throw

7 m throw

Inset: the robot's onboard camera view.

Reflex handles different box sizes and payloads

Egocentric only

Small parcel (22×21×15 cm)

Regular box (31×26×21 cm)

Large carton (46×41×32 cm)

Payloads of 0.5, 1 and 2 kg, each weighed before the throw

Reflex steps to catch off-target throws

Egocentric only

Two off-target throws, seen from behind the robot

From catching to carrying and handing over

Egocentric only

Catch, carry, hand over, and return for the next throw

Catch success in simulation as throws land farther to the side

Simulated catch success versus lateral landing offset (from the paper; 640 throws per point, 95% intervals). The shaded bands mark the offsets seen in training.

Interactive demo: throw a box at Reflex

Pick one factor at a time (box size, throw distance, payload, or lateral landing offset) while the others stay at their defaults, press Launch, and orbit the 3D scene to watch the whole-body catch. Precomputed Isaac Sim rollouts of the full visual pipeline (rendered RGB-D → visual encoder → policy), visual encoder trained with throws up to 7 m. Each Launch replays one of the recorded successful catches; overall success rates are reported in the paper.

drag to orbit · scroll to zoom · right-drag to pan · t = 0 at release

BibTeX

@misc{jia2026reflex,
  title   = {Reflex: Visually Grounded Reactive Humanoid Control},
  author  = {Jia, Taoyang and Huang, Weikai and Song, Linxin and Darlington, Jared and
             Zhao, Jieyu and Wang, Yue and Duan, Jiafei and Ren, Zhongzheng and Krishna, Ranjay},
  year    = {2026}
}