Google AgentHands Gives AI Assistants Virtual Gestures
Google researchers have developed AgentHands, a prototype that generates synchronized 3D hand gestures for XR assistants to help users complete complex physical tasks.

Google researchers have introduced AgentHands, an AI-powered prototype designed to give virtual assistants expressive, synchronized hand gestures in extended reality environments. Developed by a team including Xun Qian, Ruofei Du, and student researcher Ziyi Liu, the project is slated for publication at the CHI 2026 conference. The system aims to bridge the gap between abstract verbal instructions and physical actions, moving beyond flat 2D screen overlays to create an embodied, spatially aware dialogue for platforms like Android XR.
The AgentHands workflow begins with a lightweight object registration module. By utilizing eye gaze and scene reconstruction, users can tag physical items to create a spatial registry of 3D bounding boxes. When a user asks a question, a backend large language model generates a response embedded with inline gesture events linked to specific trigger words. A local parser on the XR headset then uses word-level timestamps to coordinate text-to-speech playback with the animation engine, ensuring the virtual hands perform co-speech gestures in perfect synchronization.
The system utilizes a gesture library categorized into deictic, iconic, and expression behaviors. To demonstrate its utility, researchers showcased the prototype in several scenarios. In an orchid-care task, the agent pointed to the base of a plant to outline its aerial roots. For a 3D printer technical walkthrough, the virtual hands demonstrated a precise turn and click sequence to navigate control knobs. In a lifestyle coaching scenario, the agent performed an interactive warning gesture, accompanied by a red visual glow, to caution a user.
To evaluate the system, the researchers conducted a within-subjects study with 12 participants, comparing AgentHands against a speech-only baseline. Participants completed procedural tasks involving orchid care and 3D printer operations. The results showed statistically significant improvements, with p-values under 0.05 for both spatial grounding and action understanding. Participants found it significantly easier to locate objects and follow complex instructions, noting that the gestural warnings and visual effects effectively reduced their cognitive load.
This is our own summary of reporting by Google Research



