Chinese robotics firm Galbot has unveiled RoboGesture, a framework for improving nonverbal communication in humanoid robots. The system generates real-time gestures that align with spoken language, helping robots communicate more naturally during face-to-face interactions. RoboGesture aims to address a common limitation in humanoid robotics: movements that can appear repetitive, mechanical, or disconnected from speech. The project is scheduled for presentation at the European Conference on Computer Vision (ECCV) 2026, as researchers explore ways to make humanoid robots more natural and effective in social settings. Makes robots expressive Galbot’s new technology could make humanoid robots more natural during conversations by allowing them to generate gestures that match what they are saying in real time. Called RoboGesture, the system combines speech processing, artificial intelligence, motion generation, and robot safety controls to help humanoids respond with movements that reflect the meaning, rhythm, and tone of spoken language. Developed by researchers from Galbot, Tsinghua University, Peking University, Beijing Institute of Technology, Harbin Institute of Technology, and Shanghai Qi Zhi Institute, RoboGesture is designed specifically for physical humanoid robots. The research team says the technology addresses three major challenges in robot communication: limited training data, the difficulty of connecting speech with appropriate movements, and the safety problems that can arise when AI-generated motions are transferred to real robots. What if humanoid robots could gesture as naturally as they speak? 🤖Introducing RoboGesture — enabling humanoids to listen, respond, and gesture with meaning.RoboGesture closes the listen → respond → gesture loop in ~2 seconds, bringing us closer to natural human–humanoid… pic.twitter.com/zrJ7vzBDlm— Galbot (@GalbotRobotics) September 5, 2026 At the center of the system is a semantic-acoustic processing system that listens directly to speech rather than relying only on its written transcription. This is important because spoken language contains information that text does not capture easily, such as changes in pitch, emphasis, rhythm, and intonation. For example, the same word can carry a different meaning depending on how it is spoken. RoboGesture is designed to capture these signals and use them when deciding how the robot should move. The system first converts incoming speech into audio tokens using the Mimi audio codec. It then processes these signals at different levels. Lower-level features capture quick changes in sound and rhythm, while deeper features identify broader semantic information. The system uses both types of information to determine when a gesture should happen and what kind of movement is appropriate. The researchers trained the semantic component using 300 gesture categories, giving the system a broader range of movements to work with. RoboGesture then sends these signals to a motion-generation model based on a diffusion transformer. Instead of producing movements as a series of fixed motion commands, the model works in a continuous motion space and generates movements based on the incoming speech and the robot’s previous movements. This allows gestures to flow from one action to another while remaining connected to the conversation. Enables safer gestures The researchers also identified a problem they call “modality eclipse.” In this situation, a motion-generation model can become too dependent on the robot’s previous movements and continue repeating similar actions instead of responding to new speech. RoboGesture uses an Anti-Inertia classifier-free guidance masking method to push the model to pay more attention to incoming audio and produce movements based on the current speech. Safety is another important part of the technology. AI-generated movements cannot automatically guarantee that a physical robot will avoid collisions or remain within its movement limits. RoboGesture therefore sends generated motion through a model predictive control (MPC) safety filter before execution. The filter checks movement constraints and adjusts the motion to reduce the risk of self-collision and unstable movements. In testing, it reduced the self-collision frame ratio from 4.16 percent to 0.13 percent. According to the team, the system has been tested on a Unitree G1 humanoid equipped with BrainCo dexterous hands. Its complete interaction pipeline combines automatic speech recognition, a language model, text-to-speech, audio processing, gesture generation, and the MPC safety filter. This allows the robot to listen to a person, generate a conversational response, produce matching speech and gestures, and execute those movements on its body.RoboGesture aims is make humanoid robots better suited for face-to-face interaction, where communication depends on more than spoken words. Researchers say that by connecting speech signals directly with physical movement while adding a safety layer for real-world operation, the technology offers a way to make robot conversations more expressive and responsive.Get the latest in engineering, tech, space & science - delivered daily to your inbox.Jijo is an automotive and business journalist based in India. Armed with a BA in History (Honors) from St. Stephen's College, Delhi University, and a PG diploma in Journalism from the Indian Institute of Mass Communication, Delhi, he has worked for news agencies, national newspapers, and automotive magazines. In his spare time, he likes to go off-roading, engage in political discourse, travel, and teach languages.
Galbot’s new system helps humanoid robots move naturally during conversations
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.