Google DeepMind debuts Gemini Robotics 2 model series for humanoid robots Alphabet Inc.’s artificial intelligence research lab today debuted a family of models optimized to power humanoid robots. Google DeepMind says that the Gemini Robotics 2 series enables multiple autonomous machines to collaborate on a task. According to the company, it can automate chores that comprise hundreds of steps. Many humanoid robots feature a so-called dual-system AI architecture. That means they use two AI models to carry out work. The first model, which is known as an embodied reasoning algorithm, crafts a high-level plan for how to perform a task. It then sends the plan to a so-called VLA model, which turns the instructions into low-level commands for the host robot’s motors. The main highlight of the Gemini Robotics 2 series is an embodied reasoning algorithm called Gemini Robotics ER 2. It enables users to describe the task that a humanoid robot should perform in natural language. According to Google, ER 2 supports tasks that comprise hundreds of steps and take several minutes to complete. The model can split a lengthy chore among several different robots to speed it up. Moreover, a tool calling feature enables ER 2 to access external cloud services. For example, it could use Google Search to clarify parts of a prompt that it doesn’t understand. Humanoid robots require the ability to redo tasks that they don’t complete successfully on the first try. According to DeepMind, its engineers equipped ER 2 with two features that streamline the workflow. The first feature enables the model to track the progress of a task using footage from the host robot’s cameras. If the robot makes a mistake, ER 2 can identify the last step that the machine completed correctly and pick up where it left off. That removes the need to redo chores from scratch, which saves time. The other new feature makes ER 2 better than its predecessor at determining when a task is complete. The faster a humanoid robot’s AI can tick off a task, the sooner it can move on to the next one. The capability also eases certain related tasks such as identifying ways to correct mistakes. “By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step,” Google engineers Steven Hansen and Peng Xu wrote in a blog post today. After ER 2 generates a plan for how to carry out a task, it can send the instructions to one of the two other models in the Gemini Robotics 2 series. They’re VLA algorithms capable of translating action plans into low-level instructions for the host robot. The first model is known as Gemini Robotics 2. Unlike certain earlier models, it can control all of a humanoid robot’s components and not just its hands. As a result, the algorithm can optimize the host machine’s center of gravity in a way that minimizes the risk of falls. Furthermore, Gemini Robotics 2 supports a broader range of robotic hands than DeepMind’s earlier software. The second VLA model that debuted today is called Gemini Robotics On-Device 2. As the name indicates, it’s designed to run directly on humanoid robots’ onboard computers. DeepMind says that the model can be adapted to a new robot with a few hours of training. Developers can access ER 2 via Google Cloud, the Gemini API and Google AI Studio. The company is rolling out the model alongside a new embodied AI safety benchmark. The ASIMOV-Agentic Benchmark, as it’s called, is designed to evaluate human robots’ ability to avoid collisions and other risks. Image: Google A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network. Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links. About SiliconANGLE Media SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
Google DeepMind debuts Gemini Robotics 2 model series for humanoid robots
Full Article
Original Source
Read the full article at Siliconangle →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.