Google introduced two new audio models that enable enterprises to deploy intelligent AI agents with reasoning depth, more interactivity and conversational speed.Launched on September 15, Gemini 3.8 Live and Live Extended thinking let developers build voice agents that can reason and execute tasks while maintaining the flow of the conversation, Google said. The models’ key capabilities include performing API and tool calls while continuing to stream audio responses to users, providing live visual inputs that help agents understand what users say and see, and supporting more than 97 languages. The 3.8 Live Extended Thinking version supports configurable thinking, a feature that lets users turn an AI model’s internal reasoning on or off.The release of the audio models comes two weeks after Google initially introduced the Gemini 3.8 Flash and 3.8 Cyber foundation models. With 3.8 Live’s capabilities, Google appears to be catching up to vendors such as Apple, which already offers real-time transcription in productivity apps such as Notes and Voice Memos, said Bradley Shimmin, an analyst at Futurum Group. Still, Google tech takes it to the next level by changing how users interact with AI, he said.Related:Calls for AI slowdown raise new challenges for open-weight modelsReal Conversations“The way it's architected … you can customize this to behave in a lot of different ways,” Shimmin said. He noted that while enterprises can use the models for traditional translation use cases, Google has also designed the models so users don’t interact with it in a prompt-response way; rather, they can have a real conversation.“It's just a natural conversation with the ability to interrupt it in real time to inject,” Shimmin said. “You can inject context into the conversation without interrupting what it's doing.”With Google’s advanced speech technology capability, the new models don’t just answer a question at face value, said Sid Nag, founder of Tekonyx.“The model is designed to do background reasoning,” Nag said. “It’s doing it during the conversation so you can keep the conversation alive while it’s reasoning.”This technology is an advancement in which interactive latency is separated from reasoning latency. Interactive latency is the delay between a user performing an action and the system responding. Reasoning latency is the total time it takes for an AI model to process information.“It’s doing it in real time, rather than halting and then coming back with another answer,” Nag said.The speech models also change how an AI agent responds and interacts because “an AI agent doesn’t necessarily have to choose between being fast and being thoughtful,” Nag continued. “It can actually maintain a real-time interaction.”Related:AI changes the ROI equation. Here’s how some have found successEnterprise applications for the new models include customer support chatbots, sales processes and multimodal agentic applications that combine text, voice and video, Nag said.About the AuthorNews Writer, AI BusinessEsther Shittu has covered AI technologies and industry trends since 2021. As co-host of the Targeting AI podcast, she talks with experts, thought leaders and practitioners exploring critical AI developments. Before AI Business, she wrote for SearchEnterpriseAI, the New York Daily News, Bklyner and the Brooklyn Daily Eagle. When she's not diving deep into the world of AI, she spends her time on passion projects and raising her three daughters.
Gemini 3.8 Live Transforms Conversational AI
Full Article
Original Source
Read the full article at Aibusiness →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.