Google has added a new text-to-speech model to Gemini 3.8, which, while not groundbreaking, refines and improves AI audio technology that already exists.Google introduced Gemini 3.8 TTS and Gemini 3.8 Flash-Lite TTS on September 23 as the tech giant’s most capable audio generation models yet.The models enable creators, developers and enterprises to create higher quality, more expressive audio, the vendor said. Gemini 3.8 Flash TTS is for creative direction and character design. It creates new voices from scratch using natural language prompts across mediums such as gaming, audiobooks, podcasts and interactive media. The model can also customize role, accent and voice characteristics across more than 100 languages and dialects. Gemini 3.8 Flash-Lite TTS is for audio dubbing, content creation and voice agents.Many of the features in Gemini 3.8 Flash TTS and Flash-Lite TTS are not new to the audio or AI speech market. Vendors such as ElevenLabs, Baseten and others provide the same type of audio technology. However, with this release, Google is improving its own AI audio technology, giving enterprise users in entertainment and marketing, among others, more choice and realism if they choose its model.Related:Anthropic, OpenAI launches show shift toward multi-model enterprise AIBetter for enterprises“What we’re seeing here is a refinement of your text-to-speech with some heavy targeting to specific use cases,” said Bradley Shimmin, an analyst at Futurum Group. He said the model is useful for users creating audiobooks and formats such as long-form and short-form narrated fireside chats.“It opens up a lot of interesting opportunities for markets like gaming and entertainment because you suddenly have the opportunity to create highly expressive, highly realistic environments where you can interact with ... and consume audio content that is extremely personable,” Shimmin said.Another selling point for Gemini 3.8 TTS is that it is part of the Gemini family of models.“The industry and the state of the art are going from individual single endpoints like what ElevenLabs offers … to one cohesive model that can do everything you need to with voice,” said Carter Huffman, CEO of Modulate, an AI voice vendor. “Having one cohesive endpoint is a win so long as that one cohesive model gives you excellent performance, ideally best-in-class performance.”For enterprises, the fact that the voice model is in the same model family enables tighter integration, Shimmin said.“Data professionals cite that the biggest challenge they have in using AI is integration,” he said. “Not data integration but integration of technologies”Related:OpenAI takes AI price war to next level with GPT-6 Sol, Luna pricingAlthough the Gemini foundation model allows for integration, the voice AI capabilities are still nascent, Huffman said. So, for Google to really stand out in the market, it needs to provide a capability no one has or an application that hasn’t been seen before, he said.About the AuthorNews Writer, AI BusinessEsther Shittu has covered AI technologies and industry trends since 2021. As co-host of the Targeting AI podcast, she talks with experts, thought leaders and practitioners exploring critical AI developments. Before AI Business, she wrote for SearchEnterpriseAI, the New York Daily News, Bklyner and the Brooklyn Daily Eagle. When she's not diving deep into the world of AI, she spends her time on passion projects and raising her three daughters.
Gemini 3.8 text-to-speech refines voice AI capabilities
Full Article
Original Source
Read the full article at Aibusiness →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.