Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS, offering studio-grade voice fidelity, regional accents, and scalable audio solutions for creators and enterprises.
Google has introduced two advanced text-to-speech (TTS) models—Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS—marking a significant leap in AI-powered voice technology. The announcement, made on September 23, 2026, builds on the Gemini Audio family, empowering creators, developers, and enterprises to deliver nuanced, expressive, and scalable audio experiences.
The flagship Gemini 3.8 Flash TTS model targets high-fidelity applications like gaming, audiobooks, and podcasts, offering granular control over voice acting, accents, and pacing. Users can design entirely new voices from scratch using natural language prompts, enabling dynamic character creation for media and storytelling. Meanwhile, the Flash-Lite variant is optimized for cost-effective, high-volume use cases such as dubbing and conversational agents, maintaining fine-grained tonal and emotional control.
Gemini 3.8 Flash TTS is positioned as a "vocal studio" for creative professionals. It supports bespoke voice design with a library of over 2,000 production-ready voices across 100+ languages and dialects, including regional varieties like Mexican Spanish and Scots English. The model also leverages voice replication technology, enabling audio cloning from a 30-second sample while including safeguards like SynthID watermarking and consent verification to ensure ethical use.
For enterprises and developers, Flash-Lite TTS excels at scaling audio solutions for localized media and AI-driven voice agents. Both models support long-form content generation, two-speaker dialogue, and realistic backchanneling, making them suitable for podcasts, virtual assistants, and multi-turn conversations.
Source link







