πŸ“° Key Takeaways

The Gemini family is adding two new text-to-speech models β€” “Gemini 3.8 Flash TTS” and “Gemini 3.8 Flash-Lite TTS” β€” upgrading voice generation from fixed preset audio files into a dynamic creative studio, available in products like Gemini Notebook and Google Vids. 3.8 Flash TTS is built for deep creative direction and character design: you can craft brand-new voices from scratch using natural language prompts, covering games, immersive audiobooks, podcasts, and interactive media characters, with line-by-line control over the performance β€” including delivery cues, pacing, accent switching, and response tone. 3.8 Flash-Lite TTS is built for large-scale, low-cost use cases β€” great for bulk voiceover work, audio content production, and expressive voice assistants, with fine-grained control over tone, pacing, and emotional nuance. On the voice library side, the number of original voices has expanded from 30 to nearly limitless, with natural-language prompt design supported across more than 100 languages and dialects. There are also over 2,000 ready-to-use commercial voices covering regional variants like Mexican Spanish, Quebec French, and Scottish English. Voice cloning needs just a 30-second audio sample to reproduce a consistent voice signature, and it comes with built-in consent verification, SynthID watermarking, and C2PA credentials to protect developers and voice actors. Users can also save and manage custom voices, keeping output stable and drift-free for long-term projects. A voice remixing feature is coming next, letting you fine-tune an existing voice’s timbre, pitch, pacing, and accent with a prompt (like adding a slight Southern American accent or slowing down the delivery). This release continues the expansion of the Gemini Audio family, following earlier products like 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.


πŸ’¬ JudyAI Lab Take

The Gemini family just dropped two new text-to-speech models β€” “Gemini 3.8 Flash TTS” and “Gemini 3.8 Flash-Lite TTS” β€” upgrading voice generation from fixed preset audio into a dynamic creative studio.

Based on the source summary, 3.8 Flash TTS is built around crafting character voices from scratch using natural language prompts, with line-by-line control over pacing, accent switching, and response tone β€” great for games and audiobooks. 3.8 Flash-Lite TTS, meanwhile, targets large-scale, low-cost applications. This points to a design shift worth watching for AI builders: voice generation is moving from “pick a preset voice” to “direct a voice performance with prompts,” with an interface logic closer to a creative tool than a plain API call. At the same time, the official rollout ships consent verification, SynthID watermarking, and C2PA credentials alongside the new capabilities β€” a sign that as generative voice products scale up, identity verification and abuse protection are being treated as standard, not an afterthought.

If you’re thinking about building with voice, here’s the thing to consider: if your product uses voice cloning or character voiceover, licensing and watermarking need to be part of the design from day one β€” not a patch you bolt on later.


πŸ“… Original Source


πŸ”— Further Reading