Gemini 3.8 text-to-speech says hello
Summary
Google introduced two new text-to-speech models: Gemini 3.8 Flash TTS, built for deep creative direction and character design, and Gemini 3.8 Flash-Lite TTS, optimized for high-volume, cost-efficient dubbing and voice agents. Flash TTS enables users to create custom voices from scratch using natural language prompts across over 100 languages and dialects, access 2,000+ production-ready voices, replicate voices from 30-second samples with built-in consent verification and SynthID watermarking, and direct performances line by line with granular control over acting cues, pacing, and dialect shifts. Both models support long-form generation, native two-speaker scene staging, and scripted vocal bursts and backchanneling for realistic conversational texture. Gemini 3.8 Flash TTS secured the #1 spot on Hume AI's Voice Design Benchmark and accent modeling rankings, while both models ranked #1 and #2 respectively on Hume AI's Overall Quality Index. The models are available through Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids, with partnerships including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang.
(Source:Gemini)