0:00 / 0:36
News
Google's Gemini 3.8 TTS Designs Voices From A Sentence And Clones Them In 30 Seconds
calendar_today Date:
schedule Duration: 0:36
visibility Views: 589
database
Summary Report
Gemini 3.8 Flash TTS designs voices from prompts, clones from 30 seconds of audio and tops Hume AI's voice benchmarks, with 2,000+ voices in 100 languages.
- 01. Voice design from plain-language prompts; consent-checked cloning from 30 seconds of audio; 2,000+ voices in 100 languages
- 02. #1 and #2 on Hume AI's quality index; SynthID watermarking; live in the Gemini API and AI Studio
Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS. Describe a voice in plain language and it designs one, or clone a voice from thirty seconds of audio after a consent check. There are more than 2,000 ready-made voices across 100 languages including regional accents, with line-by-line direction and multi-speaker conversations. Flash TTS takes first place on Hume AI's voice design benchmark and quality index, with Flash-Lite second. Output is watermarked with SynthID.
Meta Data
Company:
LLM: