Google has released two new text-to-speech models that let developers build custom AI voices from scratch using nothing but text descriptions. Gemini 3.8 Flash TTS and Flash-Lite TTS represent a significant shift in how voice generation works, moving away from selecting from preset options toward generative design.

The Flash TTS model generates new voices directly from text prompts. Instead of choosing from a library of pre-recorded voices, users describe the voice they want. A prompt might specify age, accent, tone, emotional quality, or gender presentation. The model then synthesizes a matching voice. This approach democratizes voice creation for smaller studios and independent creators who previously needed either voice actors or access to expensive commercial voice banks.

The lighter Flash-Lite variant offers a stripped-down version for cost-conscious developers. Both support over 100 languages, making global content creation feasible without sourcing voice talent across regions. This matters for game studios, audiobook producers, education platforms, and accessibility tools that need voices in multiple languages quickly.

Stage directions represent a second innovation. Users can annotate individual lines with instructions like "whisper this," "say angrily," or "pause here." The TTS engine respects these cues, adding emotional variation and dramatic timing without requiring multiple recordings. A single script becomes dynamic and expressively varied.

Two-voice dialogue generation creates conversations from a single script file. The system assigns voices to different characters and handles back-and-forth exchanges. This feature speeds up production for podcasts, audiobooks, and interactive fiction where character differentiation matters.

Google also embedded voice cloning into the system. Users provide a 30-second audio sample from any voice, and the model builds a profile to reproduce that speaker's characteristics. This enables several use cases: creating companion narration matching an existing voice, preserving voices for cultural or historical preservation, or matching new dialogue to existing character performances.

The combination of these features lowers barriers across content production. A solo creator can now produce fully voiced dialogue scenarios without hiring actors. Indie game developers can voice multiple characters affordably. Educational platforms can generate explanatory content in target languages without translation bottlenecks.

Technical limitations remain. Voice quality depends on description specificity and the model's training data diversity. Cloning from short samples works better with clear recordings than noisy or heavily accented audio. Generated voices still lack some of the subtle humanity in professional performances, though the gap narrows with each generation.

The announcement arrives as major AI labs compete to dominate voice synthesis. OpenAI's voice features for ChatGPT, Microsoft's voice options in Copilot, and ElevenLabs' specialized voice API all compete in similar space. Google's advantage lies in integration with Gemini and its scale. Developers already building with Gemini can layer in voice generation without switching platforms.

Deployment happens through Google Cloud's APIs, making the tool accessible via simple integration. Pricing follows Google's standard consumption model, charging per character generated. This means costs scale with usage, attractive for variable workloads but potentially expensive for high-volume production.

The models represent synthesis, not recording technology. Ethical considerations around deepfakes and voice misuse apply. Google likely has safeguards, though the specifics remain unclear from the announcement. Consent for voice cloning and disclosure requirements for synthetic voices in commercial content remain evolving legal territory.