From capability to creation · Character voices

Give a character a voice, then direct each line: Gemini 3.8 TTS

One knight can have different voices; a monologue can carry directions for each line. Official demos show what creators can control when voicing a short film.

Released September 23; reviewed September 24. These are Google demos. We checked the visuals and playback access, but have not generated our own audio or tested Chinese delivery.

In this article3 chapters

Design the voice separately from the character

In the game demo, an armored knight stands beside a voice description and playback button. At 48 seconds, the description asks for a deep, gravelly warrior voice. At 73 seconds, it asks for a high-pitched, regal princess voice while retaining the knight. Open the video to hear those two moments.

The useful creative idea is that appearance and voice can be adjusted separately. If you already have a character or animated shot, voice becomes another way to shape personality beyond choosing a fixed preset.

An armored knight in Google’s demo, with a Voice Design description and playback control on the right.
Original frame at 48 seconds in Google’s demo. The character and voice description sit together. This is a demo application; the TTS model is not shown generating the game visuals. · Open full-size image
Sources and further reading

After voice design comes the performance

A second official demo divides a detective’s monologue into performance beats. Emotion and volume instructions sit beside the dialogue, with sighs, short pauses and breaths inside it. This shows two layers: who is speaking, then how this particular line is delivered.

Google launched Flash TTS and Flash-Lite TTS on September 23. Flash emphasizes character voice design; Lite targets high-volume generation. Both support line-level direction. The examples show an approach; natural Chinese emotion and consistency across repeated generations still need comparable samples and real projects.

The official monologue demo places performance directions and sigh, short-pause and breath cues beside a line.
Frame at 40 seconds in Google’s monologue demo. The video itself says it is shortened; it illustrates controls, not our test of continuous long-form audio. · Open full-size image
Sources and further reading

Find an entry point and bring a short scene

At launch, both models begin rolling out in AI Studio and the Gemini API. Product distribution pairs Flash with Gemini Notebook and Lite with Google Vids. This does not mean the regular Gemini chat window has every demonstrated control; check the product and account you use.

Try two or three lines of your own dialogue. Keep the script and character fixed, change one pause or emotion, and listen for whether it serves the scene. When returning to our AI film cases, notice how sound helps establish a character. This is a creative direction to explore, not a claim that those older films used the new model.

Sources and further reading