Text-to-speech generation (TTS)  |  Gemini API  |  Google AI for Developers

Out of the box, the model will natively interpret a transcript and determine how your words should be delivered. Simple transcripts without any additional prompting sound natural. But Gemini TTS also comes with tools you can use to steer it. The purpose of this guide is to offer fundamental direction and spark ideas when developing audio experiences. We’ll start with Tags for quick inline control, and then explore advanced Prompting structures for full performance direction.

Source: Text-to-speech generation (TTS)  |  Gemini API  |  Google AI for Developers

Leave a Reply