Gemini TTS
Product website: Create speech with Gemini TTS
Gemini TTS is an independent browser-based text-to-speech application for narration, read-aloud audio, and two-speaker dialogue. Write a script, choose preset voices, direct the delivery, then play or download the completed audio.
The website explicitly states that it is not a Google product. This Hub repository introduces the application at geminitts.app; it does not represent Google, publish a speech model, or provide hosted inference.
Product description
The studio combines script entry with voice and delivery controls. In single-voice mode, one selected voice reads the script. In two-speaker mode, individual lines are assigned to speaker one or two, with separate voice settings for each speaker.
You can listen to neutral voice samples before submitting a task. Completed takes can be played in the studio, downloaded as WAV, and found in the account's library. The website also presents example recordings for a single voice, a Mandarin reading, and a two-voice conversation.
These recordings illustrate the product; they are not a benchmark or a guarantee that a different script will have the same pronunciation, timing, or expressiveness.
Available controls
At the time this card was written, the studio showed the following options. Check the live interface for current availability and limits.
| Control | What it does in the studio |
|---|---|
| Model selector | Offers options labeled Gemini 3.8 Flash and Gemini 3.8 Flash Lite. |
| Voice selection | Provides 30 preset voices with neutral sample playback. |
| Speaker mode | Selects one voice or a two-speaker script with line assignments. |
| Accent | Offers selectable delivery presets for each voice. |
| Style | Includes options such as newscaster, whisper, empathetic, and deadpan. |
| Pace | Provides pacing presets for directing the reading. |
| Vocal cues | Inserts cues such as <laugh>, <sigh>, and <short pause> into the script. |
| Advanced controls | Accepts optional voice descriptions and scene direction. |
The model names above are labels verified in this product's interface, not a statement of official Google release status or an affiliation claim. The website positions Flash for expressive readings and dialogue, and Flash Lite for straightforward or repeated drafts; actual generation time and delivery can vary.
Intended use
- Video narration: prepare a voiceover for an explainer, product walkthrough, or draft edit.
- Stories and read-aloud drafts: listen to a passage and assess its rhythm or tone.
- Two-speaker exchanges: assign a short conversation to two contrasting preset voices.
- Script review: hear whether a sentence is too long, unclear, or awkward before recording a final version elsewhere.
These are suggested workflows. Listen to the completed take and check permissions for the script before publishing or incorporating it into a deliverable.
How to create a take
- Open the studio. Choose a currently available model option and decide whether the script needs one or two speakers.
- Preview and select voices. Listen to a neutral sample. For dialogue, choose a voice for each speaker.
- Write or paste the script. The current interface accepts up to 5,000 characters per generation; in two-speaker mode, the combined lines count toward that limit.
- Direct the performance. Adjust accent, style, and pace. Add a small number of vocal cues where useful, or use the optional voice and scene descriptions.
- Review the estimated credits. Check your balance, sign-in requirements, and the displayed estimate before confirming generation.
- Play the full result. Listen for pronunciation, pauses, speaker assignments, and delivery. Revise the wording or controls if needed.
- Keep the take. Download the WAV from the studio or revisit the completed audio in your library.
Illustrative scripts
The examples below are original writing examples, not generated recordings or tested performance claims.
Single-voice narration
Welcome to the workshop. <short pause> Today, we will turn one small idea into a clear plan. Keep your notebook nearby, and take each step at your own pace.
For a calm reading, begin with restrained settings and listen to the result before adding more directions. The cue indicates the requested pause; it does not specify an exact pause duration.
Two-speaker exchange
Assign the following lines to the indicated speakers using the studio's line controls:
| Speaker | Script line |
|---|---|
| 1 | What should we explain first? |
| 2 | Start with the problem. Then show one useful example. |
| 1 | Good idea. Let us keep the explanation short. |
Check the speaker assignments and listen to the complete exchange. A dialogue draft is easier to review when each line has a clear speaker and purpose.
Tips and limitations
Voice samples use neutral settings, so a sample does not demonstrate every possible combination of accent, style, pace, and script. Controls guide the performance but do not guarantee precise pronunciation, emotional delivery, or timing.
Names, abbreviations, unusual punctuation, and language changes deserve particular attention during review. A generated reading may need script revision or audio editing before it fits a finished video or other timed production.
Use a short, coherent passage when exploring delivery. Adding many competing cues or directions can make a script harder to assess. Check the current character limit rather than treating the studio as an unlimited long-form narration service.
The application uses credits. Starter-credit availability, generation estimates, account requirements, and commercial-use terms are governed by the live service. Public access to this card does not include audio credits or grant rights to every generated recording.
Frequently asked questions
Is Gemini TTS an official Google application?
No. The website describes itself as an independent application built around Gemini speech models.
Can I create two-speaker dialogue?
Yes. The current studio has a two-speaker mode with line assignments and separate voice controls.
What can I download?
The website offers completed audio as WAV and keeps finished takes in the library.
Can I insert pauses or vocal cues?
The script controls include cues such as <laugh>, <sigh>, and <short pause>. Review the generated delivery rather than assuming exact timing.
Is generating audio free without limits?
No such claim is made here. Review the service's current starter credits, pricing, balance, and estimate before submission.
Does this repository contain speech weights or an inference endpoint?
No. It is a README-only reference to the external application.
Get started
Open Gemini TTS to preview a preset voice and prepare your own script.
Repository scope and license
This is a public README-only product-reference repository. It contains no model weights, training code, audio corpus, or hosted inference endpoint. Speech generation takes place on the external website under its own terms.
The MIT license applies to the README content in this repository. It does not license the external application, Google's models or trademarks, voice samples, or generated recordings; their respective rights and terms continue to apply.