gemini-speech
Generates natural speech audio (WAV) with Gemini text-to-speech through the gemini-media MCP server - voiceovers, narration, audiobook passages, announcements, explainer and ad reads, and two-speaker podcast or dialogue scripts, with controllable delivery style, 30 prebuilt voices, inline sounds like sighs and pauses, and automatic language detection. Use when the user wants text read aloud, a voiceover or narration track, TTS, a spoken intro, or a conversation rendered as audio. Not for music or singing (use gemini-music), transcription or speech-to-text, dialogue inside a generated video clip (gemini-video), or live voice chat.
- Version
- 1.0.0
- License
- Apache-2.0
- Compatibility
- Requires the gemini-media MCP server (https://github.com/mordor-forge/gemini-media-mcp) with a Gemini API key or Vertex AI project.
Pinned to revision 747f3b876f54, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/gemini-speech/SKILL.md
- skills/gemini-speech/evals/evals.json
- skills/gemini-speech/evals/files/onboarding_it.txt
- skills/gemini-speech/evals/trigger-queries.json
- skills/gemini-speech/references/voices.md
Every link opens the file at its source, pinned to the revision this page describes.