Skip to content

mordor-forge/gemini-media

v1.0.0Apache-2.0

Generate images (Nano Banana), video (Veo, Gemini Omni), speech (Gemini TTS) and music (Lyria) with Google's generative media models, with cost estimates and spend budgets.

gemini-speech

Generates natural speech audio (WAV) with Gemini text-to-speech through the gemini-media MCP server - voiceovers, narration, audiobook passages, announcements, explainer and ad reads, and two-speaker podcast or dialogue scripts, with controllable delivery style, 30 prebuilt voices, inline sounds like sighs and pauses, and automatic language detection. Use when the user wants text read aloud, a voiceover or narration track, TTS, a spoken intro, or a conversation rendered as audio. Not for music or singing (use gemini-music), transcription or speech-to-text, dialogue inside a generated video clip (gemini-video), or live voice chat.

Version
1.0.0
License
Apache-2.0
Compatibility
Requires the gemini-media MCP server (https://github.com/mordor-forge/gemini-media-mcp) with a Gemini API key or Vertex AI project.
Read SKILL.md at the source

Pinned to revision 747f3b876f54, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.