Skip to content

audiopodai/audiopod

v1.0.0MIT

Audio AI for agents: text-to-speech, voice cloning, music generation, stem separation, speaker separation, transcription, denoise and media conversion over AudioPod's hosted MCP server, plus 14 task skills.

change-voice

Convert speech from one voice into another while preserving the words and timing (voice-to-voice).

clone-voice

Clone a voice from a 5–30s reference clip and reuse it for narration or TTS.

convert-media

Convert audio between formats (MP3, WAV, FLAC, OGG, M4A, AAC).

create-podcast

Turn source documents or a topic into a multi-speaker podcast episode with a public RSS feed.

denoise-audio

Remove background noise, hiss, and hum from a recording while preserving the voice.

design-voice

Create a brand-new voice from a text description, preview candidates, and publish one as a reusable voice.

generate-music

Music generation, powered by AudioMusic β€” compose royalty-free songs, instrumentals, and rap from a text prompt.

narrate-audiobook

Convert a manuscript (PDF/EPUB/text) into an ACX-spec audiobook with AI narration, chapter splits, retail sample, and multi-platform export presets.

read-article

Turn an article or long text into narrated audio with a shareable player.

separate-speakers

Diarize a recording and split it into per-speaker audio tracks.

separate-stems

Split a song into vocals, drums, bass and more, or isolate one instrument from a 45-instrument catalog.

text-to-speech

Text to speech, powered by AudioSonic β€” synthesize natural speech from 200+ voices across 200+ languages, with inline directing and optional word-level timestamps.

transcribe-audio

Transcribe audio with speaker diarization and word-level timestamps.

translate-audio

Translate or dub spoken audio into another language.