audiopodai/audiopod
Audio AI for agents: text-to-speech, voice cloning, music generation, stem separation, speaker separation, transcription, denoise and media conversion over AudioPod's hosted MCP server, plus 14 task skills.
Convert speech from one voice into another while preserving the words and timing (voice-to-voice).
Clone a voice from a 5β30s reference clip and reuse it for narration or TTS.
Convert audio between formats (MP3, WAV, FLAC, OGG, M4A, AAC).
Turn source documents or a topic into a multi-speaker podcast episode with a public RSS feed.
Remove background noise, hiss, and hum from a recording while preserving the voice.
Create a brand-new voice from a text description, preview candidates, and publish one as a reusable voice.
Music generation, powered by AudioMusic β compose royalty-free songs, instrumentals, and rap from a text prompt.
Convert a manuscript (PDF/EPUB/text) into an ACX-spec audiobook with AI narration, chapter splits, retail sample, and multi-platform export presets.
Turn an article or long text into narrated audio with a shareable player.
Diarize a recording and split it into per-speaker audio tracks.
Split a song into vocals, drums, bass and more, or isolate one instrument from a 45-instrument catalog.
Text to speech, powered by AudioSonic β synthesize natural speech from 200+ voices across 200+ languages, with inline directing and optional word-level timestamps.
Transcribe audio with speaker diarization and word-level timestamps.
Translate or dub spoken audio into another language.