Skip to content

aim-it4/voicestudio-mcp

v0.1.0

Connect VoiceStudio speech, voice cloning, voice design, transcription and its wider API to an MCP client.

VoiceStudio MCP for Desk2Quant

A deployable MCP gateway and private plugin source package for debpalash/VoiceStudio. It keeps inference in VoiceStudio and exposes its existing native tools plus its broader HTTP API over Streamable HTTP or local stdio.

Start here: DEPLOY.md. The ZIP is source, not a running server. After deployment your private MCP endpoint is https://<your-gateway>/<secret>/mcp. No live domain or credentials are embedded in this package.

What is included

CapabilityHow it is accessed
Speech / narrationNative generate_speech; advanced parameters via HTTP API
Voice cloning / voice designNative clone_voice, describe_voice, design_voice
TranscriptionNative transcribe; multipart API for larger recordings
Voice / personality / language discoveryNative list_voices, list_personalities, list_languages
Video dubbing, translation and subtitlesDiscover, inspect and call the live /dub/* and related operations
Audiobooks and long-form productionLive audiobook / longform operations, including finite SSE rendering
Batch jobs, projects, conversion and pronunciationLive API discovery and execution
Models, engines, watermarking, settings and workersLive API discovery; upstream permission and hardware rules apply
Audio, video, subtitles and manuscript filesSigned uploads, repeated multipart file fields and timed downloads
Long operationsBounded background jobs and persistent status without automatic retry
MCP voice/history resourcesvoicestudio_read_resource

There are 9 native tools + 12 gateway tools in the bundled schema. Native discovery also picks up additional tools exposed by your backend. The source catalog contains 337 HTTP routes inspected at the commit in UPSTREAM.json; actual operation IDs, schemas and availability always come from your running backend's OpenAPI document. This is API coverage, not a claim that every desktop feature or every engine is usable on every host.

Package files

  • Dockerfile, railway.json: deploy the small gateway.
  • docker-compose.yml, docker-compose.gpu.yml: gateway + official backend image.
  • .env.example: configuration; scripts/init_env.py generates fresh secrets.
  • plugin.json, mcp.json: portable local plugin with a real stdio entrypoint.
  • scripts/export_remote_plugin.py: verify your deployed endpoint and make a separate remote plugin ZIP containing its actual URL.
  • scripts/smoke_test.py: safe deployed check of discovery and status.
  • scripts/upload_file.py: upload a local file without putting base64 in a chat.
  • CAPABILITIES.md: coverage and limitations.
  • VERIFICATION.md: what was actually tested.
  • uv.lock: frozen Python dependency resolution.

Important limits

  1. A running VoiceStudio backend and installed models are required. The gateway does not contain model weights and cannot synthesize audio by itself.
  2. A small Railway/Render instance can host the gateway. Running the complete VoiceStudio inference workload there is a separate resource decision; do not assume a 1 GB/free instance will support large speech models or video dubbing.
  3. Native desktop controls (microphone widget, OS hotkeys, native file pickers, reveal-in-folder, desktop runtime management) and continuous WebSocket sessions remain in the original app. Finite HTTP/SSE workflows are supported.
  4. The initial mcp.json is local. A ChatGPT web/mobile connection needs the deployed HTTPS endpoint or the remote plugin exported after verification. Installing a ZIP does not start hosting. Availability depends on the host's current custom-MCP/plugin support; this package does not promise mobile-only installation.
  5. Secret URLs are a private single-user access method. They are not OAuth. Treat the full MCP URL as a password. Use an OAuth gateway for shared/public distribution. Optional bearer authentication is supported for clients that send headers; this package does not implement an OAuth authorization server.
  6. Hosting, GPU resources, optional paid providers, telephony and some models can incur costs. Open source does not make these services free.
  7. Upstream is AGPL-3.0; model licenses differ. This gateway source is provided under the included AGPL-3.0 license. Only clone a speaker's voice with permission.

No changes to Desk2Quant, its repository, payments or existing services are needed to use this separate package.

Development

uv sync --frozen
uv run pytest

For local stdio, set VOICESTUDIO_URL to your backend, then run uv run voicestudio-mcp --transport stdio. Logs go to stderr.

Sources

The package pins MCP SDK 1.28.1 to match the inspected upstream integration; it does not depend on the current SDK main branch's v2 interface.