text-to-speech
Generate local Kokoro narration as one WAV file per JSON segment and build measured timing manifests. Use for offline voiceover, narration clips, mixed speakers, or timing metadata for OfficeKit videos and slides.
- Version
- 0.2.0
- Compatibility
- Requires an OfficeKit workspace with uv. Initial Kokoro model/voice downloads need network access; some languages require espeak-ng or Misaki extras.
Pinned to revision dc4f73ca376e, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/text-to-speech/SKILL.md
- skills/text-to-speech/evals/evals.json
- skills/text-to-speech/evals/files/mixed-voices.json
- skills/text-to-speech/evals/files/protected-clip/intro.wav
- skills/text-to-speech/evals/files/segments.json
- skills/text-to-speech/scripts/synthesize.py
- skills/text-to-speech/scripts/tts_manifest.py
- skills/text-to-speech/voices.json
Every link opens the file at its source, pinned to the revision this page describes.