Skip to content

hopper-inc/hopper

v1.2.3MIT

LLM, text-to-speech and speech-to-text for voice agents and AI agents: set Hopper up in a voice-agent project, measure time to first token on its own prompt, find what slows it down, and speak or transcribe audio.

hopper-benchmark

Measure a voice agent's LLM time to first token on Hopper with its own system prompt and tools over a simulated 10-turn call, and report first-turn and later-turn latency and the prompt-cache hit rate. Changes no code. Use when the user asks "how fast would my agent be on Hopper", "benchmark TTFT", "is prompt caching working" or wants before/after numbers. Not for STT/TTS latency or load tests; to switch the agent to Hopper, use hopper-integrate.

hopper-diagnose

Find what delays a voice agent's first spoken word: timestamps or IDs at the top of the prompt, a new client per call, HTTP/1.1, no warm-up, thinking left on, tools named but not sent. Ranks fixes by the latency they recover. Works with any LLM provider, needs no key, and edits nothing until the user picks fixes. Use when the user reports "long pauses", "dead air", "the first reply is slow" or asks for a latency review of a voice agent.

hopper-integrate

Switch an existing voice agent's LLM to Hopper, an OpenAI-compatible endpoint built for low time to first token. Gets a key for the project with no human step, benchmarks the agent's own prompt and tools, and edits code only after the user says go. Use when the user says "make my agent respond faster", "set up Hopper", "try Hopper" or "swap the LLM" in a Pipecat, LiveKit Agents, Vapi or OpenAI-SDK voice agent. For numbers only, use hopper-benchmark; to find what's slow, use hopper-diagnose.

hopper-speak

Turn text into speech with Hopper text-to-speech: play it aloud, save a WAV file, or preview and pick a voice. Needs no setup: the first use registers a key for the agent itself. Use when the user says "read this aloud", "say this out loud", "generate a voiceover", "make an audio file of…", "TTS" or "which voice should I use". Not for writing a speech (plain text), transcribing audio (hopper-transcribe) or a voice agent's LLM (hopper-integrate).

hopper-transcribe

Transcribe speech in an audio or video file with Hopper speech-to-text, with word timestamps. Takes WAV, MP3, M4A, voice memos and video (decoded with ffmpeg). Needs no setup: the first use registers a key for the agent itself. Use when the user says "transcribe this", "what does this recording say", "speech to text", "STT", "summarize this call" or "caption this". Not for generating audio (hopper-speak) or measuring LLM latency (hopper-benchmark).