Skip to content

robstrayer/voice-foundations

v1.0.0MIT

Thirteen provider-neutral skills for building, debugging and reviewing AI voice agents: stack selection, turn taking, conversation design, speech pipeline, audio frontends, latency, media debugging, call reliability, data capture, evaluation, cost, phone compliance and security.

voice-agent-evaluation

Design and run proportionate voice agent evaluations covering conversation success, turn taking, tool correctness, audio behavior, failures, and lifecycle cleanup.

voice-agent-security

Review and harden an AI voice agent against abuse specific to phone and voice. Use for a voice agent security review, caller verification and spoofed caller ID, prompt injection or social engineering by phone, PII or card data in transcripts and recordings, provider API keys in a browser voice app, toll fraud and cost abuse, voice cloning consent, or a pre-launch red team.

voice-audio-frontends

Select, configure, and evaluate microphone and incoming-call audio processing for voice agents. Use when echo, background noise, competing speakers, gain changes, or speech suppression affect transcription, VAD, turn detection, and interruptions.

voice-call-reliability

Diagnose and recover voice call failures across signaling, media, transfers, and business actions. Use for silent calls, failed handoffs, duplicate effects, stranded sessions, capacity incidents, or production readiness reviews.

voice-conversation-design

Write and evaluate voice agent prompts for spoken turn taking, concise responses, clarification, tool progress, interruptions, and safe handoffs.

voice-cost-estimation

Estimate what a voice agent costs per minute, per call and per month, and find what drives the bill. Use when someone asks how much a voice agent will cost, compares a managed platform with a self-built stack, sets pricing for clients, sees a bill higher than expected, or plans for scale.

voice-data-capture

Capture phone numbers, emails, names, addresses, dates, amounts, confirmation codes and yes or no answers correctly on a voice call. Use when a voice agent mishears or misspells what callers say, treats a transcript as truth, flips a negation, or fails on long digit strings or spelled codes, or when spoken data is read back, validated or written to a CRM, booking or payment system. Also for slot filling, E.164 phone numbers, spelling alphabets such as NATO, keypad or DTMF entry, SSML say-as read-back, and keyterms or phrase lists for speech recognition. Covers per-field protocols, asking and read-back, recognition aids, validation in code, test fixtures and keeping card data off the model path. Needs no tools. The optional script needs Python 3.

voice-latency-audit

Investigate voice agent response delay using raw turn events, consistent clocks, component timings, and audio playback evidence; use for latency regressions and benchmark audits.

voice-media-debugging

Diagnose silent, distorted, one-way, delayed, or obsolete audio in voice agents by checking codecs, sample rates, transport, queues, and playback. Use when a connection succeeds but the media path fails, or audio continues after an interruption.

voice-phone-compliance

Keep AI phone agents legal and their numbers deliverable. Use when building or reviewing inbound or outbound voice agents, outbound campaigns or dialers, call recording or transcripts, AI disclosure scripts, opt-out and do-not-call handling, HIPAA or PCI voice flows, or when calls show "Spam Likely" or get blocked. Covers TCPA and FCC rules for AI voices, consent, calling hours, state recording and AI-disclosure laws, STIR/SHAKEN and number reputation, and EU, UK and Canada basics. Not legal advice.

voice-speech-pipeline

Select, integrate, and evaluate streaming speech recognition and synthesis for a voice agent, including transcript revisions, language support, pronunciation, text buffering, and playable audio delivery.

voice-stack-selection

Choose a voice AI architecture, speech engine, transport, and hosting model from real constraints. Use when comparing platforms, planning browser or phone agents, estimating total cost, or migrating an existing stack.

voice-turn-taking

Design, debug, and tune voice-agent turn boundaries and interruptions. Use for callers being cut off, slow endpointing, ignored corrections, backchannels stopping speech, unstable streaming transcripts, or audio and tool state diverging after barge-in.