using-system/vllm-on-tap
Serve vLLM presets on demand - locally, Kubernetes (KServe), Azure (Container Apps), AWS (ECS) or GCP (Cloud Run) - with their traces over OTLP, and tear them down again.
The vllm-on-tap preset - its format, where a preset name resolves (custom in .vot/presets/, then builtin in the official repository), how it is validated and how its per-stack overrides merge. Read when a preset is served, listed, written or checked.
How vllm-on-tap runs vLLM on each stack type, supported or planned - one reference per type with its prerequisites, config fields, serve, readiness, destroy and traps. Read the current stack type's reference only.
vLLM capabilities a vllm-on-tap preset or serve needs - the serve flags, tracing over OTLP, the API key, the OpenAI-compatible API and tuning for a GPU - each linked to vLLM's documentation. Read when writing or adjusting a preset, or when a serve adds tracing or an API key.
Create, update or select a vllm-on-tap environment - its stack type, its config and its optional OTLP traces endpoint - checking and offering to install the tools it needs, and preparing what the stack type needs once. Use when the user wants to configure where presets are served.
Destroy the unit vot-<preset> a vllm-on-tap serve created on the current environment (a process, a container or a cloud app), keeping the environment itself. Use when the user wants a served preset stopped and removed; the argument is the preset name.
Create or edit a vllm-on-tap custom preset in .vot/presets/ from a free-form request - a model of interest, a Hugging Face link, a GPU, a constraint. Use when the user wants a preset written or changed; the argument is the request.
Serve a vLLM preset on the current vllm-on-tap environment - resolve the preset (custom or builtin), start vLLM on the environment's stack as unit vot-<preset>, wait until it answers, and print the endpoint with a ready-to-paste request. Use when the user wants a preset running; the argument is the preset name.