Skip to content

using-system/vllm-on-tap

v0.2.3MIT

Serve vLLM presets on demand - locally, Kubernetes (KServe), Azure (Container Apps), AWS (ECS) or GCP (Cloud Run) - with their traces over OTLP, and tear them down again.

vllm-guide

vLLM capabilities a vllm-on-tap preset or serve needs - the serve flags, tracing over OTLP, the API key, the OpenAI-compatible API and tuning for a GPU - each linked to vLLM's documentation. Read when writing or adjusting a preset, or when a serve adds tracing or an API key.

Read SKILL.md at the source

Pinned to revision 2046ecc85212, so it is the text this page describes rather than whatever the author pushed since.

Files

Every link opens the file at its source, pinned to the revision this page describes.