serving-llms-vllm
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/serving-llms-vllm/SKILL.md
- skills/serving-llms-vllm/references/optimization.md
- skills/serving-llms-vllm/references/quantization.md
- skills/serving-llms-vllm/references/server-deployment.md
- skills/serving-llms-vllm/references/troubleshooting.md
Every link opens the file at its source, pinned to the revision this page describes.