llamaguard
Classifies LLM prompts and responses as safe or unsafe using Meta's LlamaGuard (7B v1, 8B v2 and v3) across six categories: violence and hate, sexual content, guns and illegal weapons, regulated substances, suicide and self-harm, and criminal planning. Covers HuggingFace Transformers and vLLM inference, a FastAPI moderation endpoint, Sagemaker deployment, and NeMo Guardrails integration. Use when filtering user prompts before an LLM call, moderating model outputs before display, serving a self-hosted moderation API, or reducing latency and GPU memory of a moderation model with vLLM or quantization. Use when tuning thresholds to cut false positives. For a simpler hosted option, use the OpenAI Moderation API instead.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
Every link opens the file at its source, pinned to the revision this page describes.