optimizing-attention-flash
Enables Flash Attention for transformer models using PyTorch native scaled_dot_product_attention (PyTorch 2.2+) or the flash-attn library, including multi-query attention, sliding window attention, and FP8 on H100 (FlashAttention-3). Covers profiling speedup, checking accuracy against a baseline, and troubleshooting install and GPU support errors. Use when training or running transformers on long sequences (over 512 tokens), when standard attention runs out of GPU memory, when attention is the inference bottleneck, when switching a PyTorch model to the flash backend, or when tuning attention on H100 GPUs. Not for CPU inference, V100 GPUs, or sequences under 256 tokens; consider xFormers for other attention variants.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/optimizing-attention-flash/SKILL.md
- skills/optimizing-attention-flash/references/benchmarks.md
- skills/optimizing-attention-flash/references/transformers-integration.md
Every link opens the file at its source, pinned to the revision this page describes.