vai-ffn-quantization
Identify and quantize Feed-Forward Network (FFN) layers in ONNX models. An FFN is any MatMul+Add (weight+bias) linear layer: MLP projections, attention QKV projections, attention output projections, and classifier heads. This skill detects all FFN patterns, groups them into blocks, and applies per-layer INT8 quantization.
- Version
- 6.3.0
- License
- Apache-2.0 WITH LLVM-exception
- Compatibility
- Requires VitisAI 6.3+
Pinned to revision 2cdc9eef1b6b, so it is the text this page describes rather than whatever the author pushed since.
Files
Every link opens the file at its source, pinned to the revision this page describes.