blip-2-vision-language
Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAVIS for image captioning, visual question answering, image-text matching, and feature extraction. Use when generating captions for images, building a VQA system, doing zero-shot image-text understanding without task-specific training, matching or retrieving images against text, or fitting a BLIP-2 model into limited GPU memory with INT8/INT4 quantization. Prefer LLaVA or InstructBLIP for instruction-following multimodal chat, and CLIP for plain image-text similarity.
- Version
- 1.0.0
- License
- MIT
Pinned to revision df088027ff23, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/blip-2-vision-language/SKILL.md
- skills/blip-2-vision-language/references/advanced-usage-2.md
- skills/blip-2-vision-language/references/advanced-usage.md
- skills/blip-2-vision-language/references/troubleshooting.md
Every link opens the file at its source, pinned to the revision this page describes.