Skip to main content

blip-2-vision-language

Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAVIS for image captioning, visual question answering, image-text matching, and feature extraction. Use when generating captions for images, building a VQA system, doing zero-shot image-text understanding without task-specific training, matching or retrieving images against text, or fitting a BLIP-2 model into limited GPU memory with INT8/INT4 quantization. Prefer LLaVA or InstructBLIP for instruction-following multimodal chat, and CLIP for plain image-text similarity.
Category: multimodal-and-emerging · License: MIT · Version: 1.0.0

Install

When to use it

Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAVIS for image captioning, visual question answering, image-text matching, and feature extraction. Use when generating captions for images, building a VQA system, doing zero-shot image-text understanding without task-specific training, matching or retrieving images against text, or fitting a BLIP-2 model into limited GPU memory with INT8/INT4 quantization. Prefer LLaVA or InstructBLIP for instruction-following multimodal chat, and CLIP for plain image-text similarity.

Full playbook

Read SKILL.md for the complete workflow, references and any scripts. The agent installer copies the full skill folder.