> ## Documentation Index
> Fetch the complete documentation index at: https://kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# blip-2-vision-language — AI agent skill for multimodal and emerging

> Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAV…

# `blip-2-vision-language`

> Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAVIS for image captioning, visual question answering, image-text matching, and feature extraction. Use when generating captions for images, building a VQA system, doing zero-shot image-text understanding without task-specific training, matching or retrieving images against text, or fitting a BLIP-2 model into limited GPU memory with INT8/INT4 quantization. Prefer LLaVA or InstructBLIP for instruction-following multimodal chat, and CLIP for plain image-text similarity.

**Category:** [multimodal-and-emerging](/skills#multimodal-and-emerging) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install blip-2-vision-language
npx skills add KalarisLabs/research-agent-skills --skill blip-2-vision-language
```

## When to use it

Explains how to use Salesforce BLIP-2 (Q-Former bridging a frozen image encoder and an LLM such as OPT or FlanT5) through HuggingFace Transformers and LAVIS for image captioning, visual question answering, image-text matching, and feature extraction. Use when generating captions for images, building a VQA system, doing zero-shot image-text understanding without task-specific training, matching or retrieving images against text, or fitting a BLIP-2 model into limited GPU memory with INT8/INT4 quantization. Prefer LLaVA or InstructBLIP for instruction-following multimodal chat, and CLIP for plain image-text similarity.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/blip-2-vision-language/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.
