> ## Documentation Index
> Fetch the complete documentation index at: https://kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# quantizing-models-bitsandbytes — AI agent skill for ml training

> Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss.

# `quantizing-models-bitsandbytes`

> Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.

**Category:** [ml-training](/skills#ml-training) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install quantizing-models-bitsandbytes
npx skills add KalarisLabs/research-agent-skills --skill quantizing-models-bitsandbytes
```

## When to use it

Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 formats, QLoRA training, and 8-bit optimizers. Works with HuggingFace Transformers.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/quantizing-models-bitsandbytes/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.
