> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# speculative-decoding — AI agent skill for multimodal and emerging

> Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques.

# `speculative-decoding`

> Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

**Category:** [multimodal-and-emerging](/research-agent-skills/skills#multimodal-and-emerging) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install speculative-decoding
npx skills add KalarisLabs/research-agent-skills --skill speculative-decoding
```

## When to use it

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/speculative-decoding/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.