> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# sglang — AI agent skill for ml inference and ops

> Fast structured generation and serving for LLMs with RadixAttention prefix caching.

# `sglang`

> Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.

**Category:** [ml-inference-and-ops](/research-agent-skills/skills#ml-inference-and-ops) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install sglang
npx skills add KalarisLabs/research-agent-skills --skill sglang
```

## When to use it

Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/sglang/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.