> ## Documentation Index
> Fetch the complete documentation index at: https://kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# tensorrt-llm — AI agent skill for ml inference and ops

> Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency.

# `tensorrt-llm`

> Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

**Category:** [ml-inference-and-ops](/skills#ml-inference-and-ops) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install tensorrt-llm
npx skills add KalarisLabs/research-agent-skills --skill tensorrt-llm
```

## When to use it

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/tensorrt-llm/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.
