> ## Documentation Index
> Fetch the complete documentation index at: https://kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# evaluating-llms-harness — AI agent skill for ml evaluation and safety

> Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag).

# `evaluating-llms-harness`

> Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

**Category:** [ml-evaluation-and-safety](/skills#ml-evaluation-and-safety) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install evaluating-llms-harness
npx skills add KalarisLabs/research-agent-skills --skill evaluating-llms-harness
```

## When to use it

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/evaluating-llms-harness/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.
