> ## Documentation Index
> Fetch the complete documentation index at: https://kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# evaluating-code-models — AI agent skill for ml evaluation and safety

> Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics.

# `evaluating-code-models`

> Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass\@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

**Category:** [ml-evaluation-and-safety](/skills#ml-evaluation-and-safety) · **License:** MIT · **Version:** 1.0.0

## Install

```bash theme={null}
npx research-agent-skills install evaluating-code-models
npx skills add KalarisLabs/research-agent-skills --skill evaluating-code-models
```

## When to use it

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass\@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

## Full playbook

Read [SKILL.md](https://github.com/KalarisLabs/research-agent-skills/blob/main/skills/evaluating-code-models/SKILL.md) for the complete workflow, references and any scripts. The agent installer copies the full skill folder.
