> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kalarislabs.com/llms.txt
> Use this file to discover all available pages before exploring further.

# AI research skills for agents and language models

> Agent skills for LLM evaluation, interpretability, safety checks and clear AI research reporting.

# Artificial intelligence research

For agent and language-model studies, define the behavior you want to measure before asking an agent to run or summarize evaluations. Preserve prompts, model versions, judge settings and raw outputs so a reader can reproduce your conclusions.

| Task | Skill | Starting point |
| - | - | - |
| Compare model or agent behavior | [`evaluating-llms-harness`](/research-agent-skills/skills/evaluating-llms-harness) | Define cases, metrics and human review criteria |
| Inspect model mechanisms | [`transformer-lens-interpretability`](/research-agent-skills/skills/transformer-lens-interpretability) | State a hypothesis and select an interpretable model |
| Assess adversarial inputs | [`prompt-guard`](/research-agent-skills/skills/prompt-guard) | Include benign and attack examples, plus false positive checks |
| Report an AI study | [`ml-paper-writing`](/research-agent-skills/skills/ml-paper-writing) | Provide methods, ablations and uncertainty estimates |

```bash theme={null}
npx skills add KalarisLabs/research-agent-skills --skill evaluating-llms-harness
```

Try: “Design an evaluation for our retrieval agent using the cases in `evals/`. Specify failure categories, scoring rules and which results require a human reviewer.” For training and reproducibility, see [machine learning research](/research-agent-skills/fields/machine-learning).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.