# Evaluations in your pipeline

> Rehearse an agent's saved conversations on every change, and fail the build when one no longer passes.

An **evaluation** is a short conversation written down: what a caller says, and what must be true
of the agent's replies. The console calls them saved conversations, and publishing already rehearses
every one and refuses while one fails. Through the API, your build pipeline can rehearse them too —
on a pull request, before a deploy, or every night — and stop when the agent no longer does what it
used to.

## The key it needs

**Run evaluations** (`evaluations:run`) rehearses an agent's evaluations and reads the results. It is
ticked on its own, because every run is real model calls, charged to your credit under Testing. A
test key can hold it: rehearsing changes nothing real.

Adding and deleting evaluations changes what publishing lets through, so it takes **Change agents**
(`agents:write`) and a **live** key, as other changes to an agent do.

## In a build step

```sh
curl -O https://developers.aigently.ai/aigently.mjs
node aigently.mjs evaluate 8a1f3c2e-…
```

`evaluate` starts a run, waits for it, prints each conversation's outcome with the reason for every
expectation it missed, and exits:

| Exit code | Means |
|---|---|
| `0` | Every evaluation passed. |
| `1` | At least one failed, or the request was refused. |
| `3` | None failed, but some could not be checked. |

A conversation that could not be checked says nothing about the agent: the provider had a bad
minute, the run ran out of time, or the credit was used up. Treat `3` as a pass or a failure,
whichever your pipeline needs. For example, in GitHub Actions:

```yaml
- name: Rehearse the front desk
  run: |
    curl -sO https://developers.aigently.ai/aigently.mjs
    node aigently.mjs evaluate 8a1f3c2e-…
  env:
    AIGENTLY_API_KEY: ${{ secrets.AIGENTLY_API_KEY }}
```

## Through the API

`POST /v1/agents/{id}/evaluation-runs` answers at once with a run whose `status` is `queued`.
Read it with `GET /v1/evaluation-runs/{id}` every few seconds until `status` is `finished`; then
`passed`, `failed` and `not_checked` count the outcomes, and `results` has every conversation's
`outcome`, its `verdicts` — one per expectation, each with the judge's `why` — and the rehearsal's
`transcript`.

- **One run at a time in an organization**, from the console or the API. Asking while one is going
  is refused with `evaluation_run_in_progress`. Wait for that run to finish, then ask again.
- **The agent as it is now** is rehearsed, published or not: replace a draft's definition, rehearse
  it, and publish only if it passes ([Agents as code](/guides/agents-as-code)).
- **A run is kept for 30 days.** Each conversation's latest result also shows in the console's
  editor, as one from the Rehearse all button does.

## Keep them in your repository

`GET /v1/agents/{id}/evaluations` lists them, each with its latest result.
`POST /v1/agents/{id}/evaluations` adds one — a `name`, the caller's `turns` and the
`expectations` — and `DELETE` removes one. An agent keeps at most 50; past that, adding one is
refused with `evaluation_limit_reached`.
