All pages

Guides

Evaluations in your pipeline

Rehearse an agent's saved conversations on every change, and fail the build when one no longer passes.

View as Markdown

An evaluation is a short conversation written down: what a caller says, and what must be true of the agent's replies. The console calls them saved conversations, and publishing already rehearses every one and refuses while one fails. Through the API, your build pipeline can rehearse them too — on a pull request, before a deploy, or every night — and stop when the agent no longer does what it used to.

The key it needs

Run evaluations (evaluations:run) rehearses an agent's evaluations and reads the results. It is ticked on its own, because every run is real model calls, charged to your credit under Testing. A test key can hold it: rehearsing changes nothing real.

Adding and deleting evaluations changes what publishing lets through, so it takes Change agents (agents:write) and a live key, as other changes to an agent do.

In a build step

curl -O https://developers.aigently.ai/aigently.mjs
node aigently.mjs evaluate 8a1f3c2e-…

evaluate starts a run, waits for it, prints each conversation's outcome with the reason for every expectation it missed, and exits:

Exit code Means
0 Every evaluation passed.
1 At least one failed, or the request was refused.
3 None failed, but some could not be checked.

A conversation that could not be checked says nothing about the agent: the provider had a bad minute, the run ran out of time, or the credit was used up. Treat 3 as a pass or a failure, whichever your pipeline needs. For example, in GitHub Actions:

- name: Rehearse the front desk
  run: |
    curl -sO https://developers.aigently.ai/aigently.mjs
    node aigently.mjs evaluate 8a1f3c2e-…
  env:
    AIGENTLY_API_KEY: ${{ secrets.AIGENTLY_API_KEY }}

Through the API

POST /v1/agents/{id}/evaluation-runs answers at once with a run whose status is queued. Read it with GET /v1/evaluation-runs/{id} every few seconds until status is finished; then passed, failed and not_checked count the outcomes, and results has every conversation's outcome, its verdicts — one per expectation, each with the judge's why — and the rehearsal's transcript.

  • One run at a time in an organization, from the console or the API. Asking while one is going is refused with evaluation_run_in_progress. Wait for that run to finish, then ask again.
  • The agent as it is now is rehearsed, published or not: replace a draft's definition, rehearse it, and publish only if it passes (Agents as code).
  • A run is kept for 30 days. Each conversation's latest result also shows in the console's editor, as one from the Rehearse all button does.

Keep them in your repository

GET /v1/agents/{id}/evaluations lists them, each with its latest result. POST /v1/agents/{id}/evaluations adds one — a name, the caller's turns and the expectations — and DELETE removes one. An agent keeps at most 50; past that, adding one is refused with evaluation_limit_reached.