Guides
Evaluations in your pipeline
Rehearse an agent's saved conversations on every change, and fail the build when one no longer passes.
An evaluation is a short conversation written down: what a caller says, and what must be true of the agent's replies. The console calls them saved conversations, and publishing already rehearses every one and refuses while one fails. Through the API, your build pipeline can rehearse them too — on a pull request, before a deploy, or every night — and stop when the agent no longer does what it used to.
The key it needs
Run evaluations (evaluations:run) rehearses an agent's evaluations and reads the results. It is
ticked on its own, because every run is real model calls, charged to your credit under Testing. A
test key can hold it: rehearsing changes nothing real.
Adding and deleting evaluations changes what publishing lets through, so it takes Change agents
(agents:write) and a live key, as other changes to an agent do.
In a build step
curl -O https://developers.aigently.ai/aigently.mjs
node aigently.mjs evaluate 8a1f3c2e-…
evaluate starts a run, waits for it, prints each conversation's outcome with the reason for every
expectation it missed, and exits:
| Exit code | Means |
|---|---|
0 |
Every evaluation passed. |
1 |
At least one failed, or the request was refused. |
3 |
None failed, but some could not be checked. |
A conversation that could not be checked says nothing about the agent: the provider had a bad
minute, the run ran out of time, or the credit was used up. Treat 3 as a pass or a failure,
whichever your pipeline needs. For example, in GitHub Actions:
- name: Rehearse the front desk
run: |
curl -sO https://developers.aigently.ai/aigently.mjs
node aigently.mjs evaluate 8a1f3c2e-…
env:
AIGENTLY_API_KEY: ${{ secrets.AIGENTLY_API_KEY }}
Through the API
POST /v1/agents/{id}/evaluation-runs answers at once with a run whose status is queued.
Read it with GET /v1/evaluation-runs/{id} every few seconds until status is finished; then
passed, failed and not_checked count the outcomes, and results has every conversation's
outcome, its verdicts — one per expectation, each with the judge's why — and the rehearsal's
transcript.
- One run at a time in an organization, from the console or the API. Asking while one is going
is refused with
evaluation_run_in_progress. Wait for that run to finish, then ask again. - The agent as it is now is rehearsed, published or not: replace a draft's definition, rehearse it, and publish only if it passes (Agents as code).
- A run is kept for 30 days. Each conversation's latest result also shows in the console's editor, as one from the Rehearse all button does.
Keep them in your repository
GET /v1/agents/{id}/evaluations lists them, each with its latest result.
POST /v1/agents/{id}/evaluations adds one — a name, the caller's turns and the
expectations — and DELETE removes one. An agent keeps at most 50; past that, adding one is
refused with evaluation_limit_reached.