Guides
Browser calls
Put a voice call with your agent in your own web page, and a chat beside it.
The widget is one script tag, with its own look. When you want the agent inside your own product
instead — your design, your buttons — your server starts the call and your page joins it with
@aigently/web.
Start the call on your server
#!/bin/sh
# Start a browser call for your page. From your server, with a live key: give the page the url and token, and it joins with @aigently/web.
curl -sS --fail-with-body -X POST "https://api.aigently.ai/v1/web-calls" \
-H "Authorization: Bearer $AIGENTLY_API_KEY" \
-H "Idempotency-Key: visit-20931" \
--json '{
"agent_id": "3c7d9e1f-5a2b-4c6d-8e0f-1a3b5c7d9e2f",
"variables": {
"first_name": "Sara"
},
"metadata": {
"session": "S-20931"
}
}'
It takes what a chat takes — agent_id, variables, metadata, external_id and language — and
the agent needs Browser voice among its channels, on its Configure tab. The values are set here,
by your server; the page never sends any. It needs calls:write and a live key: a browser call is
a real conversation, on your organization's credit. An Idempotency-Key makes a retry return the same
call.
The answer is the conversation, pending until the page joins, a
url and a token. The token joins this one call for 15 minutes. Give it to the page that will
talk, and to nobody else — and never put an API key in a page.
| Refusal | Why |
|---|---|
browser_voice_off |
Browser voice is off for this agent. |
agent_not_published |
Publish the agent first. |
concurrency_limit_reached |
As many browser calls are going as the agent or your organization may have. |
daily_limit_reached |
Today's conversations on the platform's own provider keys are used up. |
insufficient_credit |
Your organization's credit is used up. |
Join it from the page
import { joinCall } from "@aigently/web";
// Your own endpoint, which calls POST /v1/web-calls and returns url and token.
const { url, token } = await (await fetch("/calls/start", { method: "POST" })).json();
const call = await joinCall(
{ url, token },
{
onState: (state) => render(state),
onTranscript: (line) => show(line),
onAudioBlocked: (blocked) => playButton.toggleAttribute("hidden", !blocked),
},
);
playButton.onclick = () => call.startAudio();
hangUp.onclick = () => call.end();
The browser asks for the microphone when the call is joined.
state |
What it means |
|---|---|
connecting |
Joining the call. |
waiting |
Joined; the agent is getting ready. |
listening |
The agent is listening. |
thinking |
The agent is working out what to say. |
speaking |
The agent is talking. |
ended |
The call is over. |
Each transcript line has a speaker — you or agent — and grows while it is spoken, then arrives
once more with final: true. Keep the latest copy of each line by its id. If the browser blocks
sound until the person clicks, onAudioBlocked(true) says so: call startAudio() from a click.
With React
import { useWebCall } from "@aigently/web/react";
function Call() {
const { state, transcript, start, end } = useWebCall();
const handleStart = async () => {
const credentials = await (await fetch("/calls/start", { method: "POST" })).json();
await start(credentials);
};
return (
<>
<button onClick={handleStart} disabled={state !== "idle" && state !== "ended"}>Talk</button>
<button onClick={end}>Hang up</button>
{transcript.map((line) => (
<p key={line.id}>{line.speaker}: {line.text}</p>
))}
</>
);
}
The call ends by itself when the component unmounts.
A chat on the same page
import { startChat } from "@aigently/web";
const chat = await startChat({ agentId, publishableKey });
show(chat.opening);
const { reply } = await chat.send("When do you open?");
A chat uses the widget's own session: the agent's publishable key — the one in your widget snippet
— from a site listed on the agent's Connect tab. Values a visitor could type go in variables;
values your server vouches for go in a signed identity token, as for the
widget.
Getting @aigently/web
@aigently/web will be published with our other SDKs. Until then, a browser call works with any
LiveKit client: join url with token, publish the microphone, and play the agent's audio.