Bevia API reference
Bevia is a conversation platform: voice and messaging agents, plus the speech and language APIs they run on. This reference covers the REST API — OpenAI-compatible chat completions, speech-to-text, text-to-speech, the model and voice catalogues, and realtime voice-agent sessions. Every service works in Arabic and English.
All requests go to a single base URL. The paths in this reference are relative to it — every documented endpoint lives under /api/v1.
https://api.beviaai.com
Bodies are JSON unless an endpoint says otherwise, and every endpoint requires authentication. A quick sanity check once you have a key:
curl https://api.beviaai.com/api/v1/models \
-H "Authorization: Bearer bv_your_key"Authentication
Send your credential in the Authorization header on every request:
Authorization: Bearer bv_your_keyAPI keys start with bv_. Create and revoke them in the dashboard under Keys. A key is shown once at creation and is stored hashed — treat it like a password and keep it out of client-side code and version control.
Each key carries a set of scopes, chosen at creation. A request to an endpoint outside the key’s scopes fails with 403. The dashboard session token carries every scope, so the dashboard itself uses the same endpoints listed here.
| Scope | Unlocks |
|---|---|
| read | GET /api/v1/models, GET /api/v1/voices |
| llm | POST /api/v1/chat/completions |
| tts | POST /api/v1/tts |
| stt | POST /api/v1/stt |
| voice | POST /api/v1/voice/sessions |
Sending WhatsApp messages, templates and opt-outs under /api/v1/whatsapp |
Errors
Errors return a JSON body. Most carry a single detail message; validation failures return detail as an array of per-field errors, and billing or service failures add a machine-readable error code alongside it.
{ "detail": "This key lacks the 'tts' scope" }
{ "error": "insufficient_credits",
"detail": "Not enough credits for this request",
"required_credits": 0.42,
"balance_credits": 0.10,
"shortfall_credits": 0.32 }| Status | Meaning |
|---|---|
| 400 | Bad request — an unknown model id, or a voice that cannot speak the requested language. |
| 401 | Missing, invalid, or revoked credential. |
| 402 | Insufficient credits, or your account’s daily/monthly spend limit has been reached. The body names the shortfall or the window that tripped. |
| 403 | Authenticated, but the key lacks the scope the endpoint requires. |
| 404 | No such path. |
| 413 | The request body is too large. |
| 422 | The request failed validation — missing fields, out-of-range values, a body that is not JSON, or an audio URL that is not publicly fetchable. |
| 429 | The service is busy. Retry shortly. |
| 500 | Internal server error. |
| 502 | The upstream request failed. Retry shortly. |
| 503 | The service is temporarily unavailable. |
Chat completions
/api/v1/chat/completionsAn OpenAI-shaped chat completion. Point an OpenAI-compatible client at the Bevia base URL with your bv_ key and the request format carries over. Requires the llm scope.
| Field | Type | Description |
|---|---|---|
| model | string | Chat model id: bevia-chat (default) or bevia-chat-pro. |
| messages | array | Required, 1–200 items. Each is { role, content } with role one of system, user, assistant, tool. |
| stream | boolean | Default false. When true, the response is a server-sent event stream of chat.completion.chunk objects ending in data: [DONE]. |
| temperature | number | Sampling temperature, 0–2. |
| top_p | number | Nucleus sampling, 0–1. |
| max_tokens | integer | Completion cap, 1–32000. |
| stop | array | Up to 4 stop sequences. |
| seed | integer | Seed for reproducible sampling. |
| tools | array | OpenAI-shaped tool definitions, up to 64. |
| tool_choice | any | Tool selection control, as in the OpenAI shape. |
| response_format | object | Response format constraint, as in the OpenAI shape. |
| reasoning_effort | string | Reasoning effort hint. Accepted values depend on the model. |
curl https://api.beviaai.com/api/v1/chat/completions \
-H "Authorization: Bearer bv_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "bevia-chat",
"messages": [
{ "role": "system", "content": "You answer briefly." },
{ "role": "user", "content": "Say hello." }
]
}'{
"id": "chatcmpl-…",
"object": "chat.completion",
"created": 1760000000,
"model": "bevia-chat",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello!" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 6,
"total_tokens": 30
}
}With stream: true, the response is a text/event-stream with Cache-Control: no-store and an X-Request-Id header. Each event is a chunk shaped like the response above with a delta object in place of message, and the stream ends with data: [DONE].
Billed per token — prompt plus completion tokens, as reported in usage. Per-model prices come from GET /api/v1/models.
Text to speech
/api/v1/ttsSynthesize speech from text. The response is the audio itself — audio/wav binary, not JSON. Requires the tts scope.
| Field | Type | Description |
|---|---|---|
| text | string | Required. The text to speak, 1–5000 characters. |
| voice | string | A voice id from GET /api/v1/voices. The service default is used when omitted. |
| language | string | ar (default) or en. |
| model | string | TTS model id. Only bevia-tts today. |
| format | string | Only wav is supported — 16 kHz mono PCM. |
curl https://api.beviaai.com/api/v1/tts \
-H "Authorization: Bearer bv_your_key" \
-H "Content-Type: application/json" \
-d '{ "text": "Welcome to Bevia.", "language": "en" }' \
-o speech.wavThe response headers report what the request drew:
| Header | Description |
|---|---|
| X-Request-Id | Identifier of this request — quote it when reporting a problem. |
| X-Bevia-Credits-Charged | Credits drawn for the request. |
Billed per character synthesized, at the voice’s own rate (each voice lists it under pricing). A voice that does not exist or cannot speak the requested language fails with 400 — regional voices speak Arabic only.
Speech to text
/api/v1/sttTranscribe an audio clip — post the raw bytes, or a URL to a publicly reachable file. Requires the stt scope.
| Query param | Type | Description |
|---|---|---|
| language | string | ar (default) or en. |
| model | string | bevia-stt (default, Arabic and English) or bevia-stt-saudi — Saudi-dialect Arabic (Najdi, Hijazi, Gulf), language=ar only. It is the model Bevia voice agents listen with on Arabic calls. |
Two body shapes are accepted:
- Raw audio — the request body is the audio file itself, with its own
Content-Type(audio/wav,audio/webm,audio/ogg, and so on). - JSON —
{ "url": "https://…" }pointing at a publicly reachable audio file over HTTP or HTTPS. URLs that resolve to private or otherwise unreachable addresses are refused.
curl "https://api.beviaai.com/api/v1/stt?language=en" \
-H "Authorization: Bearer bv_your_key" \
-H "Content-Type: audio/wav" \
--data-binary @clip.wavcurl "https://api.beviaai.com/api/v1/stt?language=en" \
-H "Authorization: Bearer bv_your_key" \
-H "Content-Type: application/json" \
-d '{ "url": "https://example.com/clip.mp3" }'{
"request_id": "…",
"transcript": "Hello, how can I help you?",
"language": "en",
"model": "bevia-stt",
"duration_seconds": 2.4,
"confidence": 0.97,
"credits_charged": 0.0012
}Billed per second of audio — on the duration the service measures, reported back as duration_seconds, with the draw shown in credits_charged.
Voice sessions
/api/v1/voice/sessionsCreate a voice-agent session: a room plus the short-lived credentials one participant needs to join it over Bevia’s realtime voice transport. Returns 201. Requires the voice scope.
The flow, end to end:
- Configure the agent in the dashboard under Agents — its language, voice, greeting, and instructions.
- POST
/api/v1/voice/sessionsfor each participant who should talk to it. The response is the connection set: serverurl,room,identity, a short-livedtoken, and when it expires. - Connect the client to
url/roomwith the token. The configured agent joins the room and the conversation begins.
The request body is optional — omit both fields for a fresh session, which is the normal case:
| Field | Type | Description |
|---|---|---|
| room | string | Room suffix, 1–64 characters of A–Z a–z 0–9 _ -. Supply it only to add a second participant to a room you already hold; room names are scoped to your account. |
| identity | string | Participant identity, same charset. A random one is generated when omitted. |
curl -X POST https://api.beviaai.com/api/v1/voice/sessions \
-H "Authorization: Bearer bv_your_key" \
-H "Content-Type: application/json" \
-d '{}'{
"url": "wss://…",
"room": "bv-…-…",
"identity": "u-…",
"token": "…",
"expires_at": "2026-10-01T12:00:00Z"
}The token is a join credential for exactly that room — it expires at expires_at, after which you mint a fresh session. A 503 means voice is temporarily unavailable on your account. Session usage is billed as the speech and model work inside the call; the per-session breakdown is visible in the dashboard.
Send messages from the WhatsApp Business numbers you connect in the dashboard under WhatsApp. Connecting a number happens in the dashboard only. Reads (GET) need the read scope; sending, templates and opt-outs need the whatsapp scope. Until WhatsApp is set up for your account, these endpoints answer 503.
/api/v1/whatsapp/messagesSend one message. Pass an Idempotency-Key header (a fresh UUID per message): a retry with the same key returns the original message with 200 instead of sending twice. A new message returns 201.
| Field | Type | Description |
|---|---|---|
| number_id | string | The id of one of your connected numbers (GET /api/v1/whatsapp/numbers). |
| to | string | Recipient in international format, e.g. +9665XXXXXXXX. |
| type | string | text, template, image, document, audio or video. |
| text | object | { body, preview_url? } — for type text. |
| template | object | { name, language, components? } — an approved template. Body variables go in components: [{ type: "body", parameters: [{ type: "text", text }] }]. |
| media | object | { url, caption?, filename? } — a public https URL, for media types. |
curl -X POST https://api.beviaai.com/api/v1/whatsapp/messages \
-H "Authorization: Bearer bv_your_key" \
-H "Idempotency-Key: 6f1c0d7e-…" \
-H "Content-Type: application/json" \
-d '{"number_id": "…", "to": "+9665XXXXXXXX", "type": "template",
"template": {"name": "order_update", "language": "ar",
"components": [{"type": "body", "parameters": [{"type": "text", "text": "1042"}]}]}}'{
"id": "…",
"number_id": "…",
"direction": "outbound",
"contact": "+9665XXXXXXXX",
"type": "template",
"template_name": "order_update",
"status": "queued",
"billable": null,
"billed_credits": null,
"created_at": "2026-10-01T12:00:00Z"
}status moves through queued, sent, delivered and read, or ends at failed with error_code and error_title. Read it back with GET /api/v1/whatsapp/messages/{id}.
The 24-hour window. Free-form messages (text and media) can only be sent within 24 hours of the customer’s last message to you. Outside that window, send an approved template; anything else fails with 409.
Opt-outs. A customer who replies STOP or إلغاء is added to your opt-out list, and every later send to them fails with 409. You can manage the list yourself too.
Billing. Template messages are billed per delivered message, at the rate for the recipient’s country and the template’s category (GET /api/v1/whatsapp/pricing, in credits). A country without a rate fails with 400 before anything is sent, and an uncovered balance with 402.
| Endpoint | Description |
|---|---|
GET /api/v1/whatsapp/numbers | Your connected numbers, with quality rating and sending limit. |
GET /api/v1/whatsapp/messages | Messages, newest first. Filter with number_id, contact, direction, limit and offset. |
GET /api/v1/whatsapp/templates | Your templates. Filter with account_id and status (e.g. APPROVED). |
POST /api/v1/whatsapp/templates | Create a template and submit it for approval: { name, language, category, components }. A template that breaks the template rules fails with 422 and a review of errors, warnings and tips. |
POST /api/v1/whatsapp/templates/review | Check a template before submitting it. Returns errors, warnings, tips and a score. |
POST /api/v1/whatsapp/templates/sync | Pull template statuses from your WhatsApp account. |
DELETE /api/v1/whatsapp/templates/{id} | Delete a template. |
GET · POST /api/v1/whatsapp/opt-outs | List the opt-out list, or add { contact }. |
DELETE /api/v1/whatsapp/opt-outs/{id} | Remove a number from the opt-out list. |
GET /api/v1/whatsapp/pricing | Per-country, per-category rates in credits per delivered message. |
Incoming customer messages can be forwarded to your own systems as the signed whatsapp.message.received webhook event — subscribe to it under Webhooks.
Inbox
Every conversation across your channels in one place: WhatsApp threads, with your calls to the same number shown in them. Reads need the read scope; replying sends a WhatsApp message, so it needs the whatsapp scope; every other change needs the inbox scope.
/api/v1/inbox/conversations/{id}/replyReply on the conversation's channel: { "type": "text", "text" } inside the 24-hour window, or { "type": "template", "template": { name, language, components? } } with an approved template. Every WhatsApp rule applies (window, opt-outs, billing), and an Idempotency-Key header works as for sends.
Live updates. The list returns a cursor. Poll GET /api/v1/inbox/changes?since=cursor every few seconds: it returns the conversations changed since then and the next cursor. A conversation can come back more than once, so merge by id; reset: true means reload the list, then keep polling from the returned cursor.
AI replies. An owner or admin links an agent to a WhatsApp number and turns on auto-reply in the dashboard: the agent's published version answers customer messages, using knowledge search and only the read-only tools allowed for that number. It stands back for 30 minutes after one of your team replies from the inbox, and a customer who asks for a person pauses it on that conversation (two hours by default) and triggers the conversation.handoff_requested webhook event. AI replies are billed like any model use.
| Endpoint | Description |
|---|---|
GET /api/v1/inbox/conversations | Conversations, most recent first. Filter with channel, status, assignee (me, none or a user id), label, unread, q, number_id, limit and offset. |
GET /api/v1/inbox/conversations/{id} | One conversation with the matching contact and, for WhatsApp, whether the 24-hour window is open. |
GET /api/v1/inbox/conversations/{id}/items | The thread, newest first: messages, calls, notes and events. Page with before=next_before. |
POST /api/v1/inbox/conversations/{id}/notes | Add an internal note: { body }. Never sent to the customer. |
POST /api/v1/inbox/conversations/{id}/assign | Assign to a team member, or unassign: { user_id | null }. |
POST /api/v1/inbox/conversations/{id}/status | open, pending, resolved, or snoozed with snoozed_until (up to 90 days). |
PUT /api/v1/inbox/conversations/{id}/labels | Replace the labels: { labels } (up to 20). |
POST /api/v1/inbox/conversations/{id}/read | Mark as read. |
PATCH /api/v1/inbox/conversations/{id}/ai | { enabled?, resume? }: turn AI replies on or off here, or resume them after a handoff. |
Models
/api/v1/modelsThe public model catalogue — every model id the API accepts, with its service and its price in credits. Requires the read scope.
curl https://api.beviaai.com/api/v1/models \
-H "Authorization: Bearer bv_your_key"{
"object": "list",
"data": [
{
"id": "bevia-chat",
"object": "model",
"owned_by": "bevia",
"service": "llm",
"name": "Bevia Chat",
"description": "Arabic and English conversation model, with tool calling and streaming.",
"languages": ["ar", "en"],
"pricing": {
"currency": "SAR",
"input_credits_per_million_tokens": 0.0,
"output_credits_per_million_tokens": 0.0
}
}
]
}The pricing shape follows the service: chat models report input_credits_per_million_tokens and output_credits_per_million_tokens, TTS reports credits_per_1000_characters, and STT reports credits_per_minute. The public ids are bevia-chat, bevia-chat-pro, bevia-tts, bevia-stt, and bevia-stt-saudi.
Voices
/api/v1/voicesNative Arabic voices from every Arabic country — Saudi Gulf and Standard Arabic voices, plus a female and a male voice for each of the UAE, Kuwait, Qatar, Bahrain, Oman, Yemen, Iraq, Jordan, Lebanon, Syria, Egypt, Libya, Tunisia, Algeria and Morocco. Filter with language (ar or en) and country (a two-letter code such as EG). Requires the read scope.
curl "https://api.beviaai.com/api/v1/voices?language=ar" \
-H "Authorization: Bearer bv_your_key"{
"object": "list",
"data": [
{
"id": "ar-EG-salma",
"name": "Salma",
"gender": "female",
"country": "EG",
"accent": "egyptian",
"languages": ["ar"],
"engines": ["bevia"],
"pricing": { "currency": "SAR", "credits_per_1000_characters": 0.05625 }
}
]
}Pass a voice’s id as the voice field on /api/v1/tts or as an agent’s voice. languages lists what the voice can speak (the Saudi-locale voices speak Arabic and English; regional voices speak Arabic), and engines which voice engines can play it.
Billing
Usage draws from your account’s credit balance — 1 credit = 1 Saudi riyal. The unit depends on the service:
| Service | Billed per |
|---|---|
| Chat completions | Token — prompt plus completion, as reported in the response usage block. |
| Text to speech | Character synthesized — the request also returns the draw in the X-Bevia-Credits-Charged header. |
| Speech to text | Second of audio, on the measured duration — returned as credits_charged. |
| Voice sessions | The speech and model work inside the call; creating the session itself only mints credentials. |
| Delivered template message — at the per-country, per-category rate from /api/v1/whatsapp/pricing. |
A request the balance cannot cover fails with 402 before any work is done. You can also set daily and monthly spend caps in the dashboard — a request that would cross a cap fails with 402 naming the window. Current balance, per-key usage, and the full ledger are on the dashboard’s Usage and Credits pages.