Reference
Chat Completions
The one endpoint that answers questions. OpenAI Chat Completions shape-compatible, with an explicitly enforced supported-parameter subset: it returns a complete body, takes its behaviour from the agent rather than the request, and refuses parameters it does not implement instead of ignoring them.
Request
POST https://api.wicarax.app/v1/chat/completions
Authorization: Bearer wcx_sk_...
Content-Type: application/json
{
"model": "wicarax-zenith",
"messages": [{ "role": "user", "content": "What is your refund policy?" }]
}Every field is either accepted or refused. Nothing is silently ignored, so a parameter you send and the API does not implement comes back as a 400 naming the field rather than quietly changing nothing.
| Field | Type | Required | Notes |
|---|---|---|---|
messages | array | Yes | Must contain at least one user message. Ordered oldest first |
messages[].role | string | Yes | user or assistant only. system, developer, tool and function are refused with 400 unsupported_parameter. Persona and instructions are configured on the agent, not sent per request |
messages[].content | string | Yes | Empty content is refused |
model | string | Yes | Must be exactly wicarax-zenith. Missing, empty, non-string or any other value is refused with 400 model_not_found |
stream | boolean | No | Only false or omitted, and it must be a real boolean. true, "true", 1 and any stream_options are refused with 400 streaming_not_supported |
session_id | string | No | Wicarax extension, not part of the OpenAI schema. Groups turns in the owner's chat log. Max 64 characters. Never put a user id or token in it |
Chat requires a wcx_sk_ secret key. A publishable key is for the non-chat public surface: GET /v1/agent/config and GET /v1/models. See API Keys.
Everything else is refused. Sampling parameters (temperature, top_p, max_tokens, stop, penalties), tool calling (tools, functions, tool_choice), response_format, n, logprobs, seed, user, metadata and store all return 400 unsupported_parameter with param naming the field. A field we do not recognise, such as a typo like mesages, returns 400 unknown_parameter.
How your messages are read
You send a transcript; the server derives the question and the context from it. Knowing the rule avoids a class of confusing results.
- The last user message is the question being answered.
- Everything before it, limited to user and assistant turns, is treated as conversation history.
- A request with no user message at all is rejected with 400 and the code invalid_request_error.
- Do not put your own error text in an assistant turn. It will be read back as something the assistant said.
{
"messages": [
{ "role": "user", "content": "Do you offer annual plans?" },
{ "role": "assistant", "content": "Yes, billed once a year." },
{ "role": "user", "content": "What is the refund window for those?" }
]
}Grouping a conversation
Use session_idto group exchanges in the owner's chat log. The x-wicarax-session header can carry the same value; when both are present, the header takes precedence. Keep the value to 64 characters or fewer.
A session id is a grouping hint, not authentication or authorization. Do not put an email address, access token, or other sensitive identifier in it. Send conversation history explicitly in messages for context on a later call.
Response
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1769500000,
"model": "wicarax-zenith",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "..." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1254,
"completion_tokens": 20,
"total_tokens": 1274
}
}| Field | Notes |
|---|---|
id | Unique per completion. Useful in your own logs |
model | The public model id that produced the answer. Read it here rather than hardcoding one |
choices | Always exactly one entry. The answer is choices[0].message.content |
finish_reason | Always stop, since responses are returned complete |
usage | Token counts for this exchange. Zero when the answer is an unavailability message |
There is no citation, source or provenance field, by design. Which document an answer came from is internal and is never returned on any surface.
Behaviour at the edges
Three cases return 200 even though they are not a normal answer. Handling them as errors would be wrong.
| Case | What you get |
|---|---|
| The material does not cover the question | A reply saying the information is not available, rather than an invented answer |
| The account is out of monthly requests | A short unavailability message, with zero usage |
| The subscription needs attention, for example a failed renewal | The same unavailability message, with zero usage |
Because those are 200 responses, an integration that only checks the status code will look healthy while every visitor reads an unavailability message. Watch Usage and Analytics(sign in required), or alert on the reply text you know that state produces.
Failures
| Status | Code | Cause |
|---|---|---|
400 | invalid_request_error | The body is not valid for this API: a malformed field, a size limit, or no user message |
400 | streaming_not_supported | A truthy stream, or any stream_options |
400 | model_not_found | model absent, empty, non-string, or not wicarax-zenith |
400 | unsupported_parameter | A recognised OpenAI field this API does not implement. param names it |
400 | unknown_parameter | An unrecognised field, including a misspelling |
401 | invalid_api_key | No Authorization header |
403 | access_denied | The credential or origin was refused. One uniform answer for every such cause, including a wcx_pk_ publishable key: chat requires a secret key |
413 | The body exceeded about 100 KB | |
429 | tenant_plan_rate_limit_exceeded | Your account's current per-minute plan limit. Carries Retry-After |
429 | auth_probe_rate_limit_exceeded | Too many failed credential attempts from your network. Carries Retry-After |
5xx | server_error | A failure on our side. Safe to retry with backoff |
Full detail, including the reasoning behind the uniform 403, is on Errors. Limits and their values per plan are on Rate Limits.
Send only the turns that matter. A long transcript is charged against the request body limit and adds nothing once the relevant context is already present.
Three cases return 200 without being a normal answer: a question your material does not cover, a spent monthly quota, and a subscription that needs attention. Status codes alone will not distinguish them, so read the reply before deciding an integration is healthy.