Reference
Rate Limits
Every agent on an account shares the same request throughput and monthly quota. Current plans limit requests per minute; they do not add a separate daily throughput cap.
Three different boundaries
| Boundary | Scope | When reached |
|---|---|---|
| Per-minute throughput | Account, across agents and channels | 429 tenant_plan_rate_limit_exceeded with Retry-After |
| Monthly request quota | Account, in monthly quota windows | A complete unavailability answer with HTTP 200 and zero usage |
| Failed-credential budget | Network block | 429 auth_probe_rate_limit_exceeded after repeated refusals |
These boundaries serve different purposes. A successful request may spend monthly quota and occupy the per-minute window. A refused credential does not spend request quota, but repeated refusals may exhaust the separate failure budget.
Current standard plans
| Plan | Requests/minute | Requests/month | Agents |
|---|---|---|---|
| Starter | 30 | 200 | 1 |
| Lite | 60 | 1,300 | 2 |
| Pro | 150 | 3,700 | 5 |
| Business | 300 | 9,000 | 20 |
| Enterprise | By contract | By contract | By contract |
The limit belongs to the account, not an individual agent or key. Two agents on Pro share 150 requests per minute and 3,700 requests in a monthly quota window. Adding agents separates audiences and knowledge; it does not increase throughput.
Enterprise catalogue values have no fixed public ceiling. Capacity for an Enterprise account is defined by its agreement and effective entitlement. Check the dashboard for the limits currently applied to your account.
Monthly windows and yearly billing
A yearly subscription is billed yearly, but its request quota still runs in monthly windows. The annual payment does not turn twelve monthly allowances into one annual request bucket.
Completed answers through the widget, SDK, REST API, and Playground use the same account quota. Model and appearance reads do not consume an answered-request quota. See Plans and Usage for billing and consumption details.
What reaches each boundary
| Traffic | Per-minute plan limit | Monthly answer quota |
|---|---|---|
| Chat answer through REST or SDK | Yes | Yes, when an answer is completed |
| Hosted widget message | Yes | Yes, when an answer is completed |
| Playground answer | Yes | Yes, when an answer is completed |
| Model or appearance read | Yes, plus a separate per-address safeguard | No |
| Widget loader script | No | No |
| Refused credential or origin | Failure budget may apply | No |
A failed or refused call does not consume an answered-request quota. A successful metadata read still occupies throughput, so fetch model and appearance data when needed rather than polling on every message.
Responding to a 429
Read the machine-readable error code and the Retry-After header. A plan limit calls for pacing requests. An authentication probe limit calls for correcting the credential or origin before trying again.
HTTP/1.1 429 Too Many Requests
Retry-After: 25- Wait at least the indicated number of seconds before retrying a valid request.
- Queue batch work so it does not arrive as one burst.
- Do not retry a 403 in a loop; verify the key class, active state, agent, and installation origin first.
Do not treat a polite HTTP 200 unavailability answer as a rate-limit success. Check the response content and usage, then review quota and billing state in Usage and Analytics(sign in required).
Plan for shared traffic
Size the integration for all channels together. If your account serves a public widget and a batch job, the batch job can occupy the same per-minute allowance the widget needs. Pace background work and leave capacity for visitor traffic.
Use separate agents when their knowledge or audience should differ, not as a way to multiply limits. If normal traffic consistently reaches the account ceiling, review the current effective plan and usage in the dashboard before changing the integration.
For response codes and refusal causes, continue with Errors. For plan pricing and billing behavior, see Plans and Usage.