Reference

Rate Limits

Every agent on an account shares the same request throughput and monthly quota. Current plans limit requests per minute; they do not add a separate daily throughput cap.

Three different boundaries

BoundaryScopeWhen reached
Per-minute throughputAccount, across agents and channels429 tenant_plan_rate_limit_exceeded with Retry-After
Monthly request quotaAccount, in monthly quota windowsA complete unavailability answer with HTTP 200 and zero usage
Failed-credential budgetNetwork block429 auth_probe_rate_limit_exceeded after repeated refusals

These boundaries serve different purposes. A successful request may spend monthly quota and occupy the per-minute window. A refused credential does not spend request quota, but repeated refusals may exhaust the separate failure budget.

Current standard plans

PlanRequests/minuteRequests/monthAgents
Starter302001
Lite601,3002
Pro1503,7005
Business3009,00020
EnterpriseBy contractBy contractBy contract

The limit belongs to the account, not an individual agent or key. Two agents on Pro share 150 requests per minute and 3,700 requests in a monthly quota window. Adding agents separates audiences and knowledge; it does not increase throughput.

Note

Enterprise catalogue values have no fixed public ceiling. Capacity for an Enterprise account is defined by its agreement and effective entitlement. Check the dashboard for the limits currently applied to your account.

Monthly windows and yearly billing

A yearly subscription is billed yearly, but its request quota still runs in monthly windows. The annual payment does not turn twelve monthly allowances into one annual request bucket.

Completed answers through the widget, SDK, REST API, and Playground use the same account quota. Model and appearance reads do not consume an answered-request quota. See Plans and Usage for billing and consumption details.

What reaches each boundary

TrafficPer-minute plan limitMonthly answer quota
Chat answer through REST or SDKYesYes, when an answer is completed
Hosted widget messageYesYes, when an answer is completed
Playground answerYesYes, when an answer is completed
Model or appearance readYes, plus a separate per-address safeguardNo
Widget loader scriptNoNo
Refused credential or originFailure budget may applyNo

A failed or refused call does not consume an answered-request quota. A successful metadata read still occupies throughput, so fetch model and appearance data when needed rather than polling on every message.

Responding to a 429

Read the machine-readable error code and the Retry-After header. A plan limit calls for pacing requests. An authentication probe limit calls for correcting the credential or origin before trying again.

response headers
HTTP/1.1 429 Too Many Requests
Retry-After: 25
  • Wait at least the indicated number of seconds before retrying a valid request.
  • Queue batch work so it does not arrive as one burst.
  • Do not retry a 403 in a loop; verify the key class, active state, agent, and installation origin first.
Heads up

Do not treat a polite HTTP 200 unavailability answer as a rate-limit success. Check the response content and usage, then review quota and billing state in Usage and Analytics(sign in required).

Plan for shared traffic

Size the integration for all channels together. If your account serves a public widget and a batch job, the batch job can occupy the same per-minute allowance the widget needs. Pace background work and leave capacity for visitor traffic.

Use separate agents when their knowledge or audience should differ, not as a way to multiply limits. If normal traffic consistently reaches the account ceiling, review the current effective plan and usage in the dashboard before changing the integration.

For response codes and refusal causes, continue with Errors. For plan pricing and billing behavior, see Plans and Usage.