Building
Playground
The Playground runs the same assistant your visitors will meet, with the same material, the same limits and the same behaviour. It is the last step before launch and the first step after changing content.
What it actually runs
Nothing is simulated. The Playground calls your real agent with your real material, so an answer you see there is the answer a visitor gets.
- Conversations are real requests. They consume your monthly allowance exactly as traffic from a website does, and they appear in your usage reporting under the playground channel.
- An inactive agent is refused here as well, with a message saying the agent is inactive. Activate it first.
- When the monthly allowance is spent, or the subscription needs attention, the Playground returns an explicit error rather than a polite reply. Visitors on your website see a short unavailability message instead, so the two surfaces read differently on purpose.
- An agent with no material will decline nearly everything, which is correct behaviour rather than a fault.
- Answers never show where the material came from, because that is never disclosed on any surface.
Open it on the Playground(sign in required) page and pick the agent you want to test.
What to test before you launch
Most teams test the questions they hope to be asked. The useful test is the opposite.
- 1Ask the five questions your support inbox receives most often. These are the ones that must be right.
- 2Ask something your material genuinely does not cover, and confirm the assistant declines instead of inventing an answer. This is the single most important check.
- 3Ask the same question in different words, including the wording a customer would use rather than your internal terminology.
- 4Ask in every language you expect visitors to use. Answers come back in the language of the question.
- 5Ask a question whose answer changed recently, to confirm the outdated version was removed rather than left in place alongside the new one.
If two sources contradict each other, the assistant may answer with either. An answer that appears inconsistent between attempts usually has that cause, and the correction belongs in the material rather than in any setting.
Reading the result
| What you see | What it usually means | What to do |
|---|---|---|
| A correct, specific answer | The material covers the question and was found | Nothing. Move to the next question |
| A decline | Nothing in the material covers it | Add a question and answer pair, or the missing document |
| Correct but vague | The material mentions the topic without stating the fact | Add the specific fact as a question and answer pair |
| Outdated | An old source is still in use | Delete or archive the old source, then retest |
| An error saying the request quota was exceeded | The account has used its monthly requests, or the subscription needs attention | Check Plans and Usage |
After launch
Come back to the Playground whenever you change material, and treat it as the place to reproduce anything a visitor reports. The Knowledge Gaps(sign in required) page lists visitor questions your agents could not answer from the available material.
When you are satisfied with the answers, choose how to deliver them: the widget for a website, the SDK for your own interface, or the REST API for everything else.
A repeatable release check
Keep a small test set for each agent: a supported fact, an unsupported question, an ambiguous question, a follow-up that needs conversation context, a recently changed fact, and a tone check. Record the expected outcome, then rerun the set after changing sources or scope.
When a result changes, inspect active sources and knowledge-base scope before changing the agent. Playground uses the same knowledge as production, but it presents quota and account-state failures more directly than a visitor-facing integration.
Keep a short, written list of test questions and rerun it after every content change. Five minutes of the same five questions catches regressions that unstructured testing misses.
Playground messages are real requests. They are answered by the same assistant, metered the same way, and counted against your monthly allowance, so a long testing session on a small plan does consume it.