Guardrails
Run safety checks before and after your agent answers: topic focus, prompt-injection detection, and content moderation.
Guardrails are safety checks that run before and after your agent responds. Turn on the checks you want, and the agent blocks requests that fall outside them instead of answering.
One honest caveat up front: guardrails are best-effort, not an absolute guarantee. The topic, injection-classifier, and content checks call an external provider, and if a check times out or errors the agent answers anyway rather than going down with the provider (it fails open). Only the built-in prompt-injection pattern match works entirely locally. Treat guardrails as a strong filter, not a hard security boundary.
Turning Guardrails On
- Go to your agent dashboard
- Click Settings
- Select Security
- Open the Guardrails card and enable the checks you want
- Save
Guardrails are configured per agent. Each agent has its own settings — there is no workspace-wide default, so enable them separately on every agent you want protected.
These settings apply to chat and generated email replies. The sections below describe chat behavior; see Email replies for the email differences.
The Three Checks
Topic focus
Blocks requests that are unrelated to this agent's public purpose — the check judges each request against your agent's site name, not against your sources. Use it when your agent should stay on your product, service, or documentation instead of answering general questions.
Prompt-injection detection
Blocks attempts to replace your instructions, reveal hidden prompts, or coerce the agent into using its tools.
Content moderation
Checks visitor input and agent output for unsafe content. This is the only check that also looks at what the agent said. On output it checks the agent's main answer text only: auto-generated suggested-message chips are not separately checked — though no chips are produced at all when the main answer was flagged.
What Happens on Input
The enabled checks run before the model writes an answer.
- A fast pattern match for well-known prompt-injection phrases can block a message immediately, before the other checks run.
- The remaining enabled checks may run together rather than one after another.
- If any enabled check triggers, the turn is blocked and the visitor never sees the model's answer.
The checks read a bounded excerpt, not the entire payload: only the first 8,000 characters of the visitor's message are inspected (recent history and attachment text are capped separately), while the full message still reaches the model. An unusually long message can therefore carry uninspected content past the input checks.
When more than one check triggers on the same message, the recorded reason follows a fixed reporting precedence: content moderation, then the prompt-injection classifier, then the topic classifier.
What the Visitor Sees When Blocked
If you filled in the Blocked-message override, the visitor gets that message. Leave it blank and the agent uses a built-in default instead:
- Off-topic requests get "I can only help with questions about your site name. What would you like to know?"
- Every other block gets "Sorry, I can't help with that. Is there anything else I can help you with?"
The override is optional and limited to 500 characters.
Two things to know about blocked turns:
- They still count against your message quota. The check runs on a real request, so the turn is charged like any other.
- They are saved. Blocked turns appear in Chat Logs with the guardrail reason attached, so you can review what was caught.
What Happens on Output
Content moderation also runs once after the model responds. Topic focus and prompt-injection detection do not run on output.
An output flag does not retract the answer — the visitor has already seen it, and it stays in the conversation. What the flag does instead:
- Records the flag on that turn so you can review it in Chat Logs
- Skips the suggested messages that turn would otherwise have offered
There is no output blocking: answers stream to the visitor before the check finishes, so the output check is a review signal, not a filter. Tightening the agent's sources and instructions makes unsafe output less likely, but cannot guarantee it is never shown.
Email replies
Email checks run before drafting and after generation. An input block leaves the message for human review without generating a reply or using message quota. An output flag holds the draft in Activity > Emails > To review and prevents auto-send; generating that draft still uses quota.
These checks also apply when you choose Draft a reply. Sending a draft yourself does not run the checks again. Review flagged drafts before sending them. See the Email guide for setup and inbox controls.
False Positives
Safety checks can occasionally block a legitimate request. Review flagged conversations in Chat Logs before enabling more broadly. Filtering Chat Logs to guardrail-flagged conversations is the fastest way to see whether a check is too aggressive for your audience.
If a check blocks too much:
- Turn it off. Topic focus judges each request only against your agent's public purpose (its site name) — it does not read your sources or instructions, so adding content cannot widen what it accepts
- Set a friendlier Blocked-message override so the dead end still reads well