Outcomes
Judge finished conversations against your own success criteria and track a resolution rate.
An outcome is a verdict on a finished conversation: success, failure, or unknown, usually with a one-sentence rationale — the model can omit the rationale, and the verdict is still stored and shown without one. You describe what a good conversation looks like, and eligible conversations are judged against that description after they end. The results show up in analytics and in Chat Logs.
Evaluation runs as part of periodic conversation analysis, so results are not instant: a conversation is picked up once it has been quiet for about a day, and only conversations where the visitor sent at least two messages are analyzed — a single-question exchange is never judged.
Setting Success Criteria
Open your agent's Settings → Outcomes page and write your criteria in plain language, up to 1,000 characters. For example:
The visitor's question was answered or their issue was resolved without needing a human.
To skip evaluation entirely, leave the field empty and define no extraction fields — either one keeps evaluation active. Criteria are applied when a conversation is analyzed, not when you save: conversations that have not been analyzed yet (including older ones, or everything once you upgrade) are judged against whatever criteria are in place at analysis time, and already-analyzed conversations are only revisited if they receive new messages.
How Evaluation Works
When a conversation is analyzed, the transcript and your criteria are passed to a model. Very long conversations are truncated before evaluation — each message is capped at 2,000 characters and the transcript as a whole at 20,000, keeping the beginning — so in an unusually long conversation a resolution that happens at the very end may fall outside what the model sees. The model returns:
- outcome -
success,failure, orunknown.unknownis used when the transcript gives no reliable signal either way. - rationale - one sentence, 300 characters or less. The model is instructed to avoid names, email addresses, phone numbers, order numbers, and other identifying values, but this is best-effort — the returned text is stored as-is, without an enforced redaction pass — so treat rationales as potentially containing conversation details.
Collecting Data From Conversations
On the same settings page you can define up to 10 extraction fields — for example an email address or an order number. Each field has:
| Property | Notes |
|---|---|
| Key | Starts with a lowercase letter, then lowercase letters, digits, and underscores; up to 40 characters; must be unique |
| Type | string, number, or boolean |
| Description | Up to 200 characters, passed to the model as instructions |
Use string for identifiers such as order, ticket, or customer numbers even when they look numeric: a number field goes through JavaScript number coercion, which drops leading zeros and can silently round identifiers longer than about 15 digits. Reserve number for genuine quantities.
Write the description as a recognition rule, not a wish. The model is instructed to record a field only when the visitor actually provided the value and never to guess or infer one, and empty values are dropped rather than stored. That is an instruction, not a verification step — extracted values are not cross-checked against the transcript, so treat them as unverified conversation-derived data rather than ground truth, and validate anything you act on downstream.
Plan Availability
Outcomes are available on every plan, including Free. Criteria, fields, evaluation and the resolution rate need no upgrade.
Reading the Resolution Rate
The Resolution rate card in Analytics shows:
- The rate - successes divided by the conversations that got a decided verdict, as a percentage. In other words, success ÷ (success + failure).
unknownverdicts are excluded from the rate. Counting them would drag the number down for conversations the model could not judge either way. They are still counted in the evaluated total shown under the rate, next to the decided count.- A "View failures" link that opens Chat Logs filtered to the failing conversations, so you can read what went wrong.
If nothing has been evaluated yet, the card is replaced by a prompt to configure outcomes.
When you narrow analytics to a period, a conversation belongs to that period by its start time, not by when it was evaluated — a conversation evaluated today updates the numbers for the period in which it started.
Practical Tips
- Describe the visitor's win, not the agent's behavior - "their issue was resolved" beats "the agent was polite"
- Name the failure case too - say what counts as a miss so borderline conversations don't all land on
unknown - Start with no extraction fields - add them once the verdicts look right
- Read the failures weekly - the rationale plus the transcript usually points at a missing source