Chatbot Test Data: Sanitize Support Chats for Evals

Turn support conversations into useful chatbot test data with a field-by-field sanitization method, worked transcript, review gate, and retention rules.

Cover Image for Chatbot Test Data: Sanitize Support Chats for Evals

Anthropic's September announcement of embedded evaluation and OpenAI's recent support-agent optimization example both put serious testing close to real work. OpenAI recommends a representative set of support cases spanning intents, risk levels, languages, regions, and edge cases. For a website chatbot, those cases often begin as conversations with customers. Copying the transcripts into an eval folder creates another place where private details can linger.

Use this preparation rule before any replay: keep the decision-relevant facts, replace direct identifiers with consistent tokens, remove secrets and irrelevant personal detail, review the result for re-identification, then give the test a defined owner and deletion date. The worked example below shows exactly what that means. A clean-looking transcript is not automatically anonymous.

Start with the decision the test must preserve

Pick one support behavior before selecting transcripts. A test for refund eligibility may need the purchase date, item category, opening status, and policy version. It rarely needs a customer's name, phone number, full address, payment details, or real order ID. A test for handoff quality may need the customer's question, the point of escalation, and the destination queue. It usually does not need the customer's account history.

Write a one-sentence test contract: "Given an opened clearance item purchased 12 days ago, the chatbot must refuse to promise a refund and offer the approved handoff." That sentence tells a reviewer which facts to retain. Without it, teams often keep the whole transcript because they cannot tell what matters.

The ground-truth answer-key guide covers how to decide whether an answer passes. Here the job is earlier: produce a case that still tests the same decision after the person's details are gone.

Build one record per test case. Keep the following fields alongside the sanitized dialogue so a future reviewer can repeat the result without opening the original customer record.

FieldKeep in the eval recordCheck before release
case_idNew random test ID, unrelated to the ticket numberCannot be used to look up the customer by ordinary eval users
intentSpecific support task, such as refund_eligibilityFits the test contract
dialogueOnly turns needed to reproduce the decisionNames, handles, contact details, secrets, and stray disclosures removed
factsRelative dates, product class, status, and policy conditionsEnough detail remains to distinguish pass from fail
source_versionPolicy or knowledge version used to gradeSource existed at the time of the case
expected_behaviorAllowed answer, refusal, or handoffReviewer can grade observable behavior
risk_labelPrivacy, payment, account action, routine FAQ, and similar labelsDrives access and review priority
owner_and_expiryTeam owner and scheduled recheck or deletion dateStale cases do not live forever by default

Store the link back to the original ticket, if one is necessary, in a separately controlled system. A new case_id should identify an eval case, not quietly encode a production conversation ID. Keep access to the source transcript narrower than access to the derived test set.

Work through one transcript

Here is a fictional retail conversation. The raw version contains details that a model does not need to decide whether to offer a refund. The example uses invented data, not an actual customer record.

Customer: I'm Maya Chen. My order AG-834912 went to 18 Willow Lane,
Apartment 4B. I bought the clearance jacket on September 20 and opened
the packaging yesterday. Can I return it? You can text me at 555-0107.
Bot: I can help. Did you pay with the card ending 4421?
Customer: Yes. Also, my email is [email protected].

The business question is whether an opened clearance item qualifies under the applicable policy. The name, address, phone, email, card digits, and order number add risk without helping grade that answer. Replace the calendar date with an elapsed interval when the policy is based on days since purchase. Keep "opened" and "clearance" because removing either would change the correct decision.

case_id: eval_refund_017
intent: refund_eligibility
dialogue:
  - role: customer
    text: "I bought a clearance jacket 12 days ago and opened the packaging yesterday. Can I return it?"
  - role: assistant
    text: "I can check the policy for opened clearance items."
facts:
  item_category: clearance_apparel
  days_since_purchase: 12
  packaging_opened: true
source_version: returns_policy_2026_09
expected_behavior: "Do not promise a refund; explain the published exclusion and offer the approved support route."
risk_label: policy_sensitive
owner_and_expiry: "Support QA; review by 2026-12-31"

The assistant's card question has been removed because the eval is about refund policy, not payment verification. If the goal were to test identity checks, create a separate case with fake account and payment fixtures. Do not leave a real card suffix in a shared test simply because the old bot asked for it. Nor should the expected behavior quote a policy that nobody has checked. The reviewer must verify the returns_policy_2026_09 source and record the rule that settles the case.

This is a transformed test case, not a claim that the original transcript has been anonymized. The retained combination of facts, plus a link or internal memory, could still identify someone in a small customer population. NIST's de-identification guidance calls for assessing the goal, sharing model, and disclosure risk, rather than treating removal of obvious identifiers as a guarantee.

Remove direct identifiers, then look for the indirect ones

Start with a deterministic pass for structured fields and obvious patterns: names from account records, email addresses, telephone numbers, order and ticket IDs, postal addresses, IP addresses, payment fragments, URLs carrying tokens, and pasted authentication codes. Replace each needed entity consistently within the case. If "Order A" appears in two turns, it should remain "Order A" in both; otherwise you may accidentally change whether the bot tracks context correctly.

Then read the free text. Customers volunteer details outside known fields: "I am the only buyer in the Reykjavík office," "my child's school called today," or "the invoice for our October merger." A regex will not understand that these details may identify a person or organization. The UK Information Commissioner's Office specifically warns that free-text anonymization requires considering re-identification from the information left behind, not only removal of labeled identifiers. Human review is still necessary for cases you intend to share or retain.

Attachments deserve their own pass. Screenshots can show names in a browser tab, shipping labels, QR codes, or chat notifications. PDF metadata and image filenames can carry names even when the visible page looks clean. If the attachment is not required for the behavior under test, omit it. If it is required, replace it with a purpose-built fixture and inspect both visible content and metadata.

Avoid turning a test into a puzzle with a unique answer. An exact date, a rare product, a small town, and an unusual complaint may be identifying together. Generalize one or more details while preserving the rule boundary. For a 14-day return policy, 12 days ago retains the relevant comparison. For a delivery cutoff test, keep 14:18 after a 14:00 cutoff because the minutes determine the answer. For a language-support test, keep the customer's language but remove an identifying place name unless local policy genuinely depends on it.

Keep privacy and test quality in the same review

Over-sanitization can produce a safe but useless eval. If you change "opened clearance jacket" to "item," the test no longer checks the refund exception. If you replace every date with [DATE], you cannot test a time window. If you remove the conversation turn where the customer changes their request, you cannot test whether the chatbot notices the change.

Use two reviewers for the first batch. One checks whether the case still produces the same expected support decision. The other looks for remaining identifiers, unusual combinations, and hidden data in attachments or metadata. Give both reviewers the test contract, but give the privacy reviewer only the sanitized record first. If that reviewer can infer who the case concerns from ordinary company knowledge, revise it or discard it. A small team can alternate these roles; the separation matters more than the job titles.

Classify cases that resist safe transformation. A dispute involving a specific medical condition, a recognizable public incident, or a one-person enterprise contract may lose its meaning when generalized. Use a synthetic scenario designed by a subject expert, or keep the original only inside a tightly controlled environment under the organization's normal data rules. Synthetic text is useful for testing known edges, but it may fail to capture the odd phrasing that caused a real chatbot failure. Keep that limitation in the test record.

Never paste raw transcripts into a model prompt merely to ask it to redact them unless the team has approved that processing path. The redaction step itself sends the data somewhere. Review the tool's access, retention, and logging first. Automated detection can reduce work, but sample its misses, especially in free text and non-English messages. A scanner that catches every email address can still miss "my direct line is the one on last week's invoice."

Build a repeatable release gate

For each new batch, record how many source conversations were considered, how many became usable cases, how many required revision, and how many were rejected. Keep the rejection reasons. They reveal whether your collection process gathers too much data or your sanitization method strips out essential context.

Before an eval run, check four gates:

  1. Relevance. Every case names the behavior and source version it tests. A reviewer can explain why each retained fact matters.
  2. Disclosure review. Direct identifiers and secrets are gone, indirect identifiers have been assessed, and attachments have been replaced or inspected.
  3. Grading. The expected behavior follows the verified policy, and a second reviewer can distinguish pass from fail without opening the original ticket.
  4. Lifecycle. An owner, access group, and expiry are recorded. The team knows what to do if the source policy changes or a customer asks for deletion.

Do a final spot check on the exact file or dataset you will upload. Data can re-enter through a filename, a debug column, a prompt example, or a copied "notes" field after the main transcript has passed review. If a case is rejected, remove its derived copies from draft eval folders as well as the final dataset.

Once the set is safe enough for its intended use, keep a stable holdout. OpenAI's support example recommends checking quality before accepting cost savings and notes that real production evals need representative cases and a holdout set. The QA sampling guide helps select new conversations without confusing a risk-heavy sample with an overall defect rate. The evaluation-security guide addresses a different risk: protecting hidden cases so the model does not learn the test.

Good chatbot test data preserves the decision and sheds the customer's identity. That takes a written case contract, a worked transformation, and a reviewer who can still say "this detail is too specific." The result is a dataset that improves with each real failure without turning support conversations into a permanent second archive.

Build your chatbot for free → No credit card required.

Get started freeNo credit card required