In September, Anthropic reported that two model companies had secretly relayed some customer requests to Claude while presenting responses as their own models. Those are Anthropic's findings, not an independent audit of the companies' systems. The allegation illustrates a problem every chatbot owner can examine: a model name in a dashboard does not establish where a customer message went.
Use this model identity evidence sheet for each production route. Record the model your team selected, the service receiving the request, the model named in its response, the party that billed it, and any fallback or reseller in between. If the evidence stops at a vendor-controlled label, call the route unverified and ask the vendor for a way to reconcile it. That is a defensible finding, even if you cannot inspect its servers.
The model identity evidence sheet
Create one row per route, then sample real requests against it. A route means the complete path from your chatbot application to the system that generates a reply. The public FAQ route and the authenticated account-support route may have different providers, retention terms, and fallback rules.
| Evidence | Record for each sampled turn | Question it answers |
|---|---|---|
| Intended route | Chatbot ID, route policy version, selected model ID, timestamp | What did your application mean to use? |
| Outbound request | Destination host, direct provider or intermediary, request ID, approved region | Who first received the message? |
| Returned response | Response ID, reported model and version, provider or gateway name | What does the receiving service say served it? |
| Route changes | Ordered attempts, reason for switch, model that answered | Did fallback or a safety route change the answerer? |
| Independent reconciliation | Direct-provider usage record, invoice line, or vendor audit export linked by request ID and time | Can another system corroborate the claim? |
| Data handling | Approved processors, logging and retention terms for every hop | Was the actual path allowed for this data? |
The last two rows require care. An invoice confirms that someone bought model usage; a monthly total alone cannot prove that a particular customer turn used that model. A provider response field is useful operational evidence, but an intermediary can copy or rewrite it. Ask what generated each record, who can alter it, and how it joins to the turn you are investigating.
For most small teams, the first improvement is simple: keep the outbound destination and provider response metadata beside the conversation event. Then ask the vendor for request-level usage exports or another reconciliation path. You do not need to publish raw prompts to build this ledger.
Separate the labels your team controls from the route it took
Four labels often get collapsed into one setting. The configured model is what an administrator selected. The requested model is what the application sent. The reported model is what the service returned. The underlying processor is the organization and system that actually handled the prompt. Each transition needs evidence.
A chatbot can drift between these labels for ordinary reasons. A product may use a model alias that points to a newer version. A gateway may retry after a timeout. A safety route may replace a refused answer. A regional policy may choose another approved deployment. The model fallback test plan covers how to test those declared switches. This audit asks whether every possible processor was disclosed and whether a customer turn can be traced through it.
Start at your own boundary. Log the exact endpoint your application called, the model ID you put in the request, and a locally generated correlation ID. Preserve the provider's response ID and model metadata without rewriting them into a friendlier name. OpenAI's response reference describes a returned model ID; the Gemini API reference describes modelVersion and responseId. These fields help match a response to a request. They do not, by themselves, independently prove how an upstream intermediary routed it.
If you buy through an aggregator, ask whether its model name describes the manufacturer, a hosted replica, an alias, or a pool of interchangeable backends. Find out whether a route can change after a timeout, a quota event, a moderation decision, or a price update. A model picker that shows one name while the contract permits undisclosed substitutions is a procurement issue before it becomes a debugging issue.
Work through a support turn
Imagine a retailer's chatbot answering a public delivery-policy question. Its configuration says model-a; the team sends the turn to an approved gateway. The gateway returns a normal answer and reports model-a, but the gateway has not supplied a per-request provider usage record.
{
"turn_id": "turn_4821",
"route_policy": "public-faq-v4",
"requested_model": "model-a",
"destination": "approved-gateway.example",
"gateway_request_id": "gw_71c2",
"reported_model": "model-a",
"gateway_response_id": "resp_92b7",
"fallback_reported": false,
"provider_request_id": null,
"provider_usage_match": "not_available",
"identity_assessment": "gateway_reported_only"
}
This record supports a narrow conclusion: the application sent the question to its approved gateway, and the gateway reported that model-a answered. It does not establish the final model provider. The right next step is to request an export linking gw_71c2 to a provider request ID, served model, region, and billing entry. If the gateway cannot provide that link, keep the route classified as gateway-reported rather than treating the model name as verified.
Now imagine the same route receives an account-recovery question that includes an email address and order number. Your data policy approves only one processor for that route. The missing provider link has a different consequence: pause that account-specific route or send the visitor to a human-supported channel until the vendor proves it stays within the approved processor list. The public shipping FAQ can have a different risk decision. Treat identity evidence according to the data and action at stake.
Do not try to identify the model by asking it, "What model are you?" Models can answer incorrectly, repeat their system instructions, or be prompted to impersonate another provider. A writing-style test is weaker still. Two providers may sound similar, and one provider's output changes with the prompt. Use behavioral probes to find unexpected changes, not to certify the serving model.
Ask the vendor for a traceable route contract
A useful vendor answer names the legal entity receiving the request, every subprocessor that may receive prompt content, the allowed model IDs and hosting locations, and all conditions that permit substitution. It says which record identifies the actual serving model and how the customer can audit that record. It also names who handles retention, deletion, incident notice, and support when a route changes.
Ask the vendor to demonstrate the contract with a low-risk test request. Have it show the customer-visible request ID, its internal gateway record, and the corresponding provider usage entry. Run a forced fallback in a test environment and confirm that both attempts appear. Then verify that the same trace works when streaming, because a partial response can conceal a second attempt if the application saves only the final message.
Do not assume direct-provider billing is always available to you. A reseller may own the upstream account, so your own invoice will show only reseller charges. In that case, request a vendor export or independent audit evidence at the granularity your risk requires. A formal audit report may establish the control design and sample its operation, but it rarely proves each of your turns. The contract should say exactly which assurance you receive.
The chatbot data residency review helps trace the locations and subprocessors that this contract must cover. Model identity and location belong together: an undisclosed processor can change both the data recipient and the jurisdiction, even when the answer shown to the visitor looks unchanged.
Test route drift without putting customer data in the probe
Make a small, harmless canary set that uses public product information. Send the same cases through each approved route on a schedule and after a vendor change. Save request IDs, response metadata, model labels, region labels where available, latency, token usage, and the answer. A canary cannot prove the model's identity, but it can expose missing metadata, new aliases, fallback spikes, or a changed endpoint that deserves investigation.
Include at least four conditions: an ordinary FAQ, a prompt that exercises a permitted refusal, a simulated timeout or quota failure, and a multi-turn follow-up. Ask the vendor to trigger the failure conditions in staging if production traffic cannot safely force them. Check whether the event record lists every attempt, not only the model that finished. Watch for a configured model that never appears as the served model and for response IDs that cannot be reconciled after an incident.
The AI model routing guide explains how to choose a model tier based on cost and quality. Its route policy only works if the execution record reflects where traffic actually went. A surprise processor can invalidate the quality test, pricing assumption, or privacy approval for that route.
Set a review trigger when the vendor adds an upstream provider, changes an alias, alters a fallback rule, migrates a region, or loses the request-level audit trail. A normal version update may need only a fresh answer-quality test. An unapproved data recipient needs a routing decision before more account-specific conversations pass through it.
Decide what the evidence can support
Use three labels in the review record. Confirmed at the application boundary means your logs show the request and endpoint. Corroborated by the provider chain means a second record joins that turn to a serving provider and model. Unresolved means the vendor's label is the only evidence for the final hop. State the scope: a sampled month, a particular route, or a particular incident. None of these labels implies that you can inspect a proprietary model's weights or prove every internal computation.
When records disagree, preserve the original request and response metadata, the vendor's explanation, and the affected time range. Stop using the route for sensitive work if the mismatch could send data to an unapproved processor. Review whether an answer-quality regression, retention promise, or customer notice also needs action. Do not quietly overwrite the old label after the vendor fixes its dashboard.
A trustworthy chatbot model setting should survive a request-level question: who received this customer message, who generated the reply, and what evidence connects the two? Keep the sheet small enough to complete for an ordinary turn, then use it whenever a vendor changes the route or a customer asks where their data went.
Agentkit lets teams select a model for a chatbot and review conversation logs. If a provider-chain audit matters for your use case, establish the request-level evidence and vendor assurances your route requires before treating that selection as proof of the final processor.
No credit card required.



