Chatbot Model Deprecation Checklist Before Shutdown

Use this chatbot model deprecation checklist to find hidden dependencies, set a cutover deadline, test replacements, and prove every route moved.

Cover Image for Chatbot Model Deprecation Checklist Before Shutdown

A new model release is a good time to check the old one. Anthropic launched Claude Sonnet 5.5 on September 28, but a replacement announcement is not itself a retirement notice. Anthropic's model lifecycle page separates active, legacy, deprecated, and retired models. It also warns that partner platforms can set different dates.

For a customer-facing chatbot, the work starts with an inventory. Record every model ID and serving platform, find every route that can call it, check the provider's dated notice, assign a cutover owner, test the replacement, and confirm that production usage of the old ID reaches zero. Keep the inventory below as the working record. A green test on the primary chat route does not clear a fallback, a scheduled job, or an old bot that still receives traffic.

The retirement inventory

Create one row per model and platform, not one row per vendor. The same model accessed through a direct API and a cloud partner may have different availability dates. Include a row even when the provider has announced no shutdown date; that is an explicit state to recheck, not a permanent exemption.

FieldRecord thisEvidence to attach
Model and platformExact request ID, provider, direct API or partner service, and regionRedacted request log or configuration export
Lifecycle statusActive, legacy, deprecated, retired, or unknownProvider status page and date checked
Last usable dateConfirmed shutdown date, or not announcedLink to the dated notice; never infer from a new release
Call sitesMain chat, fallback, evaluation, batch, voice, test, and archived bot routesCode/config search plus observed usage
ExposureDaily calls, affected chatbots, traffic share, and risky intentsUsage export over a stated window
ReplacementCandidate model ID and any API or behavior changesProvider migration guide and test results
Owner and gatePerson, test deadline, cutover window, rollback ruleChange record and approval
CloseoutOld-ID calls after cutover, error rate, and final disable dateUsage export and production logs

Google's Gemini deprecation table makes the distinction concrete: some models have a shutdown date, while others say no shutdown date has been announced. OpenAI maintains a separate API deprecations page with notices, removal dates, and replacements. Use the page for the platform you actually call. A search result snippet or an old internal spreadsheet is too easy to mistake for the current schedule.

Find the model behind every route

Start at the provider's usage export, then work back toward configuration. Anthropic documents a Console export broken down by API key and model. Other platforms offer their own usage and billing views. A code search alone misses IDs stored in a database, environment variable, dashboard, workflow tool, or customer-specific setting. Usage alone misses a dormant fallback that will wake up during the next outage. You need both views.

Search for exact IDs, aliases, and the configuration keys that select them. Look at production, staging, scheduled jobs, evaluation scripts, support macros that invoke an API, and old chatbot versions still embedded on a page. Ask who owns each API key before treating a usage row as harmless. A key with no clear owner is a finding, not evidence that nobody uses it.

Then map the paths a visitor can trigger. A support bot may answer normal questions with model A, summarize the transcript with model B, and call model C only when the main provider times out. If B is retired, the customer might still see a normal reply while the human handoff loses its summary. If C is retired, the failure appears only on a bad day. The fallback testing guide gives that hidden route a separate test matrix.

For each path, capture one recent redacted request with the requested model ID and the model that served the response when the provider exposes it. Keep the timestamp and platform. Do not rely on the model name shown in an admin picker; it may be a friendly label, a default, or an alias. If you cannot determine the ID from a real request, mark the row unknown and investigate before setting a cutover date.

Put the deadline on the calendar correctly

A lifecycle page can say "not sooner than" without announcing a retirement. That phrase is a floor, not a deadline. Likewise, the release of a successor does not mean the older model has been deprecated. Record the actual notice date and shutdown date separately. When neither exists, record not announced and set a recurring check of the official status page.

Anthropic says it gives active customers at least 60 days' notice before retiring a publicly released model on its operated platforms. That is a provider commitment, not a universal migration window. Its direct API, Claude Platform on AWS, and Microsoft Foundry use the dates on Anthropic's page; Amazon Bedrock and Google Cloud set their own schedules. Google and OpenAI use their own terminology and tables. A team using two providers should not combine their notices into one generic calendar entry.

Work backward from a confirmed shutdown date. Leave room for the replacement's technical migration, a test run against real support questions, a controlled production shift, and observation before the old route disappears. If a vendor gives short notice, reduce scope or move traffic to a tested alternative. Do not compensate by deleting the observation window and calling a smoke test a migration.

Assign one person to monitor vendor notices and one person to approve the customer-facing cutover. Those can be the same person on a small team, but write the name down. Add the notice URL and a screenshot or dated export to the change record. A calendar reminder with no source link will be hard to verify six weeks later.

A worked retirement record

Suppose Northline Gear has a returns chatbot that uses support-model-old through a direct API. The provider announces on October 1 that this exact ID will stop serving on December 1. Northline observes 6,000 calls to it during the last 30 days. Five thousand come from the main chatbot, 700 from a fallback route, and 300 from a transcript summary job.

Its first inventory row looks like this:

model_id: support-model-old
platform: direct-api
notice_checked: 2026-10-01
status: deprecated
shutdown_at: 2026-12-01
usage_window: 2026-09-01..2026-09-30
calls: 6000
routes:
  main_chat: 5000
  fallback: 700
  transcript_summary: 300
owner: support-engineering
candidate: support-model-new
cutover_gate: policy-cases-pass-and-old-id-calls-zero

Those numbers are hypothetical, but the arithmetic matters: a migration of only the visible chat route leaves 1,000 calls per month on the retiring ID. At the observed rate, roughly one in six calls would still be exposed after December 1. The 300 summaries are particularly easy to overlook because they run after the visitor has left.

Northline tests the candidate on saved returns conversations, including an opened item, a final-sale item, and a request that needs a human. It verifies both the answer and the action payload. A prompt that still produces good prose but calls the refund tool with a missing order ID fails the gate. The team then moves a small share of traffic, checks logs and support outcomes, moves the rest, and watches the old-ID usage by route. The migration closes only after every row shows zero old calls for a full normal traffic cycle and the fallback has been exercised deliberately.

If the shutdown is unexpectedly accelerated, Northline disables the unsupported route and sends affected chats to a human or a previously tested model. It does not leave a broken fallback in place in the hope that the provider will keep accepting requests.

Test the replacement as a new dependency

A replacement is rarely a string change. The Claude Sonnet 5.5 migration guide, for example, documents settings that can return HTTP 400 after a model switch. It also says the same text can produce about 30% more tokens than on several earlier Claude models. This is why the inventory links to a model-specific migration guide, and why cost, token limits, and error handling belong in the test run. The guide's details apply to its listed starting models; do not copy them into an unrelated provider migration.

Freeze the source snapshot and prompt version while comparing the old and candidate routes. Use redacted conversations that represent your actual support traffic. Include high-volume questions and the small set of policy cases where an incorrect answer does damage. Grade required facts, forbidden claims, handoff behavior, action payloads, and language coverage. The model evaluation scorecard gives a repeatable way to compare those outputs without rewarding a longer answer merely for sounding confident.

Run at least one request through every real entry point: website widget, direct API, human handoff, scheduled task, and fallback, where applicable. Some failures are configuration failures, not answer-quality failures. A deprecated parameter may cause a 400 before the model writes a word. A changed output format may pass a chatbot demo but break the downstream summary parser. Check response status, parsed fields, total time, and billed usage alongside the transcript.

Keep two release gates separate. The technical gate checks that every route can call the new ID and parse its response. The customer gate checks that the bot gives correct policy answers, escalates when required, and stays within service and cost limits. Both must pass before the main route moves. A provider's suggested replacement is a useful candidate, not a guarantee that it follows your refund policy.

Cut over, then prove the old ID is quiet

Move traffic in a way you can reverse while the old model still works. Record the switch time and configuration version. Compare production errors, latency, cost, handoffs, and negative feedback against the baseline. If a critical policy answer fails, use the rollback checklist to restore the previous route during the remaining support window. Once the provider retires the old ID, that rollback option disappears; prepare a second tested route or a human handoff before that day.

After the cutover, query usage by model ID and API key. Look for nonzero calls from an old environment, a customer-specific bot, or a fallback that runs only when the main service is slow. Exercise the fallback on purpose and verify that the old ID stays at zero. Check the same report again after a normal traffic cycle, including any weekly scheduled job. A zero in a ten-minute sample is weak evidence for a job that runs every Sunday.

Archive the notice, inventory, test results, change approval, cutover timestamp, and final usage export together. Remove the old ID from selectable defaults and deployment configuration only after you know which saved records still need to load. Keep a historical label for old conversations if deleting it would erase audit context. The goal is a production route that no longer depends on the retiring endpoint, with evidence that covers the quiet paths too.

A model retirement is finished when the provider's deadline can arrive without surprising a customer or a support agent. The inventory gives you a way to prove that: every route has an owner, the replacement passed real policy cases, and the old ID stopped receiving calls.

Build your chatbot for free → No credit card required.

Get started freeNo credit card required