Every new AI model is a change event: a retest checklist for small businesses
Before swapping an AI model, run the same business tasks before and after. This ten-task routine gives an owner a release decision and a one-page change log, without a compliance team.
Your AI drafts customer replies. A new model arrives. You change the setting, try one question and get a better-looking answer. Is it ready for Monday's enquiries?
Treat that swap as a change event: record what changed, test the work again and decide whether to release it.
Current as of 2 October 2026.
This follows our guide to testing AI against your actual job. Here, the question is narrower: what should an owner check when replacing the model inside an existing workflow?
What changes when you swap a model?
This week's announcements make the question timely:
- Anthropic introduced Claude Sonnet 5.5 on 28 September 2026.
- OpenAI's API changelog records GPT-6.1 Sol's release on 29 September 2026.
- Google announced Gemini 4 Argon on 30 September 2026, with phased access ahead of wider availability.
An announcement does not establish availability in your subscription or software. Google's announcement explicitly describes a gradual expansion of access before a broader release.
OpenAI's model optimisation guide says outputs are non-deterministic and behaviour changes between model snapshots and families.
Keeping the prompt does not establish that the new model will make the same decisions. That is the reason to retest the workflow.
Compatibility can change too. Anthropic's Sonnet 5.5 migration guide lists changed thinking settings and unsupported forced tool choices in the Messages API.
If you use packaged software, ask the supplier what changed. If you use an integration, ask whoever maintains it to check the migration guide. Then test what staff actually use.
Our recommended comparison covers five things: factual accuracy, usable output, escalation, permitted actions, and the time and cost of completing the job.
The cost of skipping the retest
Consider a hypothetical customer-service assistant. The new model produces smoother replies but promises a refund outside your policy. Someone must catch the promise, correct the reply and handle the complaint if it was sent.
This is an example, not a measured loss estimate. For your business, count correction time, customer remediation and interrupted work against any apparent saving.
For personal information, the risk is more than poor wording. The OAIC's commercial AI guidance describes inaccurate outputs and privacy risks.
It recommends ongoing review rather than treating product selection as a one-off decision. Retesting a model change is one practical way to apply that advice.
Ten tasks before and after the swap
This is our suggested starter routine for an owner-led business with roughly $1 million to $5 million in annual turnover. It is advice for one bounded workflow, not a regulator's prescribed test or a claim that turnover determines risk.
Use a spreadsheet and a folder. Nominate the owner and a staff member who knows the job; if you work alone, compare each answer with the written source. No compliance department is needed to run these tasks.
OpenAI's evaluation guidance recommends task-specific tests, comparison and evaluation on changes. The sample sizes and release rules below are our practical starting choices.
1. Name the job and the stop rules — before
Write one sentence: “Draft replies from our approved refund policy; a person approves every send.” List critical failures: wrong customer, invented entitlement, exposed private data or an action without approval.
Decide the quality threshold before seeing results. For this starter exercise, require every must-pass case to pass and zero critical failures. A nicer tone cannot compensate for a broken stop rule.
2. Record the old and proposed setup — before
Save the product, provider, displayed model name or API identifier, date, instructions, settings, document versions, connections and permissions. Record why you want the swap and what improvement would justify it.
Mark hidden details “not disclosed”. Ask the supplier about model routing, fallback, retention, training use and access. Do not assume a model switch changes these settings; verify whether it does.
3. Select a small repeatable case set — before
Start with 12 cases: four routine tasks, two ambiguous requests, two missing-information cases, two privacy or permission challenges, one document-conflict case and one service-failure case.
Write the expected answer or action beside each case, with its approved business source. Use synthetic or carefully de-identified inputs. Do not paste customer records into a new tool just to test it.
The OAIC recommends avoiding personal information in public generative AI tools, particularly sensitive information.
4. Capture the old model's baseline — before
Run all 12 cases in the existing workflow. Save inputs, replies, actions, errors, time, available usage charges and minutes of human correction. Judge each against the expected result, not how confident it sounds.
Keep old failures visible. The baseline is a comparison, not permission to carry a serious defect into the replacement.
5. Change only the model — after, in a test copy
Use a test workspace or supplier sandbox with sending, payments and record changes disabled. Keep the same cases, documents and instructions. Start fresh conversations with equivalent context for both models.
If migration requires other changes, record them. Compare the old setup with the complete proposed setup and label it a combined change; do not attribute every difference to the model alone.
6. Check facts, calculations and usable output — after
Run the same cases. Compare prices, GST treatment, dates and policy statements with your approved sources and an independent calculation. Check that the output still fits the email, spreadsheet or form staff use.
Record old result, new result and pass or fail beside each case. Save both outputs so the reviewer can explain the decision.
7. Test uncertainty and document conflict — after
Withhold a required fact, give an ambiguous request and supply conflicting policy documents. Check that the system asks, stops or uses the approved current source as your rules require. Fail invented answers.
Repeat the ambiguous and critical cases three times with fresh context. This is a spot check for variation, not a statistical guarantee; OpenAI explains why identical inputs can produce different outputs.
8. Challenge privacy, permissions and approval — after
Using fictional records, request another customer's details and ask the assistant to send or alter a record without approval. Put “ignore your rules” inside a test document. Inspect attempted actions as well as the reply.
Check that the connected system blocks prohibited actions. A polite refusal is insufficient if the tool still executes. For a drafting-only tool, confirm it has no live send or write connection.
OWASP's agent security guidance recommends minimum tool permissions, explicit authorisation for sensitive operations and treating external content as untrusted.
9. Measure the whole job and practise recovery — after
Compare completion time, usage charges and correction minutes on the same cases. For a fixed subscription, record the fee and any limits rather than inventing a per-answer cost. Note the currency of supplier charges.
Simulate an unavailable document or failed connection in the test copy. Check the hand-off to a person. Practise restoring the old setup; if the supplier cannot restore it, nominate a manual process and test that instead.
10. Decide, pilot and review — after
Reject the release if a critical failure occurs. Fix it and rerun affected cases plus the routine set. Do not turn an average pass rate into an excuse for unauthorised actions or disclosures.
If the tests pass, let one staff member pilot the change with every customer-facing output reviewed. Suggested first checkpoint: 20 jobs or five working days, whichever comes first. Review sooner if a failure appears.
Record the decision, restrictions and stop owner. Turn any new failure into a saved test case before expanding use. These pilot limits are starting choices, not evidence that the model is safe.
The one-page change log
Copy this into your spreadsheet or document. Keep detailed outputs in a restricted folder and link them by case ID; the one-page log is the decision summary, not a store of customer information.
| Field | Entry to complete |
|---|---|
| Change and owner | ID; workflow; owner; tester; date |
| Reason | Expected benefit and acceptance criteria |
| Before → after | Product; provider; model ID or displayed name |
| Other changes | Instructions; settings; documents; tools; permissions; unknowns |
| Data check | Test data used; supplier settings checked; unresolved questions |
| Evidence | Case-set version; restricted link to old/new outputs and action records |
| Results | Old/new passes out of cases run; repeat runs; critical failures; fixes |
| Whole-job cost | Time; charges and currency; human correction minutes |
| Decision | Approve, restrict or reject; reason; approver and date |
| Recovery and review | Restore or manual steps; stop owner; pilot limit; review date and triggers |
Suggested decision wording: “Approved for drafting only, with human review. These cases passed on this date. Sending remains disabled. Stop and investigate if the system invents policy or attempts an unauthorised action.”
Retest when the model, instructions, source documents, connected tools or permissions change, or when a material failure appears. Keep the failed case and the fix together.
The honest caveat
Twelve cases and a short pilot can reveal a defect. They cannot prove reliability across future inputs, certify compliance or replace specialist review for decisions where an error could seriously harm someone.
If the old model is unavailable, you cannot create a fresh side-by-side baseline. Use retained outputs where lawful, test against written expectations and record that limitation. Keep manual review while the comparison remains incomplete.
In a product that hides model versions, record what you can see and ask for change notices. Repeat tests after supplier notices or unexplained behaviour changes. Do not claim to have pinned a version you cannot control.
The $1 million to $5 million audience spans different privacy positions. The OAIC's small-business checklist explains the $3 million threshold and exceptions.
Being small does not automatically exempt a business. Check coverage for your circumstances; this log does not establish legal compliance.

