This kit lets you run the same test I ran on Claude, ChatGPT and Microsoft Copilot: one broken programme budget, one loose prompt, and a scorecard that measures whether the tool told you what it did.
Who is this for
Anyone who gets handed spreadsheets they didn't build and is being told to use AI on them. Project and programme managers, finance business partners, operations leads. You don't need to be an Excel expert. You need to be able to compare two numbers.
What problem does it solve
Vendor demos use clean data. Real handovers have merged cells, hidden rows, dead dashboards and one number that was pasted in by someone who has since left. This kit shows how a tool behaves on that, and in particular whether it changes your inputs without saying so. That is the failure that gets past review.
What is inside
The workbook: ten sheets, 729 formulas, 40 deliberate faults, one planted error on the travel row worth exactly £3,500.
The prompt: one sentence, plus the only follow-up line you're allowed to give.
The scorecard: five criteria at 20 points each, plus an unscored list of the other faults so you can see how thorough the tool was.
The answer key: every fault, cell by cell, sealed until you've scored.
How to use it
Copy the workbook, hand the copy to one tool, send the prompt, take what comes back without iterating, then diff the returned file against the original before you score. Full steps and the diff method are on the Notion page.
What it will not do
It won't tell you which tool to buy. One file, one run, one day. It will show you how a tool treats your numbers when nobody is watching, which is the thing worth knowing before you trust it with a real one.
Related
Subscribers get the template
Enter the email you use for the Cliffinkent newsletter. If it is confirmed, the worksheet opens here.
No account. No download portal. Just a subscriber check.