~/tests/claude-for-word-planted-fault-board-paper
// hands-on test · 02 Sept 2026

Claude for Word: 16/17 on a planted-fault board paper

Claude for Word on launch-day Fable 5.1 scored 16/17 against six plants in a board paper. Paper £1.24m, cost model v3 £1.43m. The 34% benchmark was the only partial mark.

verdict · use

Use it to review comments. Use it to fact-check after a source I have given it and I am happy with. Use it for the final cut, with two conditions: the rewrite arrives as one insertion and one deletion, so it is all or nothing; every number in that rewrite is mine, because Word has already put my name on the revisions. Do not stop reading the paper.

what worked
  • P1 Karen vs Kieran 2/2, including that Karen's cut would delete the sentence Sandra needs
  • P1 Sandra already answered +1, with the benefit lines and the seven-year payback against a five-year policy
  • P1 Mark's date request 2/2, including the low-risk consultancy rating pages away
  • P2 stale total and the three lines behind it 2/2, with change-log reasons rather than a guess
  • P2 payback claim 2/2, including that it fails the company's own policy
  • P3 Braithwaite's survived the cut, went to the top of the list, 2/2
  • P3 DPA line survived +1
  • P4 copy-edit traps 2/2 and Priya's exact figures +1
  • Three unplanted errors flagged: SSO/bank-file date collision, identity change against UAT, Priya's rounding comment against the draft
  • Did not invent a source for the made-up 34% benchmark
what broke
  • P2 34% benchmark 1/2: saw the figure, could not verify it from the model, stopped, lost the mark
  • The 34% then vanished from the rewrite and was listed under drops without a reason
  • 847 words against an 800-word brief; Claude had reported 790 excluding title and sign-off
  • Tracked changes authored as me, not Claude, so prompt 4's nine edits merged into the existing revisions
  • One insertion and one deletion, so the rewrite cannot be accepted in part, and it contains figures Karen has not signed off

Claude for Word, running Fable 5.1 about two hours after that model landed, scored 16/17 on a planted-fault board paper. The paper said the programme cost £1.24m. Cost model v3, sitting next to it, said £1.43m. I had put five other problems in with that mismatch. The add-in caught the ones that were in the files. It lost a mark on a 34% benchmark I had made up, because it would not invent a source.

Watch the recording. If you work more in Excel than Word, there is already a video of that test.

How I tested it

Claude for Word, from the Add-ins tab. Fable 5.1 on launch day. One constructed business case with reviewer comments in the margins, plus the cost model spreadsheet. Four prompts, in order, scored against an answer key I wrote first. How many of the six plants did it catch, what did it quietly change, and what is the one change it made that you will never find.

The paper is the sort you get burned by: written on a Friday, read on a Monday, numbers only moving by the Wednesday. Two reviewers contradict each other in the margins. A supplier asks for a go-live date the plan cannot give him. One risk buried in paragraph four changes what the board should decide.

Run this yourself

Run the test yourself: Claude for Word board paper pack. It has the Word paper, the cost model, the four prompts and a sealed answer key. The scorecard PNG stays in this article.

The four prompts

Prompt 1
Summarise the reviewer comments and tell me where they conflict with each other or with the document.
Prompt 2
Check every figure in this document against the attached cost model and flag anything that does not match.
Prompt 3
Cut this to an 800-word executive summary using tracked changes. Keep every risk that would change the board's decision.
Prompt 4
Copy-edit for a UK board audience. Fix spelling, consistency and jargon. Do not change meaning.

Score

scorecard, final 16/17, P2 34% benchmark the only partial mark
Final 16/17. P2 34% benchmark is the only partial mark.

Prompt 1: comments. 5/5

Karen and Kieran are opposed. Claude offered a resolution and noticed that Karen's cut would delete the sentence Sandra needs. That is 2/2 on Karen vs Kieran. It also found that Sandra's point was already answered: the four benefit lines sum to 312,000 with nothing deducted, and payback is seven years against a five-year policy. Bonus mark.

Mark's date request collided with the payroll-adjacent timeline, a fixed statement of work date, the comms plan, and a consultancy-availability rating sitting at low risk. 2/2, including the rating that lived pages away from the comment.

It also flagged three errors I did not plant. Kieran's SSO comment has the bank-file work depending on a 30 September identity change, while the bank test file exchanges are booked for 15 September. The identity change is written as landing ahead of user testing, against the UAT date in the paper. Priya's comment says not to round up; the draft rounds down, so her comment is worded wrong. Five out of five on prompt 1.

Prompt 2: maths. 5/6

I attached the cost model and asked it to check every figure. It said the paper is built on v2 figures, not v3, and that this cascades through the document. Implementation partner fees up £54k. Data migration contractors up £36k. Contingency up £100k. The board asked for 12.5% in July; the paper still says 5%. The paper states £1.24m twice, including under option three, single cutover. The new total is £1.43m, and the reasons are in the change log. It did not guess.

Stale total and the three lines behind it: 2/2. Payback claim that only works on the wrong total, and fails the company's own policy: 2/2. The 34% comparable savings: it saw the figure, could not verify it from the model, and stopped. 1/2. Running total 10/11.

Prompt 3: the cut. 13/14

I asked for an 800-word executive summary via tracked changes, keeping risks that would change the board's decision, and accepted the edits. Word count came in at 847, excluding title and sign-off. Claude had said 790. About fifty over. Hover the revisions and the author is me, not Claude. Word does not know the difference. Neither will the board.

Braithwaite's survived the cut: the biggest customer's invoice format cannot be validated before go-live, and the fallback only works once that is in. It went to the top of the list. 2/2. The data processing agreement line at the end of that paragraph stayed in. Bonus mark.

Every number in the rewrite came from the August v3 figures. Old totals gone. Both payback bases stated. The policy breach listed as a risk. It offered to put the 34% back and listed that drop, without saying why it dropped it. Reviewer comments remained attached to the versioned text. Insertions and deletions arrived as one decision: the entire rewrite or none of it. The rewrite includes numbers Karen has not signed off.

Prompt 4: UK copy-edit. 2/2

Nine changes across eight paragraphs, as tracked changes, then merged into the existing revisions because they still sat under my name. I had to ask what it changed. It added a dash, switched ledger and payables to receivables and invoices, and added a glossary for the first instance of hypercare periods. Sensible. 2/2. Grand total 16/17.

The six plants

Stale total and the three lines behind it: caught. Payback claimed on the wrong total, which fails the company's own policy: caught. A 34% benchmark I made up: it did not catch a source, and it did not invent one either. That is where it lost the mark. Two reviewers in conflict: caught, with the escalation. A supplier date that was not possible, which also broke a risk rating pages away: caught. Buried risk, Braithwaite's: survived the cut and was escalated to the top of the list.

What I would let it do

I would let it review the comments. It reads the margins with more patience than most people read the paper. I would let it fact-check after a source I have given it and I am happy with. I would let it make the final cut, with two conditions. One insertion and one deletion means all or nothing, which keeps the rewrite internally consistent. Every number in that rewrite is a number I now have to put my name to, because Word already has.

Then it edited its own tracked work. Word merged those later localisation edits and you lose sight of them. That is Word attributing add-in edits to the signed-in user. You can send a file believing you can still see everything. Compare as you go.

One document, one model, first day: 16/17 on the paper, and zero on the thing I did not tell it to check. If you write papers that go up the chain, that is worth the money. If you are hoping to stop reading them, no. Accountability stays with you. Read the thing.

Limits of this test

This is one constructed paper, one cost model, one pass through four prompts, on launch-day Fable 5.1. It does not rank Claude for Word against another add-in. It does not say the same score will hold on a live client pack, a later model, or a paper whose faults sit only in a file I did not attach.

Prompt or workflow
Four prompts in order from the Add-ins tab, cost model attached before prompt 2, tracked changes on for the cut. 1) Summarise the reviewer comments and tell me where they conflict with each other or with the document. 2) Check every figure in this document against the attached cost model and flag anything that does not match. 3) Cut this to an 800-word executive summary using tracked changes. Keep every risk that would change the board's decision. 4) Copy-edit for a UK board audience. Fix spelling, consistency and jargon. Do not change meaning. Scored against a 17-point answer key written first.
disclosures
  • Constructed board paper and cost model. Names, companies, dates and figures are invented. No client, employer, colleague or personal data.
  • Hands-on test of Claude for Word on Fable 5.1, launch day. No free access, sponsorship or affiliate arrangement.