GLM 5.2 Matches Human Bookkeeper Accuracy on UK VAT Returns - With Some Caveats

TL;DR
A new benchmark shows GLM 5.2 processing 59 transactions and producing VAT returns off by only 7 pence - at $2.73 versus typical accounting fees of $1,000+.
Last updated: August 14, 2026
Update (August 14, 2026): This benchmark used GLM 5.2. GLM-5.3 launched today - same base model with scaled-up post-training and selectable reasoning levels, at the same price. The VAT results below have not been rerun on 5.3; treat them as a floor for the model line rather than the current ceiling.
A benchmark published by Toot Books today showed GLM 5.2 preparing quarterly VAT returns for a UK small business with near-human accuracy - and at a fraction of the cost. The Hacker News discussion hit 132 points and 79 comments, with the conversation quickly pivoting from "wow this works" to "but who goes to prison when it doesn't?"
The benchmark is one of the most concrete demonstrations yet of LLMs performing structured financial compliance work. But the details matter more than the headline.
What the benchmark actually tested#
The setup: GLM 5.2 ran on an isolated Google Cloud instance with access to accounting software and a command-line tool. It received bank feeds, receipt PDFs, and two user notes providing context - the same inputs a human bookkeeper would receive.
The task: process 59 transactions and produce a quarterly VAT return.
The numbers:
| Metric | GLM 5.2 | Human Accountant |
|---|---|---|
| Processing time | 68 minutes | Variable (hours to days) |
| Cost | $2.73 | $1,000-2,800/quarter |
| Net position accuracy | Off by 7 pence (~10 cents) | Ground truth |
| Transactions processed | 59 | 59 |
| Total checks evaluated | 354 (6 criteria x 59) | - |
The model achieved this cost efficiency partly because 93% of prompt tokens hit the provider's cache at reduced rates.
What the model handled well:
- Correct account classification for standard transactions
- Invoice matching to bank entries
- Disambiguating complex scenarios: splits, transfers, duplicate entries
What the model got wrong:
The benchmark documented 20 failures across 18 transactions. The most serious:
- Misclassified founder capital: A $10,000 founder share capital entry was logged as "Capital Account" instead of "Unpaid Shares" - a distinction with potential legal audit implications
- VAT category confusion: 14 instances of mixing up zero-rated versus exempt VAT categories
- Split-transaction VAT errors: 3 cases of incorrect VAT allocation on split entries
The benchmark authors acknowledge a key scope limitation: "The job performed by the humans was broader than what was requested of the model. Humans also had to find the relevant invoices (searching through mailboxes, or requesting them from providers) and reason through circumstances which cannot be inferred from the bank feed and invoices alone."
In other words: the model got the easy version of the task.
What HN is saying#
The Hacker News discussion focused less on whether the tech works and more on what happens when it does not.
The liability question dominated. As one commenter put it: "This is a prime example of a problem space where accuracy matters, but it also matters who ultimately goes to prison. I'm going to go out on a limb and guess it's not the LLM."
The distinction matters. If you hire an accountant and they commit fraud, your liability is limited to some extent - you acted in good faith by engaging a professional. If your LLM decides to commit tax fraud, you are in uncharted legal territory.
The "nearly as accurate as a human" framing drew pushback. One commenter noted: "Humans aren't exactly known for perfect recall" - implying that human bookkeepers make mistakes too, so matching their error rate is not necessarily impressive. Another referenced the classic "60 percent of the time, it works every time" line.
But practitioners were already doing this. Several commenters shared that they are actively using Claude Code, DeepSeek, and other models for bookkeeping in production:
- One commenter uses Claude Code with FreeAgent to match PDFs to invoices and handle VAT
- Another built beansync, a "vibe-coded deepseek bookkeeping" system that parses emails, extracts numbers, and correlates transactions
- A third uses Claude Code with Opus to keep beancount ledgers up to date from Mercury bank feeds
The trust question remained unresolved. "I'd be scared shitless to even try something like this," one commenter wrote, noting that the company behind the benchmark has minimal public presence - "just a company Vineyard Finance LTD that was incorporated last year."
Where this actually matters#
The benchmark makes a compelling case that bookkeeping as pure classification work is largely solved. Take a bank feed, match it to invoices, assign categories, calculate VAT - this is pattern matching with well-defined rules. LLMs are good at this.
But the benchmark also reveals the limits:
Edge cases require domain expertise. The founder capital misclassification would not be caught by someone reviewing outputs casually. You need to know that "Capital Account" and "Unpaid Shares" have different legal meanings.
VAT rules are surprisingly complex. Zero-rated versus exempt is not obvious from the transaction itself - it depends on the nature of the goods or services and the specific regulatory category. The model confused these 14 times out of 59 transactions.
The human loop matters. Every commenter using LLMs for bookkeeping mentioned review steps. The model generates candidates; a human approves. This is not autonomous bookkeeping - it is assisted data entry with smart defaults.
The cost math#
The cost comparison is dramatic on its face: $2.73 versus $1,000-2,800 per quarter. But that comparison elides several factors:
- The human accountant also does invoice retrieval, which the model did not
- The human accountant takes liability for errors
- The human accountant knows when to escalate unusual situations
If you factor in a human review step, the LLM approach still wins on cost - but the margin narrows. You are not eliminating the accountant; you are giving them a first draft that is usually right.
For small businesses with simple books, this might be transformative. For businesses with complex VAT situations, cross-border transactions, or audit risk, the human accountant is not going away.
The bigger picture#
This benchmark is part of a broader pattern: LLMs getting good enough at structured compliance work that the question shifts from "can it do this" to "should it."
The technical capability is clear. GLM 5.2 processed 59 transactions with a 7-pence error on net position. That is better than many humans would do on their first pass.
The harder questions are institutional:
- Who is liable when AI-prepared returns contain errors?
- How do you audit AI-assisted financial records?
- What happens when HMRC (or the IRS) starts using AI to audit everyone?
As one commenter put it: "It's not hard to imagine tax authorities using AI to audit everyone's tax returns every year."
The asymmetry is notable: if the tax authority uses AI to catch errors, and you used AI to make errors, the human in the middle is you.
Continue Reading#
- Cheap subagents are better when their work is visible
- The DD Stack Cookbook: Five Recipes That Compose
- Deep Research Agents Need Constraint Ledgers
Sources#
- Toot Books GLM 5.2 VAT Benchmark - full methodology and results
- Hacker News discussion - 79 comments as of this writing
- Digits AI vs Human Bookkeeper Benchmark - referenced in discussion
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on AI coding tools
GLM 5.2 and the AI Margin Collapse Thesis
Martin Alderson's argument for why open-weights models like GLM 5.2 will compress frontier lab margins is sparking debate on HN. Here is what the thesis actually says, where HN agrees and disagrees, and why it matters for developers choosing models.
7 min readColibri: Run GLM 5.2 on a 32GB Laptop With Disk Streaming
A solo developer built a 1,300-line C inference engine that runs the 744B GLM 5.2 model on consumer hardware by streaming routed experts from disk. Here's how it works.
6 min readArmin Ronacher on The Coming Loop and Why Agent-Driven Code Still Needs Human Comprehension
Armin Ronacher's new essay explores the tension between letting AI agents loop autonomously and maintaining the engineering comprehension that makes software maintainable. The Hacker News discussion adds practical caveats worth reading.
9 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








