Month·End·Close
← Month-End Close

Agents & Automation

Cold tests of agents on real close work

A standing log, not a single post. Each entry: an agent that's never seen this workbook or this client, a real close task, and a reference answer computed independently before the run. Graded by formula against the completed file — not by asking the agent how it thinks it did.

This isn't a benchmark and I'm not claiming it's a standard anyone else should be measured against. It's one practitioner running one task through whatever's available, the same way I'd size up any tool before trusting it near a real close. New entries get added as I run more tests. Nothing here is accounting advice; facts, materiality thresholds, and control environments vary by company. All data is synthetic — see About.

Every entry states plainly what the agent was given, what it wasn't told, how it was graded, and what this one run doesn't prove. That's the whole method.

Run the same test yourself. The kit below is exactly what the first entry's agent received — the workbook, the client's two exports, the engagement note — plus a formula-only grading sheet with the reference answer already in it, so you can score your own agent's result without sending anything back to me. Nothing uploaded, nothing tracked.

Download the test kit (.zip)

Entries: September 2026 — fixed-asset-to-GL tie-out


Logged 2026-09-23

What happened when I ran an agent on a fixed-asset tie-out

Result: matched the reference answer exactly. Every control check built into the workbook passed before anything downstream would compute. 7 of 10 GL accounts tied, the same 3 accounts broke by the same dollar amounts as the reference, same file-level disposition. Graded 24 of 24 on a formula-driven grader that reads the completed workbook's cells directly, not the agent's own summary. Zero formulas in the template were altered — every number stayed derived, nothing got typed over.

TaskFixed-asset subledger-to-GL tie-out, one entity, one period end. Synthetic client, synthetic figures.
AgentA general-purpose AI assistant with file and spreadsheet access — not a purpose-built accounting product. That distinction matters; see limits, below.
Inputs givenExactly what a paid human tester would receive: a blank formula-driven workbook with its own embedded instructions, the client's fixed-asset register as exported, the full trial balance, and a one-page engagement note.
WithheldThe reference answer. Any indication this was a test, beyond what the engagement note itself said. No prior context on the workbook's design or history.
GradingA reference answer computed independently, before the run, from the same source data by a separate process. A formula-driven grader reads the completed workbook and compares cell by cell.

What it got right

The engagement note mentioned, in passing, that one asset had been scrapped. The agent used that to trace a paired $8,750 break across two accounts — the asset's cost and its accumulated depreciation, both still sitting in the register after the GL write-off went through. It didn't clear the entries itself; it flagged them, consistent with the engagement note's own instruction not to.

A separate, smaller break ($1,240, below the file's own escalation threshold) had no traceable cause anywhere in the data it was given. It reported that plainly instead of inventing an explanation — the same standard commentary gets held to applies here, and it held.

Only one fixed-asset report existed to serve as a control total. The agent used the report's own printed totals and said so in the required source field, rather than treating the single-report situation as a blocker.

A limitation it noticed in the workbook

Nothing in the workbook itself stops one person from typing a name into both the preparer's and the reviewer's fields — they're plain text, not an enforced control. Worth noting as a gap in the template, not a finding about anyone's real controls: in practice, segregation of duties is usually enforced upstream, by the source system's own record of who uploaded and who approved, not by two fields inside a workbook. The same gap shows up elsewhere when a review only tests that something exists rather than testing it independently.

Limits of this one run

One task, one scenario, one general-purpose assistant rather than a dedicated accounting product — a specialized tool might do better, or worse, in ways this doesn't test. I directed and observed this run; it wasn't handed to a stranger end to end the way a real handoff would be. And a result that matches a reference answer on one synthetic file says nothing about volume, about messier real-world exports, or about the exceptions a threshold-based check won't catch in the first place. One entry is one data point. For a related but different exercise — the same checks a review like this depends on, scripted and run against a full flux review rather than graded against a reference answer — see the flux analysis demo, and before any of this, check the export is complete.


Most of this surfaces again at flux review. The free flux template flags which balances moved enough to explain. The CloseOps Flux & Variance System ($79) starts from your trial balance instead: confirm each account's classification and it builds the income statement and balance sheet flux statements, then ranks what to investigate.

Related: getting started with agents in month-end close, what agents can and cannot do in the close, the commentary test — the same computed-grading idea, for flux commentary, the formula-only workbook these tests actually run on, and why some accounts can be explained from the detail and others can't.

Found something wrong or missing? Tell me here — anonymous, thirty seconds.