Logged 2026-09-23
What happened when I ran an agent on a fixed-asset tie-out
Result: matched the reference answer exactly. Every control check built into the workbook passed before anything downstream would compute. 7 of 10 GL accounts tied, the same 3 accounts broke by the same dollar amounts as the reference, same file-level disposition. Graded 24 of 24 on a formula-driven grader that reads the completed workbook's cells directly, not the agent's own summary. Zero formulas in the template were altered — every number stayed derived, nothing got typed over.
| Task | Fixed-asset subledger-to-GL tie-out, one entity, one period end. Synthetic client, synthetic figures. |
| Agent | A general-purpose AI assistant with file and spreadsheet access — not a purpose-built accounting product. That distinction matters; see limits, below. |
| Inputs given | Exactly what a paid human tester would receive: a blank formula-driven workbook with its own embedded instructions, the client's fixed-asset register as exported, the full trial balance, and a one-page engagement note. |
| Withheld | The reference answer. Any indication this was a test, beyond what the engagement note itself said. No prior context on the workbook's design or history. |
| Grading | A reference answer computed independently, before the run, from the same source data by a separate process. A formula-driven grader reads the completed workbook and compares cell by cell. |
What it got right
The engagement note mentioned, in passing, that one asset had been scrapped. The agent used that to trace a paired $8,750 break across two accounts — the asset's cost and its accumulated depreciation, both still sitting in the register after the GL write-off went through. It didn't clear the entries itself; it flagged them, consistent with the engagement note's own instruction not to.
A separate, smaller break ($1,240, below the file's own escalation threshold) had no traceable cause anywhere in the data it was given. It reported that plainly instead of inventing an explanation — the same standard commentary gets held to applies here, and it held.
Only one fixed-asset report existed to serve as a control total. The agent used the report's own printed totals and said so in the required source field, rather than treating the single-report situation as a blocker.
A limitation it noticed in the workbook
Nothing in the workbook itself stops one person from typing a name into both the preparer's and the reviewer's fields — they're plain text, not an enforced control. Worth noting as a gap in the template, not a finding about anyone's real controls: in practice, segregation of duties is usually enforced upstream, by the source system's own record of who uploaded and who approved, not by two fields inside a workbook. The same gap shows up elsewhere when a review only tests that something exists rather than testing it independently.
Limits of this one run
One task, one scenario, one general-purpose assistant rather than a dedicated accounting product — a specialized tool might do better, or worse, in ways this doesn't test. I directed and observed this run; it wasn't handed to a stranger end to end the way a real handoff would be. And a result that matches a reference answer on one synthetic file says nothing about volume, about messier real-world exports, or about the exceptions a threshold-based check won't catch in the first place. One entry is one data point. For a related but different exercise — the same checks a review like this depends on, scripted and run against a full flux review rather than graded against a reference answer — see the flux analysis demo, and before any of this, check the export is complete.