Agents & Automation
What agents can and cannot do in the close
Software that runs a sequence of steps on its own — pull the export, match it against a rule, draft the note — is showing up inside close work now, usually under the label "agent." Some of that is genuinely useful. Most of what's written about it isn't written by anyone who runs a close.
"Agentic accounting" gets used loosely enough that it's worth being precise once, here, and then dropping the word. An agent, in the sense that matters for this article, is software that takes a multi-step task — pull the data, apply a rule, produce a draft, flag an exception — and runs the whole sequence without a person doing each step by hand. That's different from a model you paste text into and get text back.
It's worth saying plainly what that actually looks like, because most of what gets written about agents undersells it to make a tidier argument. I've watched one design and build a full reconciliation workbook end to end — deciding how a tie-break should rank when two results come out the same size, reworking a control-total requirement the original spec couldn't actually satisfy, catching its own formula bug before it shipped. That's reasoning, not lookup. Supervised closely, with someone checking the work before it goes anywhere, an agent does real judgment-shaped work today.
None of that is the question this article is actually about, though. The question that matters for a close isn't whether an agent can reason. It's this one:
Everything below follows from which side of that line the task sits on, not from what the agent is capable of in the abstract.
What agents are already good at in a close
Pulling two periods of an export and lining them up. Matching a trial balance to a chart of accounts. Flagging a balance that moved outside a threshold, or one that didn't move at all when it should have. Assembling a first-pass reconciliation from a subledger and a GL extract. Watching a checklist of dependencies and telling you which account groups are actually ready to work.
What these have in common isn't that the agent is only capable of simple things — it clearly isn't. It's that the result is checkable after the fact, against a formula, a threshold, a control total, so a wrong answer surfaces as a broken check instead of a debate. That's what makes it safe to let run without someone watching every step. The bar isn't whether the agent can do more than this. The bar is whether you can verify it fast enough for that not to matter. Set up well, an agent doing this kind of prep is faster and more consistent than a person doing it by hand at 11pm on close day. This is the same "Group 1" distinction that governs whether a model can write flux commentary for a given account — the reason has to already be in the detail. Agents extend how much of that detail gets assembled automatically. They don't change what the detail has to contain. Here's the same checkable-prep idea, scripted and visible step by step — no agent involved, just the same logic an agent would need to pass.
Where the drafts still need a person
An agent can produce something that reads like finished commentary — a sentence with a number in it and a plausible-sounding reason attached — without the reason being the actual driver. The same things that get commentary sent back from a human preparer get agent-drafted commentary sent back too: restating the movement instead of naming a cause, an unexplained residual, a driver that sounds right but wasn't checked against the source that would confirm it.
The failure mode is specific to agents in one way, though: a person who doesn't know an account tends to write vaguely, which is easy to catch. An agent that doesn't have the real driver tends to write confidently anyway, because fluency isn't gated on being right. That makes the review step more important, not less, exactly when teams are most tempted to skip it because "the agent already did it."
The exceptions a size-based threshold misses — a sign flip, a new account with no prior period, an account that should have moved and stayed flat — are also usually the cases where an agent working off a standard rule set misses the same things a rule-based review would miss. An agent is not, by itself, a fix for a control gap that a computed rule already couldn't close.
Reviewing agent output is not the same job as reviewing a person's
A review that only tests whether an explanation exists, not whether it's good, was already a weak control before agents were involved. It gets weaker once the volume of drafted work goes up, because more gets produced and the reviewer's actual bandwidth to check each item doesn't. If anything, the standard needs to get more explicit, not less, the moment prep speeds up: coverage against the movement, direction, a source you could actually go verify, specificity instead of a plausible-sounding restatement. Those are checkable whether a person or an agent wrote the first draft. That's the point of them.
Practically, that means treating an agent-prepared workbook the way you'd treat one from a new hire: formulas rather than typed values, every number traceable to its source, nothing you'd have to take on faith. An agent that produces a hardcoded figure with no trail is harder to review than a person's mistake, because there's no one to ask what they were thinking.
Where this actually helps right now
Prep and assembly, mainly — the parts of close that were always mechanical but took real hours to do by hand: pulling exports, running the comparison, flagging what crossed a threshold, drafting a first pass at the accounts where the driver is already sitting in the detail. That's real time back. It is not, yet, something you'd let decide on its own whether a stated cause is the actual cause, or what an unusual pattern across several accounts means for the close as a whole — not because it can't reason about those questions, but because nobody has built the equivalent of a formula check for "is this explanation actually true," and until that exists, a person has to be the one who verifies it. That's a narrower gap than "agents can't do judgment." It's closer to: the close doesn't yet have a way to catch it fast if the judgment is wrong.
I've started testing this directly — a standing log of agents run cold on real close tasks, each graded against a reference answer computed in advance, not a self-report. First entry: a fixed-asset-to-GL tie-out.
If you work this in Excel: the free flux template flags what needs explaining and shows the residual your drivers don't cover. The CloseOps Flux & Variance System ($79) starts from your trial balance instead: confirm each account's classification and it builds the income statement and balance sheet flux statements, then ranks what to investigate.
Related: getting started with agents in month-end close, checking a ledger export is complete before anything reads it, cold tests of agents against a reference answer, why AI can explain some accounts and not others, when automated drafting helps, and when it costs you more, why commentary gets sent back, reviewing work against a real standard, and the exceptions a threshold won't catch.
Found something wrong or missing? Tell me here — anonymous, thirty seconds.