Month·End·Close
← Month-End Close

Agents & Automation

Getting started with agents in month-end close

Most people looking into this are exactly where I was a few months ago: aware it's coming, unsure where to actually start. This is the order I'd do it in — built from what's already worked here and what's already broken.

This is the hub for everything on this site about agents in close work. Read this first, then go deeper on whichever piece applies to you. Four things, roughly in the order they actually matter: get the vocabulary straight, know which accounts are even worth trying, test one before it touches anything real, and know what still needs a person no matter how the test goes.

The question that matters isn't whether an agent can reason well. It's whether you'd let this run unattended against this period's real numbers — or whether the result needs a person to check it before anyone relies on it. Everything below follows from which side of that line a task sits on.

1. Get the vocabulary straight

"Agentic accounting" gets used loosely enough that it's worth being precise once. An agent, here, is software that takes a multi-step task — pull the data, apply a rule, produce a draft, flag an exception — and runs the sequence without a person doing each step by hand. That's different from a model you paste text into and get text back.

What that actually looks like in practice, and where the line sits between "genuinely useful today" and "not yet, and here's specifically why": what agents can and cannot do in the close.

What "agentic" adds over plain automation

A fixed script or macro runs the same steps every time, regardless of what it finds. What makes something agentic, in the sense used here, is that it decides step to step what to pull, compare, or flag next — the sequencing itself is not hardcoded. That decision-making is limited to orchestration and drafting, though: the arithmetic behind a flagged variance is still computed by rule, not guessed by the agent. That's also why the checks in section 4 below don't get lighter just because the sequencing got smarter.

2. Know which accounts are even worth trying

Before you point an agent at anything, sort your accounts into three groups. This is the single biggest time-saver in the whole exercise, because most of the frustration people report comes from starting on the wrong account, not from the tool being weak.

Full version, with the check that comes before any of this (tie the total to the external source first): why AI can explain some accounts and not others.

3. Test one before it touches anything real

And before an agent reads any ledger export, check the export is complete: opening + activity = closing, for every account. An agent working from an export with a missing entry can't tell it from a quiet month. For what a checked process looks like when the data is broken, try the flux analysis demo.

Don't hand an agent a live client file as its first run. Run it cold on a synthetic task with a reference answer computed in advance, graded by formula against the completed file — not by asking the agent how it thinks it did. That's the only way to find out what it actually does before the outcome matters.

I run these myself and log them as a standing series: cold tests of agents on real close work. The first entry is a fixed-asset-to-GL tie-out, matched to the reference answer exactly — along with what it got right, and the limits of what one run like that proves.

You can run the same test against your own agent. The kit is the exact packet the agent in that entry received — the workbook, the client's two exports, the engagement note — plus a formula-only grading sheet with the reference answer already in it, so you score the result yourself.

Download the test kit (.zip)

4. Know what still needs a person, every time

An agent can produce something that reads like finished commentary without the reason being the real driver — fluently, which makes it easier to accept than a person's vague guess would be. Reviewing agent output isn't the same job as reviewing a person's, and it needs to be more explicit, not less, the moment prep speeds up: coverage against the movement, direction, a source you could actually verify, specificity instead of a plausible-sounding restatement.

Those tests are checkable whether a person or an agent wrote the first draft. Why a review that only tests that commentary exists is a weak control, agent-drafted or not — and the specific reasons commentary gets sent back, which don't change just because the draft came from a tool.

A starting sequence

  1. Read what agents can and cannot do in the close.
  2. Sort your own accounts into the three groups above, before you try anything.
  3. Download the test kit and run it against your own agent before it touches real data.
  4. Grade the result yourself, against the reference answer that ships with it. Would you sign it, in under an hour of review?
  5. If it passes, start on one Group 1 account, on real but non-critical work, reviewed in full every time, for at least one full close cycle.
  6. Log what breaks. That's the point of the standing log below — and worth keeping your own version of it, even privately, before you trust an agent on anything that matters.

The log, updated as it grows

New entries get added to the cold-tests log as I run more of them — different tasks, different agents, same method each time: a reference answer computed independently before the run, graded by formula against the completed file. One entry is one data point. The log is what starts to mean something.


Most of this surfaces again at flux review. The free flux template flags which balances moved enough to explain. The CloseOps Flux & Variance System ($79) starts from your trial balance instead: confirm each account's classification and it builds the income statement and balance sheet flux statements, then ranks what to investigate.

CloseOps is productivity software and a documented working method. It is not an audit, a review, a compilation, or any other assurance service, and nothing it produces is accounting advice or an opinion on your financial statements. You remain responsible for your accounting, your judgments, and your numbers.

Related: what agents can and cannot do in the close, why AI can explain some accounts and not others, the cold-tests log, and when AI drafting helps, and when it costs more.

Found something wrong or missing? Tell me here — anonymous, thirty seconds.