Flux & Variance
Why AI can explain some accounts and not others
If you've tried using a model for variance commentary, you've probably found it excellent on a few accounts and useless on the rest. The dividing line is more specific than "complex" or "simple."
The usual diagnosis is that the prompt needs work. It normally doesn't. Whether a variance can be explained by a tool depends almost entirely on one question:
By detail, read whatever transaction-level report supports the account: GL or sub-ledger detail, a completed reconciliation, an operational report. Whichever one you'd pull anyway.
That question sorts most schedules into three groups.
Group 1: the reason is in the detail, already labeled
Cash, once the reconciliation is done. Receivables, where the story is invoices against receipts. Payables, where it's purchases against payments. Straightforward FX.
What these have in common is that the transactions describe themselves. Bank activity carries payees. Sub-ledger detail carries customers and vendors. The tool isn't determining anything — it's aggregating labeled data and phrasing the result. No judgment is being delegated, which is exactly why it works.
Build here first. These are the accounts where a repeatable skill earns its setup cost and keeps earning it every month.
Group 2: the reason is retrievable, but needs work
Inventory movements, where transactional flow sits alongside journal-driven adjustments. Contra-revenue accounts, where you're pivoting journal detail by counterparty and comparing periods. Revenue by category, where the data shows which customers moved but the why requires seasonality context that only exists in your head.
The workable split here is: let the tool compile, and write the interpretation yourself. Ranking by size, pivoting by counterparty, comparing period over period — that's mechanical and it's genuinely time-consuming by hand. The judgment stays with you.
The trap: this is where a model produces something that sounds finished. It will describe the movement fluently without naming a cause, and fluent restatement is easy to accept when you're tired. If the draft doesn't name a driver you could verify, it isn't done.
Group 3: the reason isn't in the system at all
The catch-all accrual account. Whatever your equivalent of "accrued — other" is: a dozen unrelated items, moved by journal entries whose memo lines say nothing useful about the underlying vendor, program, or decision.
Also here: a bill that doesn't indicate what it was for. A journal with a counterparty and nothing else. Answering these means asking another department, and the answer typically never makes it back into the ledger.
Don't automate these. Not because the tool is weak — because the input is absent. Prompting harder cannot retrieve detail that was never posted. Recognizing this quickly is worth more than any prompt: it's the difference between twenty minutes of re-prompting and one message to the person who knows.
The check that comes before all of this
Where an account is fed by an upstream or third-party system, there's a failure mode sitting underneath everything above: the data didn't flow in correctly and the balance doesn't reflect reality.
Nothing in the detail will tell you this, because the detail is internally consistent. It's just wrong. Tie the total to the external source before you explain anything — otherwise you're writing a careful narrative for a number that shouldn't exist. The same goes for the ledger export itself: tie it to the trial balance, account by account, before anything reads it.
What this means practically
Most people automate the account that annoys them most. That's usually the messiest one, which is exactly the one where automation can't work — and when it fails, the conclusion drawn is that AI is useless for close.
The account was the problem, not the tool. Build for Group 1 first, get it genuinely reliable, then extend outward.
One more thing worth saying: an account that lands in Group 3 every single month isn't a hard account. It's an upstream documentation gap. Memo standards, a supporting schedule maintained alongside the entry, or a sub-account split will move it permanently — which is a better return than any amount of prompt engineering.
And whatever you automate has to stay reviewable
Faster preparation means more gets built — more supporting calculations, more granularity, bigger workbooks. But the reviewer's job doesn't get faster. It gets harder: more logic to check, often structured in ways they wouldn't have written.
So: formulas rather than typed values, everything referenced, nothing hardcoded. State that explicitly every time you direct a tool. A figure nobody can trace has no value in a review process even when it happens to be right. A review where every number is computed and traceable looks like this.
If you work this in Excel: the free flux template flags what needs explaining and shows the residual your drivers don't cover. The CloseOps Flux & Variance System ($79) starts from your trial balance instead: confirm each account's classification and it builds the income statement and balance sheet flux statements, then ranks what to investigate.
Related: getting started with agents in month-end close, what this means for which account you hand over during close, when the operational report and the GL don't agree, and why the accounts you booked yourself are the fastest to explain. Also why flux commentary gets sent back, what AI agents can and cannot do in the close, and cold tests of agents against a reference answer.
Found something wrong or missing? Tell me here — anonymous, thirty seconds.