Month·End·Close
← Month-End Close

Flux & Variance

Four of the five tests your flux commentary fails are arithmetic

We argue about whether commentary is good enough because it arrives as prose, and prose cannot be measured. Change the container and four fifths of the problem stops being a matter of taste.

There is a specific hour I would like back.

It is the one where a flux package lands, and I read eighteen paragraphs of commentary, and I form an impression. Some of it is fine. Some of it is not, and I send it back, and the note I write is some version of needs more detail — which is a polite way of saying I could not tell you exactly what is wrong, only that I would not sign it.

That hour repeats every month, and it is treated as the irreducible cost of reviewing someone else's judgment. It mostly is not. Most of what makes commentary bad is measurable, and the reason we argue about it instead of measuring it is that we insist on receiving it in the one format that cannot be measured: prose.

Variance commentary has five properties, not one

When a reviewer says commentary is not good enough, they are reacting to a failure in one of five specific things. It is worth naming them, because once they are named it becomes obvious that they are not all the same kind of problem.

Good commentary is quantified — the stated causes account for the movement. It is directional — the causes push the balance the way it actually went. It is sourced — a reviewer can go and look. It is specific — it names an event rather than gesturing at a category. And it is causal — it explains why the number moved rather than restating that it moved.

Read that list again and notice something. Four of those five are arithmetic or string matching. Exactly one requires judgment.

We treat all five as judgment because of the container. You cannot test whether a paragraph's causes sum to the variance, because a paragraph does not have addends.

The change: driver lines, not paragraphs

Instead of a comment box, the preparer fills in rows. One row per driver, three fields each: what happened, how much of the variance it accounts for, and where a reviewer can go to see it.

So instead of:

Cash decreased significantly during July driven by working capital movements, capital expenditure and scheduled debt service, partially offset by non-cash addbacks.

You get:

DriverAmountSource
Finished goods build ahead of the Q3 production ramp(742,500)Inventory rollforward IR-2026-07
Trade receivable growth net of July collections(785,450)AR aging 2026-07
Trade payable increase funding part of the purchase build518,400AP aging 2026-07
Scheduled term loan principal payment made 7/01(100,000)Loan amortization schedule TL-2024-A
Net loss of $739,800 less non-cash D&A and bad debt provision(553,800)TB IS section; FA register

The paragraph and the table say roughly the same thing. The difference is that the table can be checked by a formula and the paragraph cannot.

Preparers dislike this for about a month. It is more typing, and it removes the place where a vague sentence used to be able to hide. Both of those are the point.

Test one: does it add up

Sum the driver amounts. Subtract them from the variance. What is left over is the residual — the part of the movement nobody has explained — and it is the single most useful number in commentary review.

In that cash example the variance is $(1,864,070) and the drivers sum to $(1,663,350). The residual is $(200,720) — money that left the account with nothing attached to it.

Now compare a different account from the same close. Trade payables moved $(518,400). The commentary reads, in full: higher raw material purchases from the two new tooling vendors. One driver, $(178,000).

Residual: $(340,400).

I want to be precise about what just happened, because it is the whole argument. Nobody exercised judgment. Nobody debated whether the explanation was thorough enough. Two thirds of a half-million-dollar movement is unexplained, and that is a fact about arithmetic, available in the time it takes a cell to recalculate, and it is not open to disagreement.

Some tolerance is needed, because drivers rarely sum to exactly the movement and should not have to. The obvious move is a percentage — I used 85% to 115% for a while — and it is the wrong one. A percentage scales with the account. Eighty-nine per cent of that cash movement sounds like most of it and leaves $200,720 sitting there. The same 89% on a $60,000 account leaves $6,600 and means nothing was wrong. One number cannot be right for both, because the thing I actually care about is not what share of the movement was explained. It is how many dollars were not.

So the tolerance is a dollar amount, and it comes from the same place every other threshold in the file comes from: materiality, reduced. If a movement is worth investigating above $110,000, an unexplained remainder is worth investigating somewhere below that. Pick the number deliberately, write down why, and let it fall out of the materiality you already set rather than being typed in on its own.

Testing the residual also makes the test two-sided for free. Commentary that explains 140% of a movement leaves a residual of the opposite sign, and a residual is a residual whichever way it points. Someone has double-counted a driver or is describing a different account.

Test two: does it point the right way

Compare the sign of the driver total to the sign of the variance.

Accrued legal fees in that same close moved $(65,000) — the accrual grew. The commentary reads: release of the accrual on settlement of the vendor dispute, $65,000.

A release reduces an accrual. The accrual increased. The explanation is not merely thin, it is describing the opposite of what happened.

This is the failure I find most interesting, because it is the one a careful human reviewer misses most often. The sentence is well written. It names a real event, with a real cause, in specific language. It reads like good commentary. It passes every impression-based test I would have applied at 6pm on day four of close. And a sign comparison catches it instantly, every time, without getting tired.

Test three: can I get to the support

Every driver row needs a source reference, and the test is whether every row has one — not whether it is a good one.

That sounds weak. It is not, for a reason that only shows up in practice: a preparer who has to name a source for each driver separately cannot attach one document to a paragraph and call the whole thing supported. The granularity does the work. In the right-of-use asset example from the same close, one of two drivers carried a source and the other did not — and it was the larger one, the $487,200 lease recognition, that had nothing behind it.

Test four: is it specific

A banned-phrase list, matched case-insensitively. Mine has fifteen entries. Due to timing. Normal fluctuation. Various. Immaterial. Per management. Increase in activity. As expected. See prior month. Routine.

This is the weakest of the four and I want to say so rather than oversell it. It is a screen, not a verdict. It catches lazy language, and lazy language correlates with lazy work, but the correlation is not the thing itself. A determined preparer can write something empty without using any of the fifteen.

Its real value is different: it makes the standard explicit in advance. Nobody has to be told their writing is vague. The list was published before the close started.

It earns its place anyway. Finished goods inventory moved $742,500 that month. The commentary: increase in activity at the finished goods warehouse ahead of Q3. Residual nil, direction correct, source cited — three tests passed. It still fails, on the word activity, and it should, because "activity increased" is a restatement of the variance wearing a business-sounding hat.

Test five: the one you still have to make

Does the description name a business event, or does it restate the number?

Inventory increased $742,500 is the variance, not a cause. Higher spend in the period names a category, not an event. Two XR-400 units shipped 7/22 and 7/26 under the Q3 backlog release names what happened, when, and to what — and a reviewer can go and look.

No formula does this. It requires knowing what the business does, and it will still require that in ten years.

But here is what changes. When four of five tests are already answered by the time you open the file, the entire review budget goes to the fifth. You are not spending forty minutes deciding whether the numbers hang together and then running out of attention before the part that needed you. You arrive at the one judgment with your judgment intact.

The thing the tests do not catch

I would rather this piece was useful than tidy, so here is the limitation that matters most.

None of these four tests knows what is important. They know what is large.

In that same close there is an account that moved $6,400. Vendor deposits — a debit-balance asset — sitting at a credit balance. Commentary: two vendor rebate credits applied against the deposit account instead of accounts payable. Residual nil, direction correct, sourced, specific, causal. Five for five.

It was escalated, and a $506,000 treasury sweep in the same file was accepted without a second look.

A credit balance in a deposit asset is a misposting whatever its size. The sweep was a well-documented movement of money between two accounts the company controls. Materiality ranked them in exactly the wrong order, because materiality measures size and the thing that mattered was structure.

So the four arithmetic tests grade whether an explanation is any good. They say nothing about whether the account deserved explaining in the first place. That is a separate problem, it is also mostly computable, and it is not this article.

Where this leaves the review

The change is not that review gets automated. It is that review stops being the place where completeness gets checked at all.

Completeness, direction, sourcing and specificity are properties of the submission. They should be settled before a workpaper is handed over, the same way a reconciliation ties before anyone looks at it. Nobody sends a rec to review to find out whether it balances. The version of this problem when two systems compute the same figure and don't tie is a different, harder case.

Get those four out of the review and what remains is the question worth a senior person's hour: does this explanation describe something that actually happened.

That one is still yours. It should be.


If you work this in Excel: the free flux template flags what needs explaining and shows the residual your drivers don't cover. The CloseOps Flux & Variance System ($79) starts from your trial balance instead: confirm each account's classification and it builds the income statement and balance sheet flux statements, then ranks what to investigate.

The four computed tests are running here — it opens with the six examples above already graded, so you can see what each failure looks like before you try it on your own accounts. Nothing you enter leaves your browser.

Related: the wider picture in month-end flux, start to finish. Also which accounts deserved explaining in the first place, and what to do when a late entry moves an account after commentary is already written. Also why flux commentary gets sent back.

This describes a working method, not accounting advice, and nothing here is an assurance service. Figures throughout are from a synthetic trial balance built for the purpose; no company data appears in this piece.

Found something wrong or missing? Tell me here — anonymous, thirty seconds.