If someone has to review the AI's work, did it actually save any time?
The short answer
Sometimes yes, sometimes no, and the difference is predictable. AI saves time when checking the output costs less than producing it. It saves nothing when a wrong answer looks exactly like a right one, because then verification costs full production. In accounting, that line runs between mechanical work and judgment work.
The honest test for AI in finance is not speed, it is what verification costs.
By Stephen Ninesling, FynScale

If someone has to review the AI's work, did it actually save any time?
Short answer. Sometimes yes, sometimes no, and the difference is predictable. AI saves time when checking the output costs less than producing it. It saves nothing when a wrong answer looks exactly like a right one, because then verification costs full production. In accounting, that line runs between mechanical work and judgment work.
Why is this the right question to ask?
Because it is the only honest test, and almost nobody applies it.
The pitch for AI in accounting is usually framed as a comparison between doing the task and not doing the task. That is not the actual choice. The actual choice is between doing the task and reviewing a machine's version of the task. Those are different activities with different costs, and the second is only worth it when it is genuinely cheaper than the first.
Once you frame it that way, the question stops being philosophical and becomes arithmetic. If a task takes 60 minutes and reviewing an AI draft takes 15, you saved 45 minutes. If reviewing takes 55 minutes because you have to check every line before you are willing to sign your name to it, you saved five minutes and added a dependency.
The trap is that both scenarios feel identical in the moment. Both produce finished-looking output quickly. The difference only appears in whether the reviewer can trust a scan or has to audit, and that distinction is invisible until something goes wrong.
What makes verification cheap or expensive?
One factor dominates everything else: whether a wrong answer looks wrong.
Consider matching 400 bank transactions to open invoices. If the system matches a $2,340 payment to a $2,340 invoice from Coastal Supply, a reviewer confirms it in about a second. If it matches that payment to an invoice from a different customer, the name does not fit and it stands out immediately. Errors announce themselves without anyone recomputing anything. Verification is cheap, and the time saved is real and large.
Now consider a revenue recognition treatment for a contract with variable consideration and three performance obligations. The AI produces a schedule. It is internally consistent, correctly formatted, and confidently reasoned. To know whether it is right, you have to read the contract and work the treatment yourself. Verification costs full production. There is no saving, and there is now a second-order risk: the polish of the output makes the review shallower than it would have been on a junior's messy draft.
The general rule is that verification cost scales with how much of the original work you must redo to be confident. Where confidence comes from pattern recognition, AI helps enormously. Where confidence requires reasoning through the problem from the source documents, it helps very little, and it may actively hurt.
What does this look like in an actual month-end close?
Take a close that runs 40 hours across a small team. Split it by verification cost rather than by task type, because task type is not what determines the answer.
Cheap to verify, roughly 24 of the 40 hours. Transaction categorization. Bank and credit card matching. Intercompany tie-outs. Prepaid and accrual schedule updates where the pattern is already established. Variance flagging against prior period. Draft commentary for recurring line items. Document collection, naming and filing. Reconciliation of clearing accounts where the pairing is mechanical.
Applying AI to this bucket and reviewing the output typically compresses it to somewhere between 8 and 12 hours. Not because the machine is clever, but because a human confirming a match is doing something fundamentally faster than a human finding a match. The cognitive work is recognition rather than search.
Expensive to verify, roughly 16 hours. Anything being treated for the first time. Judgment on whether an accrual is required and at what amount. Deciding what a variance means as opposed to how large it is. Anything a lender, buyer or auditor will read closely. The conversation with the owner about what the numbers imply for the decision in front of them.
Applying AI here saves close to nothing, because the reviewer ends up reasoning through the problem regardless. What it can still do is prepare the ground: assemble the relevant contracts, surface how comparable situations were treated before, lay out the options with their consequences. The judgment stays human, but the human arrives at it with the material already gathered rather than spending two hours gathering it.
Net effect on a 40 hour close is something like 24 to 28 hours. That is a real gain, worth real money, and it is a completely different claim from the ninety percent reductions that get advertised.
Does confident wrongness make review harder?
Yes, and this deserves far more attention than it gets.
A junior accountant who is unsure produces work that signals uncertainty. Blank cells. A question in the margin. A note saying they were not sure how to handle the third quarter and left it for you. Those signals are enormously valuable to a reviewer, because they point directly at where to look.
AI output does not signal uncertainty the same way. A wrong schedule and a right schedule look equally finished, equally formatted, equally confident. This removes the reviewer's single most efficient tool, which was knowing where the soft spots were before starting.
The practical consequence is that AI-assisted work requires a different review posture. Instead of scanning for signals of doubt, the reviewer has to decide in advance which parts of the output carry risk and check those regardless of how good the work looks. That is a discipline, and it has to be deliberate, because the natural human response to polished work is to review it less carefully rather than more.
Teams that get this wrong develop a specific and predictable failure pattern. The first several outputs are correct. Trust builds. Review gets lighter. The error that eventually arrives passes through untouched, because by then nobody was really looking. The error rate never changed. The review did.
What does an error actually cost?
This is the question that should set review intensity, and it almost never does.
A miscategorized $40 software subscription is a cleanup item. Someone finds it next quarter, moves it, and nothing else happens. Review it lightly, or not at all, and accept the occasional miss.
A misapplied revenue treatment on a contract is a restatement. It reaches the tax return, the covenant calculation, and any diligence process that happens in the next three years. The cost is not the correction, it is the credibility damage at the exact moment credibility matters most.
Between those extremes sits most of the work, and the useful habit is to ask what happens if this specific output is wrong and nobody catches it for six months. That question sorts tasks faster than any framework, and it produces a review plan that matches consequence rather than matching how the output looks on the screen.
So does the review requirement ever go away?
Not for anything carrying accountability, and this is worth stating plainly rather than treating as a limitation to be engineered away eventually.
When a lender relies on a covenant calculation, someone is representing that it is correct. When a board reads a forecast, someone is standing behind the assumptions. When a return is filed, a person signs it. That representation is not a formality left over from a manual era. It is the product.
Framing the review as a temporary inconvenience misunderstands what a client is actually buying. Nobody hires a finance function to receive numbers. They hire it to receive numbers that someone will defend when questioned.
What changes with better tools is not whether review happens but what the reviewer spends attention on. When the mechanical middle is drafted rather than produced from scratch, the reviewer's time shifts from assembly to judgment. That is a much better use of an experienced person, and it is the actual promise. But it is a redistribution of effort rather than an elimination of it, and firms that sell it as elimination are setting up a disappointment they will have to manage later.
How should a firm decide where to apply it?
Three questions, in this order, before automating anything.
Would a wrong answer be obvious? If yes, this is a strong candidate and you should move aggressively. If a wrong answer would look exactly like a right one, proceed with much more care and much heavier review.
What does an undetected error cost? Match review intensity to consequence, not to how finished the output looks. This is the question that prevents both over-reviewing trivial work and under-reviewing dangerous work.
Does the task depend on context the model cannot have? A great deal of accounting judgment rests on things nobody ever wrote down. Why this client always accrues that particular expense. What the lender actually cares about versus what the covenant technically says. Which of the founder's numbers are real and which are aspirational. Work that depends on undocumented context will produce confident output built on a foundation that is simply missing.
Tasks clearing all three should be automated hard. Tasks failing any of them stay human, with AI used to prepare rather than to decide.
What is the honest version of the pitch?
AI does not remove the accountant. It changes what the accountant spends the day doing.
The mechanical middle of the work compresses substantially. Matching, categorizing, drafting, assembling, checking arithmetic, chasing documents. This is genuinely most of the hours in most finance functions, and compressing it is worth real money to anyone paying for those hours.
The judgment does not compress at all. If anything it expands, because when assembly stops consuming the entire day there is finally attention available for the questions that actually change outcomes.
What that produces is not a cheaper version of the same service. It is a different mix: less time producing the number, more time on what the number means. For a business owner, that second part is what was always missing, because the traditional arrangement spent every available hour on production and left nothing for interpretation.
The review never goes away. It is not overhead sitting on top of the value. It is the value.
Common questions
Is there a task where AI review is genuinely near zero? Close to it for high-volume matching where both sides carry an identifier. Bank feed to invoice, payment to remittance. A reviewer spot-checks a sample rather than every line, because errors in that work are visible on sight.
How do I know if my team is under-reviewing? Track where errors are caught. If most are found downstream, by a client, a lender or an auditor rather than internally, the review is happening too late or too lightly.
Does this mean AI is not worth it for small finance teams? The opposite. Small teams spend the highest proportion of their hours on mechanical work, because they lack the scale to have specialists. That is exactly the bucket that compresses.
What about using AI to review AI? It helps for consistency checks, arithmetic, and formatting. It does not help for the judgment calls, because a second model has the same missing context as the first and will often agree confidently with the same wrong answer.
Should clients be told when AI is used in their work? Yes, and it is easier than most firms expect. Clients care that someone is accountable for the number, not which tools produced the draft. Being direct about it early prevents an awkward discovery later.
Does this change what a finance hire should look like? It raises the value of judgment and lowers the value of throughput. The person who can decide what a variance means is worth more than the person who can produce twenty schedules quickly, and that gap widens every year.
How do I test this in my own business? Pick one recurring task. Time it done manually. Time it done with AI plus the review you would genuinely be comfortable signing. Compare the second number, not the first, to what you have now.
FynScale is a boutique AI consulting firm for accounting and finance. AI speed. Human judgment.