Blog

What does AI actually cost to run inside a finance function?

The short answer

It is metered, not licensed, which is the part that catches people. There is no fixed seat count. Cost rises with enthusiasm, document size, and how many times a tool retries. The bill arrives about thirty days later with no explanation of what drove it, and by then the behavior that caused it is a month old.

It is a meter, not a subscription, and the bill explains nothing when it arrives.

By Stephen Ninesling, FynScale

Why does the seat license model break here?

Because every piece of software finance has bought for twenty years was priced per person per month.

You know the number before the month starts. Twelve seats at $40 is $480, and it is $480 whether the team works flat out or takes a week off. Budgeting is arithmetic and the variance is zero.

AI tools are priced by consumption. You are billed for the volume of text going in and coming out, which means the same tool costs different amounts every single month depending on what people did with it.

That distinction sounds academic until the first surprise invoice. A team that spent $340 in June spends $2,800 in July because one person discovered they could run the tool across a full year of documents rather than a single month. Nobody did anything wrong. Nobody broke a rule, because there was no rule.

The mental model has to move from a subscription to a utility. You would not be surprised that a factory's power bill moves with production. This is the same relationship, and almost nobody has set it up to be visible.

What actually drives the cost?

Five things, and only one of them is obvious.

How much text goes in. Feeding a tool a single invoice is cheap. Feeding it a full general ledger export, a policy document, and eighteen months of statements so it has context is not. Cost scales with the volume of material, and long context is where the money goes.

How much comes back. A one line answer is cheap. A full drafted memo is not.

Retries. When the first answer is wrong and someone asks again with better instructions, that is a second charge. Nobody counts these, and in practice a difficult task can involve four or five attempts.

Agent loops. Tools that work autonomously, reading, checking, and revising before producing a result, may run many cycles for a single request. From the outside it looks like one action. On the bill it is not.

Which model. The more capable models cost several times more per unit than the lighter ones. Most tools default to the expensive one, and most users never look at the setting.

The uncomfortable pattern in all five is that cost correlates with effort and thoroughness. The most engaged person on your team is your most expensive one, and nothing in the interface tells them so.

What does this look like in practice?

A twelve person finance team running a close automation tool and general document work. Baseline consumption sat around $340 a month for the first quarter, which nobody questioned because it was small.

Then two things happened in the same month.

Someone ran a bulk extraction across three years of vendor contracts to build a commitments schedule. Legitimate work, genuinely useful, and it consumed more in one afternoon than the team had used in the previous six weeks.

Separately, the tool was pointed at the close and left to run through the reconciliations autonomously. It worked. It also looped far more than anyone expected, because the underlying data was messy enough to require repeated attempts.

The month came in at $2,800. Nobody could explain it from the invoice, because the invoice showed a total and a unit count and nothing about which activity produced them.

The real problem was not the $2,800. It was that the team could not tell whether that number was a one-time event, a new baseline, or the beginning of a curve. Without that answer there is no decision to make, only anxiety.

Why is this specifically a finance problem?

Because it is a spend visibility question, and that has always been your job rather than IT's.

An unmeasured, uncapped, consumption-based cost that grows with usage is exactly the category of spend a finance function exists to control. If it were any other line, telecoms, freight, cloud hosting, nobody would accept an invoice with no breakdown and no forecast.

The reason it slips through is that it arrives labelled as a technology line and technology lines are somebody else's problem. But no CIO is going to be asked what the AI returned. That question lands on the person who signs off on the numbers.

There is also a governance gap underneath it. When something is cheap enough that nobody notices, no approval process gets built. Then it stops being cheap, and there is still no approval process, because nothing ever triggered one.

How do you instrument it without slowing anyone down?

Four things, none of which require a project.

Tag spend to a process, not to a department. Knowing that finance spent $2,800 is useless. Knowing that the close consumed $900, document extraction consumed $1,600, and ad hoc questions consumed $300 is actionable, because each of those has a different answer. Most platforms support tagging or separate keys per workflow. Set it up before you need it.

Get a per person or per team consumption number, then cap it. Not to punish anyone. Caps exist so that a surprise becomes a blocked request rather than an invoice. Set the cap generously, above the current baseline, and treat hitting it as a signal to investigate rather than as a failure.

Set an alert threshold at a fraction of monthly budget. The goal is finding out on the eighth of the month rather than the fifth of the next one. Most platforms offer this and most accounts have never turned it on.

Track a unit cost that means something. Cost per close. Cost per client served. Cost per thousand documents processed. A total tells you nothing about whether the spend is healthy. A unit cost tells you immediately whether it is improving or drifting, and it is the only version of this number a board will find useful.

What about vendor contracts?

This is where the exposure is largest and least visible.

Many AI tools are sold with a platform fee plus usage, and the usage component is frequently open-ended. The contract commits you to a floor with no ceiling.

Three things worth insisting on before signing.

A hard cap or an overage alert written into the agreement rather than left to a dashboard setting somebody can change. A commitment that pricing per unit will not rise during the term, since underlying model pricing moves and some vendors pass that straight through. And clarity on whether you are billed for the vendor's own internal retries, which is a genuine question worth asking directly.

Also ask what happens to your consumption when the vendor upgrades their underlying model. A more capable model usually costs more per unit, and a silent upgrade can raise your bill without anyone on your side making a decision.

Does the same logic apply if you build rather than buy?

Yes, and more sharply, because you own every inefficiency.

A poorly constructed process that sends the same context repeatedly will work correctly and cost several times what it should. Nothing will alert you, because the output is fine. This is the software equivalent of leaving a light on in every room, and it is invisible until someone reads the meter.

The discipline that works is treating each automated process as a line item with an expected cost, then checking actuals against it monthly. Not because the amounts are large at first, but because the ones that grow quietly are the ones that eventually matter.

What about the cost that never appears on the invoice?

Worth naming, because it is usually larger than the metered spend and it is nowhere in the budget.

Review time is the obvious one. Work that arrives drafted still needs checking, and where the checking is thorough that is real hours from experienced people. Those hours are more expensive than the tokens by a wide margin. A tool costing $340 a month that consumes eight hours of a controller's time monthly is not a $340 tool.

Rework is the second. When output is wrong in a way nobody catches immediately, the cost lands later, in an amended filing, a corrected report, or a conversation with a lender. That cost never gets attributed back to the tool that caused it, so the tool keeps looking cheap.

Then there is the setup and maintenance nobody budgets. Prompts get written, refined, and broken by a model update six months later. Someone owns that, and in a small team that someone is usually the person with the least spare capacity.

None of this argues against the tools. It argues for comparing them against the right baseline. The honest comparison is not tool cost versus zero, it is total cost including review against what the work cost before. Teams that make that comparison honestly tend to keep the tools and drop about a third of the use cases, which is the correct outcome and a much better one than discovering the same thing eighteen months later.

What does good look like?

You can answer four questions without opening an invoice.

What did AI consumption cost last month in total. How that splits across your main processes. What the cost per close or per client is, and whether it is moving. What the ceiling is if every cap were hit at once.

That last one is the number to know before a board meeting, because the honest answer to "what is our exposure" is the ceiling, not the current run rate.

None of this requires new software. It requires deciding that a consumption line gets treated like a consumption line, which is a thing finance already knows how to do and has simply not been asked to do here yet.

Common questions

Is a fixed price AI tool safer? Simpler to budget, and usually more expensive at low volume because the vendor has priced in the risk you are transferring. Worth it if predictability matters more than the absolute number, which for many small teams it does.

How much should a small finance team expect to spend? It varies too much with usage to give a single figure honestly. The more useful benchmark is your own first three months, treated as a baseline, with anything above a fifty percent jump investigated rather than absorbed.

Why is the invoice so hard to read? Most billing is expressed in units of text processed rather than in tasks completed. Nothing on the invoice maps to work anyone recognizes. Tagging by process on your side is what closes that gap, and the vendor is unlikely to do it for you.

Who should own this line? Finance, in the same way finance owns any other consumption spend. The failure mode is leaving it with whoever bought the tool, because they have no reason to look at the meter.

Does capping usage hurt productivity? Only if the cap is set too low. Set above the working baseline it functions as a smoke alarm rather than a lock, and it converts a surprise invoice into a conversation.

What is the single biggest driver people miss? Retries and agent loops. Both are invisible in the interface and both can multiply the cost of one apparent action several times over.

Should this be capitalized or expensed? Consumption spend is an operating expense in the ordinary case. Where meaningful development work is involved the treatment question is real, and it is worth asking your accountant rather than deciding by default.


FynScale is a boutique AI consulting firm for accounting and finance. AI speed. Human judgment.