A new non-executive director asks a reasonable question at her second board meeting: why did we approve last year's price rise? The chair remembers the meeting. The CFO remembers the spreadsheet. The analysis that settled it lived in a long conversation with an AI assistant, in an account that belonged to a manager who has since left.

Nobody did anything wrong. The company simply kept its thinking in a place it did not own.

That has been tolerable while every team used the same one or two hosted AI tools. It will not stay tolerable as the cost of AI changes, and that cost is moving in a direction most boards have not planned for.

Tokens are getting cheaper. The bill is not.

The price of a token for a given level of capability keeps falling. Researchers at MIT FutureTech found that the price of reaching a fixed level of benchmark performance has been dropping by around five to ten times a year.

The same paper found the cost of running frontier models rising by between three and eighteen times a year, because the newest models are larger and reason at greater length before they answer. The BenchLM index, which tracks published API prices, puts its frontier tier 72 per cent higher than a year earlier, while its mid-tier fell by 44.5 per cent over the same year.

Agents make the gap wider. Gartner estimates that an agentic task uses five to thirty times more tokens than a standard chatbot exchange, and expects inference costs per agentic workflow to rise more than fivefold through 2028. Its analyst's advice to product leaders is plain: do not count on better token economics to offset AI costs.

So the cheaper tokens arrive, and the bill still goes up, because the work an organisation asks of AI grows faster than the price falls.

Where that pushes companies

The sensible response is the one many finance directors are already making: send the frontier models only the work that needs them, and run the routine volume on smaller models, often open-weight ones hosted privately or on the company's own hardware. Gartner calls the result "complex multimodel ecosystems". In practice it means that within a year or two, the AI your company uses will be several models at once, changing every few months as prices and capabilities move.

That is a reasonable cost decision with an awkward consequence.

A small local model knows nothing about your company. It has not read your board papers, it was not in the room for last year's pricing decision, and its working memory is short. Whatever it needs to know, something has to hand it. The same is true of the frontier model you rent for the hard questions, and of whichever model replaces both next spring.

Memory becomes the part that has to stay put

When the model changes every quarter, the organisation's memory cannot live inside any one of them, or inside one vendor's chat history. It has to be somewhere the company owns, in a form any model can read.

That is why organisational memory moves from a nice idea to table stakes. It is the one part of the system that has to outlast the model.

It also has to be the right kind of memory. A pile of every transcript and every agent's output is not a memory; it is a larger version of the problem. For a board, a useful memory has four properties:

  • A person has confirmed it. An agent's draft is a draft until someone who did not write it says it is right.
  • It points to its source. An answer about last year's price rise should come back with the paper and the minute it rests on.
  • The organisation holds it, not a model's training data and not a former employee's account.
  • Any model can read it, through an open standard such as the Model Context Protocol, so changing models does not mean rebuilding what you know.

What a board can ask now

Three questions are worth a few minutes at the next meeting.

Where does the reasoning behind our last three major decisions live, and who could retrieve it if the person who did the work left tomorrow?

If we changed AI provider next quarter, what would we lose?

When an agent adds something to our record, who confirms it before we rely on it?

If the answers are uncomfortable, the fix does not need a new AI tool. It needs a place for the memory, and a rule that a person confirms what goes in.

This is the problem we built Boardverse to solve: your team keeps the AI it already uses, and the confirmed record stays with the company, readable by whichever model you choose next.

Sources: Gundlach, Lynch, Mertens and Thompson, MIT FutureTech, The Price of Progress. BenchLM Token Price Index, October 2026. Gartner, 25 March 2026 and 17 August 2026.