Measuring AI Value: The Economics of New Capability

In 1984, Eliyahu Goldratt published a novel about a factory manager with three months to save his plant. The book sold in millions, and its argument has outlived its fiction. A system's output is governed by its constraint. Improving anything that is not the constraint produces activity, cost and a great deal of local satisfaction, and leaves the output where it was.

Goldratt went further and attacked the accounting. Traditional cost measures rewarded local efficiency, so a manager who ran a non-bottleneck machine at full capacity looked productive while filling the floor with work in progress that the bottleneck could not absorb. He proposed measuring the system instead: what it actually delivers, what is tied up inside it, and what it costs to run.

Four decades later, the same argument is being had about artificial intelligence, mostly by people who have not read him. It is worth borrowing his discipline, because the mistakes are identical and the stakes are larger.

Three ways of measuring, and what each is for

The first is expenditure. Tokens consumed, seats licensed, infrastructure billed. Every conversation about AI starts here because every invoice does, and the figures have become harder to wave away as agentic workloads consume considerably more compute than the chat interfaces most budgets were written for.

Expenditure deserves attention, and it tells you nothing about value. A cap applied without a value measure falls on the best deployment and the worst at the same rate, which is a tax rather than a discipline. Goldratt would have recognised it immediately as an operating expense measure being asked to do a job it cannot do.

The second is unit cost. What does it cost to resolve a claim, review a contract, produce a quotation? This is a real advance, because it pairs spend with something the business already values. A claim that costs six units of effort to handle manually and one through a well-built system is a claim you can reason about. The number that matters is the cost of the resolved case, not the compute consumed reaching it.

Unit cost has a ceiling, though, and the ceiling is conceptual. It describes a world where the work stays as it is and only the price of performing it moves. The system was designed around human capacity, and it remains designed around human capacity, now with a cheaper participant inside it.

The third is rate. How many cycles does the organisation complete in a period? This is the measure closest to Goldratt's own, and it is the one most enterprises are not yet using. Shaving each step of a four-month process yields a three-and-a-half-month process. Rebuilding the process so it runs end to end at machine speed changes what the organisation is capable of attempting. Requests that were once too small to justify a team become worth doing. Opportunities that used to expire during the planning cycle survive it.

Rate is the right frame, and we would encourage more organisations to adopt it. What we want to add is a fourth question, and it comes from the work rather than from theory.

The work nobody performs today

Expenditure, unit cost and rate share an assumption so basic that it usually goes unstated: the work is currently being done. There is a cost to reduce, a unit to price, a cycle to accelerate.

A good deal of what enterprises now need has no such history.

Consider a manufacturer asked for assured water consumption at a named plant. The plant is metered, as are the eleven others. Readings sit in gate logbooks, a utilities spreadsheet, the maintenance system, and tanker notes filed by accounts. One meter was replaced in March, and its predecessor's register was retired. Producing a single defensible figure means knowing which meter covers what, which readings can be relied on, and who can confirm it. Nobody has ever done that. There is no baseline because the exercise has never been performed.

Or consider a group whose managers present cost savings in spreadsheets each quarter, against a gross margin target that does not move. Every claim is plausible. None has ever been tested against the ledger, because testing it would take a person weeks and no such person exists in the structure. The audit function is absent rather than slow.

In both cases the question is not what the work costs, how much each unit costs, or how many cycles are completed. The work has never happened. What a system does here is not accelerate anything. It brings a capability into existence.

We have written separately about why those answers are so difficult to assemble, and the short version is that the evidence exists while the interpretation of it lives in people. The economic point is narrower. This category of value is invisible to all three conventional measures, and it is growing fastest, because the obligations arriving from regulators, customers and auditors are precisely obligations to account for things nobody previously had to account for.

Why this is the hardest value to sell

There is a commercial consequence, and it is worth stating plainly because it costs firms like ours business.

Procurement asks what a process costs today. For work nobody performs, the honest answer is nothing. No hours are saved, because no hours are spent. A buyer applying an efficiency yardstick to a capability that did not previously exist will conclude the case is weak, and by the terms of that yardstick they are correct.

The unit here is different. It is exposure closed, and margin recovered, measured against a target the business has already set for itself. A manufacturer chasing twenty per cent gross margin who cannot locate where the leakage sits does not need hours back. He needs the leak found. Establishing that measure at the outset, before an efficiency benchmark from an unrelated project gets picked up and applied, is one of the more useful things either party can do in a first conversation.

Why the capability compounds

Rate, as usually described, sounds like a new plateau. The organisation moves from four cycles a year to forty and settles there.

Our experience suggests it does not settle, for a reason that has more to do with what accumulates than with speed.

Assembling those answers properly the first time produces something the organisation has never held: a map. Which meter covers which line. How the estimator substitutes and on what basis. Which requirements bind which site. Where the evidence actually lives. That map was built to satisfy an external party, and its more durable value is internal. The second cycle does not start from a blank sheet. The tenth starts from a considerably better one.

In Goldratt's terms, this is the part that matters. A faster process runs the existing system harder. An accumulating record raises the constraint itself, because the scarce resource was never the hours. It was the small number of people who knew how the operation actually worked, and their knowledge is now partly outside their heads.

The caution, which we learned by getting it wrong

New capacity has to be received by an organisation ready to use it, and this is where the theory meets an ordinary Tuesday.

In one engagement, a quotation cycle that had taken around three days was compressed to minutes. The system performed exactly as specified. The commercial processes around it did not move, because the follow-up cadence, the pipeline assumptions and the expectations about what a faster quotation would achieve had all been built for a three-day world. Speed turned out to be an input. Converting it required redesigning the rhythm around the new tempo and deliberately resetting what the system was expected to change on its own.

The general lesson is structural, and we now treat it as part of scope. When a workflow accelerates by an order of magnitude, the workflows next to it were designed around a constraint that has just been removed, and they need re-examining. Committees, stage gates and sign-off chains were sensible when execution was expensive. As execution becomes cheap, much of the delay an organisation attributes to technology turns out to belong to its own operating model, and that is a redesign question rather than a procurement one.

Where we stand

Work that has never been performed cannot be priced on efficiency. The unit has to be exposure closed, and margin recovered. And we have seen from experience that far less operational knowledge is genuinely irreducible than most organisations assume. A good deal of what gets called tacit is simply unasked, because no mechanism existed that would have used the answer. Where the asking is done well, a great deal comes out.

In closing

Goldratt's contribution was to insist that measuring the parts tells you very little about the system. The same holds now. Expenditure is hygiene. Unit cost is progress. Rate of execution is the measure most organisations should be adopting and are not.

The question we would add is simply asked. Of the things your organisation is now expected to know about itself, how many can you currently establish, and how many have you never once produced? For the second group, there is no efficiency to gain, because there is nothing to make more efficient. There is only a capability to build, and an organisation that can answer for itself is worth considerably more than one that merely runs faster.

Mitochondria is an agentic AI product company based in Amsterdam and Pune. ISO 27001:2022 certified. Designed to fall within the limited and minimal risk tiers of the EU AI Act, with controls aligned to the GDPR, UK GDPR and India's DPDP Act.

Next
Next

When Systems Cannot Produce the Answer