← All articles
Economics5 min read

AI Can Create Value. The Return Still Has to Be Proved.

Imagine two agents handling the same customer query. One costs 5p to run but needs three attempts and ten minutes of checking. The other costs 50p and produces an acceptable answer with one minute of review. At an illustrative £30 an hour for human time, their costs are £5.15 and £1 respectively, before shared overheads.

The cheaper model produced the more expensive task.

These are hypothetical figures, but the accounting question is real. It also helps explain why the debate about whether AI is a bubble so often talks past the decision a business actually faces.

The Medium article about Jon Gray’s Blackstone presentation makes a forceful case that AI is already delivering economic value. That deserves attention. It does not settle whether particular investments, business models or transactions will earn a return.

In Blackstone’s original presentation, Gray reports that annualised spending with Anthropic across roughly 1,400 portfolio companies, GP-stakes portfolio companies and borrowers rose from $25 million to $525 million. That supports a demand story. The Medium author’s inference that businesses would not spend without returns goes further: expenditure can reflect successful deployment, experimentation or competitive pressure.

Gray also reports concrete operating benefits: Phoenix Tower processing leases five times faster, and Enverus achieving an 18-fold return on spending with model companies. Those deserve investigation. They remain selected management-reported examples, without enough disclosed methodology to reproduce a fully loaded return calculation.

There is a material transcription error, too. Blackstone’s revenue chart places the combined $105 billion annualised run-rate for Anthropic and OpenAI in July 2026, not July 2025 as the article states. A run-rate is also different from revenue already earned over a full year. Original chart, 3:58

The broader demand signal is supported elsewhere. Google’s own keynote reports monthly token processing rising from roughly 480 trillion in May 2025 to more than 3.2 quadrillion in May 2026. That measures activity across its services. It does not measure customers’ profits or the value of each token.

We need to keep four measures separate. Infrastructure spending buys capacity. Model evaluations test capability. Customer revenue records sales. Customer profit reflects what remains after costs. Progress in one can support another; it cannot substitute for it.

For an organisation buying automated work, the useful unit is a completed, dependable task: the right outcome, within the agreed authority, at an acceptable quality and time. A plausible answer or successful API response is only an intermediate step.

Follow that task from request to acceptance and its costs become clearer. Model charges include input, output and repeated context. Cached tokens can reduce the bill, but cache creation, expiry and reuse matter. Anthropic’s pricing documentation distinguishes cache writes from reads and explains separate charges for some tools. A token count without its billing categories is an incomplete cost record.

Then come searches, browser sessions, code execution, databases and other tool calls. Retries repeat some of those costs. Human review, correction and escalation add time. Failed work still consumed resources, even when it produced nothing acceptable. Recovery from an incorrect external action may cost more than generating the original answer.

Local execution changes who receives the invoice. It does not make compute free. Electricity, hardware depreciation, maintenance and the capacity held available for peak demand all belong somewhere in the calculation. A machine waiting overnight is still capacity somebody financed. Allocate shared costs consistently and avoid counting electricity twice when a hosted price already includes it.

I would measure total attributable cost across the workload divided by accepted outcomes, alongside completion rate, turnaround time and quality. Otherwise a system can appear cheaper simply by abandoning difficult cases. Compare equivalent work, including what remains unfinished.

This is one question we have been exploring through Agent Control’s experimental Digital Labour Exchange. A work order specifies the outcome, scope, verifier and budget reservation. Eligible workers submit offers; routing considers measured performance and resource estimates. The transaction record connects the worker, execution backend, model, tools, attempts and verification result.

That provides a basis for comparing local resources with paid providers once their respective costs and quality are measured. An estimate helps choose a route; a budget limits exposure; the retained record shows what happened. None alone proves a saving.

Our actual qualification was narrower. Three local Qwen2.5 workers handled twelve held-out synthetic tasks following calibration. Brokered routing and the cheapest-estimate strategy verified the same nine outcomes. Including failed attempts, the broker used 13.4% fewer measured CPU accounting units in that single run. On two matched tasks, retries made the cheapest route consume more CPU than the fixed larger-model route.

This was our own bounded experiment, not independent market evidence. There were no paid API calls or financial settlements. Electricity was not measured; controller overhead was incompletely metered; other host activity could introduce noise. The conclusion on monetary savings remained inconclusive. The useful result was a comparison that retained failures rather than hiding them.

The strongest objection to the wider investment story therefore survives: useful AI and growing demand can coexist with overbuilding, subsidised prices, weak margins and disappointing returns.

There is primary evidence for that tension. Microsoft’s 2025 annual report says scaling AI infrastructure reduced its Intelligent Cloud gross-margin percentage, partly offset by efficiency gains. Revenue growth and margin pressure can happen together. The report also discloses substantial datacentre lease commitments; outsourcing infrastructure does not make its obligations disappear.

Financing relationships deserve scrutiny too. Anthropic’s Amazon agreement combines an investment from Amazon with a commitment exceeding $100 billion of AWS spending over ten years. That does not establish artificial demand. It does mean investment flows, supplier revenue and independent customer demand must be distinguished.

Better efficiency may keep prices falling. Investors may also demand higher margins, discounts may end, and hardware may need replacement sooner than expected. Buyers should test their economics under those alternatives. Investors need to examine utilisation, contract durability, replacement costs and the price paid for the asset.

Finally, hours released are not automatically cash saved. They may become better service, additional output, shorter working hours or unused capacity. The benefit needs an owner and an observable destination.

AI can create substantial value while particular investments fail. The concrete test for buyers and investors is the same: What does one accepted outcome cost, who captures the saving, and does that saving persist without subsidy?

Also published on Medium ↗.

This article reflects the evidence and development status at its original publication date.