Industry & Business MIT Technology Review (AI)

Making AI an asset, not an expense

AI economicsenterprise AIinference costsDeloitte report

When customers discuss AI costs, the conversation usually starts with token prices and access to the latest, most capable cloud model. But they may not always need that level of capability. As AI moves from experimentation to production, model choice is only part of the equation: steady, business-critical demand can turn a consumption-only approach into a variable monthly line item that is difficult to forecast as usage, workloads, and model requirements change. The question is no longer simply which model to consume or which provider offers the lowest token price, but how to run AI economically, predictably, and at sustained scale.

AI is moving from isolated pilots into production portfolios: assistants, retrieval-and-knowledge systems, and agentic applications. Customer-service, IT, research, and business-process agents can execute multi-step workflows across enterprise systems, creating recurring demand across models, data, and tools. This is already starting to happen. Deloitte’s 2026 State of AI in the Enterprise reports that worker access to AI rose 5% in 2025, and the share of companies with at least 40% of their AI projects in production is expected to double within six months.

When AI becomes a portfolio of always-on workloads rather than a collection of experiments, the economics change. Consumption pricing gives teams flexibility and limits commitment, but when usage becomes steady, predictable, and large enough to keep capacity productive, leaders need to ask whether it still makes economic sense to buy AI one request at a time or to invest in capacity they can optimize and control. This is not an abstract cloud-versus-on-premises debate; it is a workload-by-workload business decision. Over the next 12 to 18 months, companies must assess how much AI demand they can reasonably expect and how consistently that capacity will be used. When multiple workloads share infrastructure, the enterprise can spread fixed costs across more productive use, improving the economics of ownership. The question is how much you run.

Ownership is not automatically the lower-cost answer. It only makes sense when an enterprise can keep capacity productive. Every organization has a crossover point, the level of sustained use at which owning capacity can become more economical than buying it one request at a time. There is no universal number: it depends on the models being used, the balance of input and output tokens, performance requirements, system design, energy costs, and the operating model required to support it. A retrieval-heavy knowledge system can have a very different cost profile from a simple assistant because it may process far more context for every interaction. Agentic workflows can be different again, since a single business task may involve repeated reasoning, retrieval, model calls, and tool use.

Read original →

← Back to home