Why Circulara · The thesis

The coming age of
compute waste.

Why the AI economy needs a circular economy, and why the conversation starts with your token bill.

The month the bill arrived

Somewhere in corporate America this year, a company spent half a billion dollars on AI in a single month. Not on data centers. Not on GPUs. On usage, an AI consultant told Axios that one client ran up roughly $500 million in one month after failing to put usage limits on AI licenses for employees.

That story is an outlier in scale, but not in kind. In the first half of 2026, the pattern repeated at every size:

  • Uber watched adoption of agentic coding tools jump from 32% to 84% across its 5,000-engineer organization in three months, and burned through its entire 2026 AI budget in four. Its CTO put it plainly: "I'm back to the drawing board, because the budget I thought I would need is blown away already." Uber's response was a hard cap: $1,500 per employee per month per agentic coding tool. Its COO told analysts in May that AI costs were becoming "harder to justify."
  • Microsoft: a company that builds AI for a living, terminated most of its internal licenses for a third-party agentic coding tool after per-engineer bills reached $500-$2,000 per month, redirecting its developers to its own tooling.
  • Forrester now finds enterprises postponing a quarter of planned AI spend to 2027 as CFOs demand justification. Anecdotes from the field are almost comic: one CTO reported employees using frontier AI models to check the weather.

The first era of enterprise AI ran on a simple argument: adopt fast or fall behind permanently. Finance departments accepted it. That era is over. The CFOs are in the room now, and the questions they're asking don't have good answers.

Everyone is spending. Almost no one can show the return.

The spending is extraordinary. Gartner forecasts worldwide AI spending at roughly $2.59 trillion in 2026, up about 47% year over year. Hyperscalers alone are on track to invest $675 billion in AI infrastructure this year, up 63%. Nearly nine in ten organizations now use AI in at least one business function.

The returns are not keeping pace, and this is no longer a contrarian claim; it is the consistent finding of nearly every major study:

  • PwC's 2026 CEO Survey: 56% of CEOs report neither increased revenue nor decreased costs from AI in the past twelve months. Only 12% report both.
  • MIT's Project NANDA: 95% of generative-AI deployments produced no measurable P&L impact.
  • Morgan Stanley: only 21% of S&P 500 companies could cite a measurable AI benefit at all.
  • S&P Global: 42% of companies abandoned most of their AI projects in 2025.
  • IBM's CEO study: roughly a quarter of initiatives delivered expected ROI; 56% of CEOs reported no significant financial benefit.
  • KPMG, remarkably, found that 65% of UK organizations would keep investing in AI regardless of tangible ROI, spending sustained by fear of falling behind rather than by measured value.

The market has started pricing this in. Citi identified a 30-basis-point credit spread penalty for companies classified as AI adopters without evidence of return, the debt market literally charging a premium for unproven AI spend.

Here is what makes this moment strange: the technology works. Individual productivity gains are real, repeatedly measured, and sometimes dramatic. The disconnect isn't capability. It's that spending is unmeasured, ungoverned, and, as we'll see, substantially wasted. The problem isn't that companies bought AI. It's that nobody is watching where the tokens go.

The token curve: why this gets much bigger, fast

To understand what's coming, follow one number.

In May 2024, Google processed about 9.7 trillion tokens per month across its products. By May 2025: 480 trillion. By October 2025: 1.3 quadrillion. At I/O in May 2026, Google reported over 3.2 quadrillion tokens per month: roughly a 330-fold increase in two years, running at about 1.2 billion tokens every second.

And this is early. Goldman Sachs Research projects global token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month: with a potential 55x by 2040 if enterprise agents reach full adoption. The driver is the shift from chatbots to agents. A chatbot answers a question: one call, done. An agent monitors, plans, calls tools, checks its own work, retries, and runs around the clock. Gartner's analysis puts agentic workloads at 5 to 30 times more tokens per task than a chatbot interaction. Nvidia's Jensen Huang has said agentic AI requires on the order of 1,000% more compute than generative AI.

Now, the standard rebuttal: token prices are collapsing, down more than 90% since 2023, with inference cost falling 60-70% per year. Doesn't cheap compute solve this?

It doesn't, and a 160-year-old economic principle explains why. In 1865, William Stanley Jevons observed that when the steam engine made coal use more efficient, coal consumption didn't fall, it exploded, because cheaper energy unlocked uses that weren't economical before. Apollo's chief economist Torsten Slok applied it directly to AI this June: as tokens get cheaper, companies don't spend less, they run more agents, automate more workflows, generate more code. The data confirms it brutally: across Ramp's 70,000-business dataset, token consumption grew 1,000% in fifteen months (January 2025 to April 2026), and total corporate AI spending doubled since late 2025, during the steepest price collapse in the technology's history. Inference now consumes roughly 85% of enterprise AI budgets.

This is the intellectual heart of our thesis: unit efficiency does not reduce aggregate consumption. It never has, not for coal, not for electricity, not for compute. The physical economy learned this lesson the hard way, and its answer was not "make each unit cheaper." Its answer was the circular economy: a waste hierarchy, measurement, and the discipline of not consuming what you don't need.

Tokens are electricity: the translation nobody does

Every token is computed somewhere, and computation is energy. The token curve above is an energy curve, and the energy system is already straining under it.

The current state, from the most credible sources available:

  • The IEA estimates global data centers consumed about 415 TWh in 2024 (~1.5% of world electricity), on track to roughly double to 945 TWh by 2030 and reach 1,200 TWh by 2035, growing more than four times faster than all other electricity demand combined. Electricity use in AI-accelerated servers specifically is growing ~30% per year, and the IEA reported that demand from AI-focused data centers grew about 50% in 2025 alone, against 3% growth in electricity overall.
  • US data centers consumed 183 TWh in 2024: over 4% of national electricity, roughly the annual consumption of Pakistan, projected by the IEA to grow 133% to 426 TWh by 2030. The Electric Power Research Institute's 2026 analysis is starker: 9% to 17% of all US electricity by 2030, up from 4-5% today. Bain's baseline says 9%. Whichever forecast you take, data centers account for nearly half of all US electricity demand growth between now and 2030.
  • The concentration is the sharpest part. Data centers already consumed about 26% of Virginia's entire electricity supply in 2023. A single AI hyperscale facility draws as much power as ~100,000 households; the largest under construction will draw twenty times that, one building consuming like two million homes.

By decade's end, per the IEA, the United States will consume more electricity for data centers than for producing aluminum, steel, cement, chemicals, and all other energy-intensive goods combined.

At the level of one company, the numbers feel innocent, and honesty requires saying so. A single typical model query runs on the order of a third of a watt-hour; a mid-market agent fleet pushing a billion tokens a month draws megawatt-hours at most, a rounding error alone. But two things break the innocence. First, reasoning and agentic workloads are one to two orders of magnitude hungrier, long reasoning prompts have been measured at 33+ Wh, over 70 times a small model's draw, and agentic tasks multiply calls 5-30x. Second, there is no "alone": multiply across hundreds of thousands of fleets, growing 24-fold, and you arrive exactly at the national curves above. AT&T alone scaled from 8 billion to 27 billion tokens per day on multi-agent systems in six months. At Google's I/O, it was disclosed that 375 global enterprises now each consume more than a trillion tokens annually. Fleet by fleet, the innocence disappears into the grid.

And the grid is answering in ways that would have been unthinkable five years ago: the IEA estimates ~20% of planned data center projects risk delay for lack of grid connection; utilities are building new gas-fired plants dedicated to data centers (Microsoft-Chevron in West Texas; a 933 MW gas facility linked to Google in Texas); and two retired nuclear plants, Three Mile Island and Duane Arnold, are being revived specifically to feed AI demand.

The carbon ledger: what the sustainability reports just admitted

The energy story became a carbon story the week this article was written.

On July 9, 2026, Microsoft reported its greenhouse gas emissions rose ~25% in a single year: to roughly 20-21 million metric tons CO2, driven by its AI data center buildout, even as it reaffirmed a "carbon negative by 2030" pledge that now points in the opposite direction of its trajectory. Its own report contains the sentence that should be read in every boardroom: "sustainability solutions are not scaling fast enough to meet demand."

Microsoft is not the exception; it is the pattern. Alphabet's emissions rose 48% between 2019 and 2024, with a further 18% jump last year to 18.8 Mt CO2. Amazon's rose 16%. These are the three companies with the most sophisticated sustainability programs on Earth, the largest clean-energy purchase agreements in history, and AI demand is outrunning all of it. The IEA now expects CO2 emissions from data center electricity to reach 300 million tonnes by 2035, nearly doubling from today.

Every serious climate pathway assumed electricity demand efficiency; AI just added a Pakistan of demand to the US alone, with a 24x usage multiplier queued behind it. And what is shocking is that a large fraction of this energy is spent producing nothing. Which brings us to waste.

Inside the machine: how much of this is simply wasted?

If every token were productive, the story above would be a difficult but honest trade-off: energy for value. It is not. Engineering analyses of production LLM workloads converge on numbers that would be considered scandalous in any physical industry:

  • 40-60% of input tokens in typical unoptimized agentic workflows are context the model doesn't need: stale conversation history, boilerplate prompts, entire files attached when three functions matter. Every one of those tokens is paid for, and computed, on every call.
  • Practitioners report 15-40% of total token spend wasted in large organizations (with a range of 5-50%), driven by shadow AI spend, redundant calls, no caching, and "model overkill", frontier models running trivial classification tasks.
  • Research on coding agents finds review-and-rework loops consume ~59% of tokens: the machine checking and redoing its own work dominates the bill.
  • Agent teams consume ~7x the tokens of standard sessions; a wasted hundred tokens in turn one of a thirty-turn session is paid thirty times over.
  • And an entire category of waste that no one even measures: duplication. Inside a single organization, different teams re-embed the same codebase, re-scrape the same sources, re-parse the same documents, re-generate the same datasets, because no agent knows what another agent already produced. Roughly 90% of agent tasks produce a persistent artifact; almost none of those artifacts are ever reused. Every duplicate is compute, cash and carbon, spent to recreate something that already exists.

Translate that to the physical world and the absurdity is immediate: a factory that throws away half its raw material before the production line, manufactures the same component in five departments because none knows the others exist, and has no meter on any of it. In manufacturing, that operation goes bankrupt. In AI, in 2026, that operation is normal, because the bill arrives as one undifferentiated line item, and the carbon arrives as someone else's sustainability report.

What companies are doing about it: and why it isn't enough

To be clear: the optimization toolbox exists, and it works. The current state of practice:

  • Prompt caching (provider-native): cache reads cost ~10% of standard input pricing, a 90% discount on repeated context, with break-even at just over two reuses. Widely considered the highest-ROI first move.
  • Model routing / right-sizing: sending each task to the cheapest model that handles it, 40-70% savings on routable traffic in published analyses.
  • Context engineering and compression: attacking the 40-60% waste directly; 50-70% token reduction reported.
  • Semantic caching: up to ~73% cost reduction on high-repetition workloads.
  • Batching: a standing 50% discount for non-urgent work.

Applied together, engineering guides consistently report 50-85% cost reduction potential. A discipline has formed around this: "FinOps for AI", 98% of FinOps practitioners now manage AI spend, up from 31% in 2024, and in June 2026 the Linux Foundation launched a Tokenomics Foundation, backed by Oracle, Google, Microsoft, and JPMorganChase, to standardize AI cost management. GitHub is moving Copilot to usage-based billing. The wake-up is real.

But look at what the toolbox is: point solutions, adopted piecemeal, by the minority of teams with the expertise and the mandate, each attacking one layer, none measuring the whole, and almost none touching duplication (the reuse problem) at all. Meanwhile the observability tools that do exist tell you what you spent, not what you wasted, dashboards without verdicts. And over all of it hangs Jevons: every efficiency gain gets reinvested in more consumption unless something changes the logic of consumption itself.

The physical economy has seen this exact movie. Efficiency programs alone never solved industrial waste. What solved it, where it has been solved, was a systematic framework: measure everything, prevent what you can, reduce what you can't prevent, reuse before you remake, recycle what remains. A hierarchy, applied as a system, with the meter running. That framework has a name.

The wrong answer: rationing the most useful technology we have

Watch what companies actually did when the bills exploded: Uber capped usage. Microsoft revoked licenses. Forrester's data shows a quarter of planned AI investment postponed. Employees who were told last year to "adopt AI or be left behind" are being told this year to use less of it.

This is the wrong answer, and everyone involved knows it. These companies didn't cap AI because it stopped working, Uber's adoption tripled because engineers found it valuable. They capped it because they could not distinguish productive spend from waste, so they rationed both together. It's the bluntest possible instrument: when you can't see waste, you can only cut volume, and cutting volume cuts the value with it. A technology whose entire purpose is to raise productivity and lower cost is being suppressed by its own inefficiency.

The right answer is the one every mature industry eventually reached: don't ration the resource; eliminate the waste. Keep every productive token; kill every wasted one; and be able to prove, in real numbers, which was which.

And there is a deadline on learning this. Today the constraint is budgets. Within a few years, on every credible forecast above, the constraint becomes physical: grid connections, gigawatts, permitted data centers. Demand is projected to grow 24-fold while the infrastructure to serve it fights decade-long interconnection queues and communities pushing back on the gas plants next door. The conversation will move, sooner than most expect, from "our token bill is too high" to "there is not enough power to run everything we want to run." When compute becomes rationed by physics rather than finance, the organizations that learned to run lean fleets will simply get more intelligence per available watt than those that didn't. Waste discipline is about to become competitive advantage, then necessity.

Circular economy for AI - the Circulara AI

This is precisely the moment the circular economy was invented for. It emerged in the physical world when waste volumes made linear consumption, take, make, dispose, untenable, and it succeeded by replacing exhortation with a system: a waste hierarchy (refuse, reduce, reuse, recycle), applied in order, with honest measurement underneath. We spent over a decade building technology for that world, reuse platforms, recycling platforms operating across 150 countries, circular logistics, and today we serve circular-economy businesses with agentic AI through our Circular Route platform. We have watched an economy learn, painfully, that you cannot manage what you do not measure and you should never remake what already exists.

Compute is the next resource to learn it. Circulara applies circular economy principles to agentic AI, inside one organization, starting now:

  • Reduce: stop over-spending on every call: right-size models to tasks (the single largest lever, learned safely from each organization's own usage), compress context, cap waste before it's computed.
  • Reuse: the principle nobody else applies: the expensive artifacts an organization's agents produce, embeddings, parsed documents, datasets, tool results, fingerprinted, governed, and reused across the whole fleet instead of regenerated. No agent pays twice for what the organization already owns.
  • Recycle: recover value from repeated work through disciplined caching, strictly quality-gated.

Everything runs behind one meter, and the meter is the point: savings shown in real time, in dollars, tokens, kilowatt-hours, and CO2: actual spend beside what would have been spent, per team and per user, computed transparently from usage metadata at published rates. For the CFO, that's operating margin. For the sustainability team, it's ESG-reportable avoided energy and avoided emissions, the same measurement discipline circular economy brought to materials, applied to compute. Organizations applying these principles can save up to 30-40% of their tokens, with environmental impact reduced accordingly, not by using AI less, but by wasting less of it.

We won't pretend the whole framework arrives at once; we begin with the principles that pay immediately and prove themselves in the meter. But the direction is not in doubt. Compute is energy. Energy is carbon. Demand is going vertical, supply is physical, and waste at this scale is a choice. The AI economy will adopt circularity for the same reason the physical economy did: first because it's profitable, then because it's unavoidable. We'd like to start that conversation now, while it's still merely profitable.

The turn against waste always begins the same way: not with a mandate, not with a pledge, but with someone finally looking at the meter. Yours is already running.

Start by measuring. It’s free

You can’t optimize what you can’t measure. So we made measuring free.

Circulara Observer is a free tool you can install today as an MCP plugin into your AI tooling. It watches your organization’s real usage and shows, in one live meter: what you’re spending, in tokens, dollars, kWh, and CO2: and what you could be spending, including how much routing simple tasks to lower-tier models would save you, per user and per task.

  • Free, indefinitely

    No trial clock, up to 3 seats per organization.

  • Counts, not content

    Observer reads usage metadata only. Never your prompts, outputs, or code. Your data stays yours.

  • No commitment

    Measure first. Decide about paid tiers only after your own numbers make the case, or don't.

Sources

  1. 01Axios, "AI sticker shock hits corporate America" (May 2026): $500M single-month AI bill; Microsoft canceling Claude Code licenses; Uber COO "harder to justify"; weather-checking anecdote. axios.com/2026/05/28/ai-spending-roi-enterprise-costs
  2. 02Cockroach Labs, "The Bill Arrives: Managing Agentic AI Costs at Scale" (June 2026): Uber CTO quote, 32%-84% adoption, budget exhausted in four months; Gartner 5-30x tokens per agentic task; inference = 85% of AI budgets; prompt-caching economics. cockroachlabs.com/blog/agentic-ai-costs-at-scale
  3. 03AI Business Weekly, "Enterprise AI Spend Scrutiny 2026": Uber $1,500/month cap; Microsoft $500-$2,000 per-engineer bills; Forrester 25% spend postponement; Gartner/Bain findings. aibusinessweekly.net
  4. 04Forbes, "56% of CEOs See Zero ROI From AI" (Jan 2026): PwC 2026 CEO Survey. forbes.com
  5. 05Terminal X Research, "AI ROI in 2026": MIT NANDA 95%; S&P Global 42%; Morgan Stanley 21%; IBM 25%/56%; Citi 30bp credit penalty; hyperscaler capex $675B. terminal-x.ai
  6. 06Unico Connect, "AI Statistics 2026": Gartner $2.59T worldwide AI spending (+47%); McKinsey 88% adoption; RAND >80% failure. unicoconnect.com/blogs/ai-statistics-2026
  7. 07CIO, "KPMG report finds enterprise disconnect between AI and its ROI" (April 2026): 65% invest regardless of ROI. cio.com
  8. 08Goldman Sachs Research, "Decoding the Agentic Economy" (May 2026): 24x token growth to 120 quadrillion/month by 2030; 55x by 2040; 60-70%/yr inference cost decline. goldmansachs.com/insights
  9. 09TECHi / TechJack, Google I/O 2026 token disclosures: 9.7T (May 2024) - 480T (May 2025) - 3.2 quadrillion tokens/month (May 2026); 375 enterprises >1T tokens/yr (vendor-reported figures). techi.com
  10. 10GreyJournal, "How Much Companies Spend on AI Tokens in 2026": Ramp data (+1,001% token growth; median $2,246 vs average $140,842/month; top 1% $7,450/employee/month); Jevons paradox / Torsten Slok (Fortune, June 2026); Silicon Data Token Expenditure Index; Tokenomics Foundation. greyjournal.net
  11. 11IEA, "Energy and AI" (2025/2026): 415 TWh (2024) - 945 TWh (2030) - 1,200 TWh (2035); US 183 - 426 TWh; data centers ~ half of US demand growth; > aluminum+steel+cement+chemicals combined by 2030; ~20% of projects at grid risk; 300 Mt CO2 by 2035; AI data center demand +50% in 2025. iea.org/reports/energy-and-ai
  12. 12EPRI, "Powering Intelligence 2026": 9-17% of US electricity by 2030; 177-192 TWh (2024) - 380-790 TWh (2030); single campus = 80,000-800,000 homes. powering-intelligence.epri.com
  13. 13Bain & Company, Data Center Model (2025): 9% of US electricity by 2030. bain.com
  14. 14Pew Research (Oct 2025): US 183 TWh ~ Pakistan; Virginia 26%; hyperscaler = 100,000 households, new builds 20x; Three Mile Island / Duane Arnold revivals. pewresearch.org
  15. 15Fortune / Forbes / Axios / Spokesman (July 9-10, 2026): Microsoft FY2025 emissions 20-21 Mt CO2e, +25-27%; "sustainability solutions are not scaling fast enough"; Chevron gas deal; Alphabet +48% (2019-2024), +18% (2025, 18.8 Mt); Amazon +16%. fortune.com, forbes.com
  16. 16Carbon Brief, "Five charts on data-centre energy" (2025): global context, ~1% of electricity, ~0.5% of CO2 today. carbonbrief.org
  17. 17Morph, "LLM Inference Optimization" & "LLM Cost Optimization" (2026): 40-60% of input tokens wasted; 70-85% combined savings; routing 40-70%; caching 90%; batching 50%. morphllm.com
  18. 18Alpacked, "LLM Deployment Cost Optimization in 2026": 5-50% token waste (typical 15-40%); shadow AI spend. alpacked.io
  19. 19Token Optimize, "LLM Token Optimization Strategies 2026": rework loops ~59% of tokens; agent teams ~7x tokens. tokenoptimize.dev
  20. 20Redis, semantic caching up to ~73% cost reduction. redis.io
  21. 21FinOps Foundation, "Token Economics" (2026): AT&T 8B - 27B tokens/day; enterprise pricing shifts; IEA AI-demand growth. finops.org
  22. 22FinOps Foundation "State of FinOps" (via 2026 reporting): 98% of practitioners managing AI spend, up from 31% in 2024.

Circulara, Circular Economy for AI. A product of Circular Route Inc.

Circulara AI

Get started.
Let the math decide.

Circulara is free to start for up to 3 seats. Within 30 days its Report shows your fleet’s real spend, what Circulara would have saved, and the carbon that comes with it. Upgrade only if the number is obvious.