From Token Burn to Token Budget

Updated: Sep 8
What a 20-year-old eBay experiment taught me about AI economics

Lately, I can't escape the phrase "token burn."
It comes up at work, in the AI communities I hang out in, and when I compare notes with other marketing leaders. The vendors are more polite about it: Anthropic and OpenAI both call the meter "Usage", a bland little dashboard tab that's quietly becoming one of the most interesting line items in the company. It definitely comes up when I check my own Usage page after a weekend of building agents and wonder where all the credits went.
The questions underneath are always the same. Where is all this usage going? What happens when the bill goes up? How do we know we're spending it wisely?
Then, within the same couple of weeks, two separate conversations dragged an idea out of my past that I hadn't thought about in years.
Funny thing about working long enough: occasionally the future hands you back an old problem wearing completely different clothes.
Back at eBay, we had a resource problem
Creative demand at eBay far outran creative capacity. Every business unit had important work. Every request was urgent. There simply weren't enough designers, writers, and hours to do all of it.
So a few of us built an allocation system. We called it CHIPs: Customer Happiness Is our Top Priority. (Yes, acronyms were serious business. No, I will not be taking questions.)
The mechanics were simple. Each business team got a finite allocation of CHIPs. One CHIP represented roughly five hours of creative capacity. The team, not Creative, decided how to spend its allocation. Which meant business leaders had to prioritize their requests instead of submitting all of them.
People hated parts of it, maybe all of it. A five-hour CHIP might cover a two-hour project one week and get swallowed by a monster the next; it ran on averages, and averages never feel fair to whoever is on the wrong side of one. There were never enough CHIPs. People complained, nicely. And the whole thing lived in spreadsheets that someone (hi) maintained by hand.
And yet the department ran on it for about a year.
Because CHIPs forced a conversation. Instead of Creative saying "we don't have capacity," the business had to say "of everything we want, this is what matters most."
CHIPs didn't eliminate scarcity. They made scarcity visible.
The idea was novel enough that eBay filed it as a patent, Marketing Allocation Request Systems, with my name among the inventors. The part that matters in hindsight isn't the paperwork. It's what the system paired together: a strategic review (is this worth doing?) with a capacity review (can we afford to do it?).
Then: finite human creative capacity. Now: finite AI capacity.
Many (too many) years later, I realized I was having the same conversation. Nobody was saying CHIPs. Everyone was saying tokens.
We spent the first AI era pretending intelligence was unlimited
And that was the right call. Early enterprise AI adoption was supposed to be about experimentation: buy the licenses, hand out access, let people build strange things and see what blooms. I call this phase "plant a thousand seeds." Redundancy is the price of discovery, and discovery isn't waste.
But look at what the seeds grew into.
Copilots became workflows. Workflows became agents. Agents started calling other agents and tools. Context windows got bigger. Reasoning got longer.
And suddenly the Usage tab matters.
AI token budgets keep getting cheaper. AI keeps getting more expensive.
That's the paradox, and it's not a fluke of anyone's pricing page.
Gartner predicts that inference costs per agentic workflow will rise more than fivefold through 2028, even as the underlying model economics improve. Better unit economics unlock more sophisticated applications, which need more reasoning, more steps, and more tokens. Gartner calls it the "Inference Paradox."
Stanford's Digital Economy Lab measured how hungry agents really are. In the coding tasks they studied, agentic runs consumed about 1,000 times more tokens than code chat or code reasoning. Running the same agent on the same task varied by up to 30x in cost. And more tokens did not reliably buy more accuracy.
Enterprises are already feeling it. In McKinsey's May 2026 AI FinOps survey, 93% of respondents said they had exceeded their AI token budgets. KPMG's Q2 2026 Pulse Survey found only 26% of organizations have real-time visibility into AI operating costs, and just 36% have direct token or usage controls.
The cautionary tale: Uber blew through its entire 2026 AI coding budget by April, and other companies are renegotiating contracts, capping usage by team, or pulling licenses back altogether.
Cheap tokens don't create cheap AI. They just make it easier to consume vastly more intelligence.
The question isn't "how much AI do we have?" It's "where should we spend it?"
Traditional software economics revolve around seats: who needs access? AI asks a different question: what work deserves the resource?
That's when the CHIPs connection clicked. A CHIP was never really about five hours. It was a unit of work competing for finite capacity, with the prioritization decision sitting with the people making the requests. Tokens are far more granular and measurable. The organizational problem is nearly identical.
Don't ration intelligence. Allocate it.
This is not an argument for token austerity.
A $50 agentic workflow that eliminates six hours of manual work is a bargain. A cheap workflow invoked 50,000 times without a meaningful outcome is a terrible deal. The question isn't "how do we use fewer tokens?" It's "where does AI create enough value to justify the burn?"
A few principles I keep coming back to:
Experiment before you optimize. Don't strangle the thousand seeds while they're still sprouting.
Make usage visible. Teams can't manage what they can't see.
Measure work, not just tokens. The meaningful unit is the workflow and its outcome.
Match intelligence to the job. Not every task needs the most capable, most expensive model.
Give teams ownership. Allocation works differently when the requestors do the prioritizing.
Optimize for value, not thrift. The goal is maximum return on intelligence, not minimum AI spend.
What CHIPs got wrong may be the most useful part
The units were rough. The averages felt unfair. The administration was painful. People complained, nicely.
AI can fix the administrative half. Consumption gets captured automatically, costs get attributed to the workflow that generated them, models get routed by task and cost. Nobody has to maintain the spreadsheet.
What technology won't fix is the other half. Someone still has to prioritize. Someone still believes their project deserves more. Someone still thinks another department's allocation is unfair. Someone always wants one more CHIP.
Resource allocation isn't ultimately a math problem. It's a management problem.
The CMO's next budget
Marketing leaders already manage several kinds of finite capital: media budgets, agency dollars, headcount, technology spend, creative capacity. AI capacity is about to join the list.
Not every department will literally get handed a bag of tokens. But usage budgets, workflow-level economics, showbacks and chargebacks, and cost-per-outcome measurement are already showing up. I suspect AI budgets end up looking a lot like media budgets: finite pools of capital, allocated by expected return, defended in the same planning meetings.
Experimentation, then standardization, then allocation. Leaders won't just need to know what works. They'll need to know what deserves to scale.
On a personal note, I became acutely aware of token spend while building agentic workflows and agents. Once I could see usage going out the door, I couldn't unsee it or stop thinking about ways to work more efficiently.
Light years later, I'm back to CHIPs
At eBay, I tracked CHIPs in a spreadsheet. Today I refresh the Usage page while an agent I built decides, on my behalf, that it needs to think a little harder.
The technology could not be more different. The question hasn't changed at all: we have more things we could do than resources to do them. So what matters enough to fund?
Maybe the real lesson of CHIPs was never how to ration capacity. It was how to make choices.
We spent the first era of generative AI asking what we could do with all this new capacity.
The next one is about deciding what it's worth spending on.
Are you interested in how to get the most out of AI for your organization? Contact Hivestir.com to get started. And thanks for stopping by!
A note on how this was written: I outlined this piece and supplied the story, the argument, and the research. Claude (Anthropic's AI) helped verify the sources, draft the prose, and tighten the structure. The opinions, the memories, and the spreadsheet trauma are all mine.



Comments