Cheap AI isn't going anywhere. What's going away is the fantasy that unlimited, frontier-level, agentic work can live forever within a single flat monthly fee.
A few weeks ago, GitHub sent an email telling every Microsoft Copilot customer their plan now comes with an allowance. Metered against the model's rate, with the option to buy more when you run out. OpenAI did the same thing with Codex back in April. Anthropic sells usage bundles now. Cursor gives you a bucket and bills you when the bucket's empty.
I've watched retail do a version of this dance for a decade, so the pattern is familiar to me. The industry ping-ponged between usage-based and flat fees and back again, depending on how badly every vendor exploited one model until grumpy CFOs pushed back (remember when Demandware was the only one taking a % of GMV? Then everyone did it and there was no margin left). Same story, different room.
This is not "AI is getting expensive." It's also not "AI is getting free." Both takes are lazy, and the truth is more interesting than either.
Three things are happening at once. The cost of reaching any given level of capability is falling fast. The work we're asking AI to do is getting longer, more autonomous, more compute-hungry. And the products wrapping that compute are quietly swapping fuzzy promises of "access" for allowances, credits, and usage-based billing. Watch only the first trend and everything looks headed toward zero. Watch the other two and AI looks like it's getting pricier by the week.
The market is splitting into two layers. Everyday intelligence (I'm using "intelligence" lightly here: the "let me Google that for you" query) is becoming abundant and nearly free. High-effort intelligence (long research, tools, agents doing real work, continuous automation) is getting metered.
Those two facts don't contradict each other. They're the whole point.
Chat is a turn, an agent is a process
Nobody has hit a wall on capability. Every few weeks, one of the big labs ships a new flagship, each one better at not just answering a question but at going and completing the work: reading through documents, using tools, checking its own output, and running other agents while it's at it. A chat response is a turn, and an agent is a process.
For your bill, that difference is everything.
The gap between chat and agentic work is not small. A 2026 study ran eight frontier models on coding-agent tasks and found that agentic work chewed through 1,000 times as many tokens as ordinary chat. Runs on the same task varied by as much as 30x, and spending more didn't reliably yield a better answer. Accuracy often peaked somewhere in the middle.
More money, worse result. AI is really powerful and can be incredibly stupid, and now you can watch it be stupid at scale.
The floor is collapsing, the ceiling isn't
There's a countertrend, because there always is one. Epoch AI estimates that the cost of hitting a fixed level of capability has lately been dropping by five to ten times a year (while warning that the evidence is thin and the rate could slow). Yesterday's capability gets cheap in a hurry. Today's frontier keeps wandering off into harder, hungrier work.
So price per token is the wrong yardstick. The right one is a question: what does it cost to get the task done right, at the quality and reliability the customer actually needs? A model that charges more per token but nails it on the first try can be cheaper than a bargain model that needs three reprompts, a human to clean up after it, and a couple of dead agent loops. Cheapest sticker price, most expensive outcome. Retail people know this one by heart.
Not apples to apples - context windows, caching, tools, latency and reliability all differ. The spread is the point.
The spread across providers right now is wild, and none of it is apples-to-apples. But the shape tells the story. The floor is falling through the basement. Classification, summarizing, translation, routine analysis: all on their way to free. Premium intelligence, especially the kind that runs for a long time, still costs real money. Both true. No contradiction.
Open models keep pounding that floor lower, and some now run locally. But "a useful model runs on your laptop" is not "the frontier runs on your laptop." Self-hosting just moves the work onto your plate: security, updates, uptime, capacity planning. Open models compress the floor. They're not the escape hatch from cloud economics people want them to be.
Two cautionary tales
Clearly, the big AI model providers followed the same VC-loved playbook as many that came before: they made access cheap and easy to gobble up market share. But then as it always does, reality sets in. Anthropic cracked down on the OpenClaw agent in April. Then the policy kept moving: in June Anthropic paused a broader proposal and said third-party usage still draws from a subscriber's limits while it reworks the plan.
Just like with the Tariff madness of the past year, the point isn't in which number sticks. It's that number changes, at any time, and enforcement shows up fast. Providers redraw the rules the moment they learn what agentic usage really costs. Enterprise AI users learn not to trust anything in their model for long. Never build a production workflow on subscription-plan arbitrage.
"Google can just afford free" is too easy
Google really does sit in a different seat. In Q1 2026, Alphabet reported more than $60 billion in Search and advertising revenue and over $20 billion in Cloud, and it runs its own chips alongside NVIDIA's. That's a lot of ways to monetize AI that a pure-play model company simply doesn't have.
That's a real advantage, but does not afford infinite patience for losses. Alphabet still spent tens of billions on infrastructure in a single quarter, and it put compute-based limits on Gemini for consumers anyway. More monetization surfaces and its own silicon let Google be more generous. The limits are the tell that even Google is managing something scarce and expensive.
There may not be a single trough
What's actually happening is a split into two layers, and neither is going away. The first is cheap, fast, everywhere: everyday questions, drafting, summarizing, routine analysis. Competition and better hardware keep dragging it toward zero. The second is the heavy stuff: long-running agents, deep research, big-context analysis, always-on automation. It gets more efficient too, and then providers spend every dollar they save on harder tasks and bigger ambitions. Yesterday's frontier becomes cheap. Today's frontier eats the capacity that just opened up. Both can run forever.
If you're building on this, change the unit
I'm not going to pretend this collapses into a tidy three-step checklist. The mental shift matters more than any checklist.
Stop modeling your AI product around a prompt. Model it around a finished customer outcome. Measure the whole path: input, cached input, output, reasoning, retrieval, tool calls, context growth, retries, failure rate, verification, human cleanup, shipping. A workflow that looks cheap at the model level can bleed you dry through sloppy orchestration. A pricey model can be your low-cost option if it finishes the job in fewer loops.
Treat a consumer subscription as a handy purchase option, never an infrastructure contract. Included usage, model availability, credit rates, third-party access: all of it can change under you. A workflow that only pencils out because a provider is temporarily eating the variable cost is not a workflow. It's a bet on someone else's generosity. (But also arbitrage the crap out of those free lunches while they last!)
Build routing, caching, spend controls, and fallbacks in from day one. Spend the expensive reasoning where it changes the answer. Use the cheap model where it doesn't. And track cost per successful task by model and workflow, not the blurry blended monthly bill that hides all your sins.
Stop measuring the prompt. Measure the completed task.
At FindMine, this is the only distinction that's ever mattered. Not one marketer or merchandiser has ever asked me about tokens. They care whether it produces accurate, on-brand, commercially useful work at a predictable price. And when there are misses in quality (and there inevitably are, in any system, AI or human, because to err is human), we fix them for free. That's the bar because it always should have been.
What's actually ending
The free lunch isn't ending because AI is about to get expensive across the board. It's ending because "unlimited frontier compute for one flat monthly fee" was never a real economic category. It was a product of a moment in time where snapping up market share was more important than margin and where the majority use case was a human typing a question, waiting for an answer, and stopping.
Agentic AI blew that world up. It runs long, uses tools, hauls around context, retries, verifies, and sometimes runs an entire team of subagents while you're asleep. The capability gains are real. The efficiency gains are real. The deflation in the price of yesterday's intelligence is real. So is the return of metering.
Cheap AI is becoming abundant. Ambitious AI is finding new ways to eat that abundance for breakfast. Build accordingly.
Michelle Bacharach is the CEO and Founder of FindMine, an AI platform that helps enterprise retailers build the content infrastructure their brands need to show up in search, in AI-powered shopping, and at every point a customer is ready to buy. Top brands use FindMine to automatically generate on-brand, inventory-aware curation and creative at scale (100M consumers a month see FindMine's content). With 15 years of product leadership at the intersection of retail and technology, Michelle has become one of the more direct voices on what it actually takes for brands to compete in an AI-first commerce environment, speaking at National Retail Federation's Big Ideas Keynote, SXSW, and HumanX, and appearing in Forbes, Vogue Business, and Daily Women's Wear. She is an Inc. Magazine Top 200 Female Founder and holds the Retail TouchPoints 40 Under 40 distinction.
