Something strange has been happening to the most sophisticated AI products on the market. Anthropic, maker of what many consider the best model available, is reportedly struggling to attract users while cheaper, scrappier tools eat its lunch.[1] OpenAI, meanwhile, has been yanking usage limits up and down for ChatGPT Plus subscribers like a slot machine operator adjusting the odds — restoring five-hour Codex caps, resetting weekly quotas “to celebrate milestones,” then tightening again.[5] These are not unrelated stories. They point to the same pricing challenge: matching subscription promises to real inference costs.
When the Meter Becomes a Mood Swing
A commenter on Hacker News nailed the consumer experience of one AI product with a line that deserves to be printed on a poster in every pricing team’s office: “You can only use Fable for a week as a part of your plan… Wait, now it’s up to half your usage… Ok, now its…”[1] The observation that follows is the important part — most people want to not care. They don’t want to think about their allowance. They want to open the app and have it work.

That is precisely what rationed usage limits destroy. When a company treats its pricing the way it treats model training — as an experiment to be continuously tuned, A/B tested, and adjusted based on live signal — it optimizes for margin capture and forgets that trust is the actual product being sold. Another Hacker News thread on OpenAI’s flip-flopping limits put it even more bluntly: “such a casino vibe.”[5] One reply described agentic programming as “variable reward” — including the dopamine hit of watching an agent finish its turn — but variable reward is a mechanic you design into a slot machine, not something you want your customers to feel about whether their subscription will still work tomorrow.
I’ve written before about the token tax and the seat being a lie — about how per-seat SaaS pricing collapses the moment a product starts doing autonomous, inference-heavy work instead of just displaying a dashboard a human clicks around in. What’s happening at Anthropic and OpenAI right now is the sharpest evidence yet that even the frontier labs, with all their pricing sophistication, haven’t solved this. If the companies building the models can’t find a stable meter, the SaaS companies building on top of those models are inheriting an even harder problem.
The Real Cost Structure Nobody Wants to Show You
The CloudZero team has laid out plainly why AI agents blow up traditional cost assumptions: agents don’t make one inference call and stop, they loop, they call tools, they retry, they hold context across multi-step tasks, and every one of those steps has a compute cost attached.[2] A single “task” a customer sees as one unit of value might represent dozens of underlying model calls, each with its own token cost, and that cost is only loosely correlated with how satisfied the customer ends up being.
This is the mismatch that pushes vendors toward rationing. If you can’t cleanly price the unit of value, you start rationing the unit of cost instead — hours, weekly caps, “Work limits” — and hope customers don’t notice the gap between what they’re being charged for and what they actually experience. They notice. The Hacker News reaction to OpenAI’s shifting Codex limits shows customers keeping score, publicly, in real time, comparing this week’s cap to last week’s and asking who benefits from the confusion.[5]
Monetizely’s analysis of AI agent pricing goes further, arguing that token and credit pricing itself is a temporary scaffold, not a destination — it solves the early problem of “how do we charge for this at all,” but it compresses as inference gets cheaper, which means today’s carefully calibrated token markup becomes tomorrow’s margin evaporation.[4] A pricing model built on a cost input that is structurally guaranteed to fall is a pricing model with an expiration date baked in from day one.
Why the Cheap Tools Are Winning Anyway
Here’s the part that should worry anyone betting on capability alone: the Financial Times reporting cited on Hacker News says Anthropic’s best model is struggling to attract users as cheaper tools thrive.[1] Users are not just price-sensitive, they are confusion-averse. A product that costs slightly more per token but never surprises you will beat a superior model wrapped in a shifting, opaque allowance every time a customer has to choose where to put their default habit.
This connects directly to something I keep returning to in this column: output is not outcome, and neither is “capability.” The best model in the world produces zero value if the customer can’t predict what using it will cost them next Tuesday. Predictability is itself a feature, arguably the highest-leverage feature a pricing team can ship, and it costs nothing to build except discipline.
Aaron Levie’s framing of enterprise AI adoption is useful here too. He argues the value isn’t only in the model, it’s in the bridge from raw model capability to the actual workflow inside a bank, a law firm, or a pharma company.[6] Pricing is part of that bridge. A workflow that a compliance officer can’t budget for with confidence is a workflow that never gets adopted, no matter how good the underlying model is.
Rationing Is Not a Pricing Strategy, It’s an Admission
There’s a temptation, when your inference costs are volatile and your margins are getting squeezed by heavy users, to reach for usage caps as a stopgap. Cap the hours, throttle the weekly quota, quietly narrow the definition of “unlimited.” It feels like a pricing lever. It is actually a confession that you never built a meter that reflects your real costs and real value delivered — you just built a wall and you’re moving it around based on this month’s compute bill.
The fix isn’t more caps, tuned more cleverly. It’s going back to the actual unit of value the customer is paying for — a resolved ticket, a shipped PR, a completed workflow — and pricing that unit with enough margin buffer that you don’t need to touch the dial every few weeks in response to a live cost shock. That buffer is not padding, it’s the price of predictability, and predictability is what the frontier labs are putting at risk.
What Founders Building on Top of These Models Should Take Away
If you’re a startup layering agentic features onto a base model from Anthropic or OpenAI, don’t assume their pricing instability is a problem confined to them. Their token costs are your token costs, one API call removed, and their unpredictability becomes your unpredictability unless you build a firewall between the volatility upstream and the price your customer sees downstream. That firewall is usually some combination of outcome-based pricing, committed usage tiers, and margin discipline that assumes the worst about your own compute bill, not the best.
The lesson from the casino-vibe backlash isn’t “usage limits are bad.” Limits are fine, even necessary. The lesson is that limits which move without warning, explanation, or a stable underlying logic destroy the one thing that makes any subscription business durable: the customer’s confidence that they understand what they’re buying. In the post-seat era, that confidence is the product.
References
- Anthropic’s best AI model struggles to attract users as cheaper tools thrive | Hacker News — https://news.ycombinator.com/item?id=49411102
- How SaaS Companies Can Profitably Price AI Agents — https://www.cloudzero.com/blog/ai-agent-pricing-models
- 5 Pricing Models That Will Break When You Add AI Agents To Your SaaS Product — https://www.getmonetizely.com/articles/5-pricing-models-that-will-break-when-you-add-ai-agents-to-your-saas-product
- OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus users — https://news.ycombinator.com/item?id=49432879
- Box’s Aaron Levie On Reinventing Yourself in the AI Age and Enterprise Diffusion | Sequoia Capital — https://www.sequoiacap.com/podcast/box-s-aaron-levie-on-reinventing-yourself-in-the-ai-age-and-enterprise-diffusion


Leave a Reply