The cost of GPT-3.5-level intelligence fell from $20 per million tokens in November 2022 to $0.07 by October 2024, a drop of more than 280-fold in about 18 months. Over the same period every major lab kept rationing its best model: OpenAI's free tier gets GPT-5.6 Luna and "unlimited everyday text chats" while GPT-5.6 Sol's top reasoning settings need a Pro, Business, or Enterprise plan; Anthropic caps API spend at $500, $1,000, or $200,000 a month by tier and states that its limits are "maximum allowed usage, not guaranteed minimums"; Microsoft told investors in July 2026 that "customer demand continues to exceed available capacity." Both facts are true at once because AI has two prices. This piece lays out the evidence for each, shows how Choice OMG routes its own work between them (including the August audit that found every strategy draft running at double the batch rate), and documents three incidents where a model fabricated a figure, one of them on the best model available with no gate in front of it. Every number is copied verbatim from a dated source listed at the end.
"AI costs are collapsing" and "AI capacity is sold out" are both in the news on the same day, and both are accurate. The first describes the price of reproducing a capability that already exists. The second describes the price of the capability that does not exist yet, plus the chips, electricity, and latency around it. An agency that buys AI as one product at one price will overspend on routine work and underinvest where model quality decides the outcome. We run both tiers in production, so we can show what the split looks like in practice rather than in theory.
The commodity floor is falling 9x to 900x a year
Stanford's 2025 AI Index measured it directly: "The cost of querying an AI model that scores the equivalent of GPT-3.5 (64.8% accuracy) on MMLU dropped from $20 per million tokens in November 2022 to just $0.07 per million tokens by October 2024 (Gemini-1.5-Flash-8B), a more than 280-fold reduction in approximately 18 months."
That is one capability level. Epoch AI, whose data Stanford's chart draws on, tracked the same question across six benchmarks and found the decline is steep everywhere and uneven by task: "we found prices declining between 9x per year and 900x per year, with a median of 50x per year." The fastest drops are the most recent ones; Epoch notes "the fastest trends (e.g. 900x per year) start after January 2024." General-knowledge and coding thresholds fell 9x to 40x a year; PhD-level science questions fell 40x to 900x a year.
The energy side moves the same way. The International Energy Agency's 2026 report states that "Software and hardware advances have resulted in the energy use per AI task dropping by at least an order of magnitude annually in recent years. Simple text queries now typically consume less electricity than running a television over the same period of time." The IEA puts the cost of running every conventional internet search as a simple AI text query at "less than 4 terawatt-hours (TWh) of electricity annually, equivalent to less than 1% of total data centre consumption today."
A business reading those three sources could reasonably conclude that AI is about to be free. For last year's capability, that conclusion is roughly right.
The frontier is rationed, and the vendors say so in their own documentation
The same labs that give away yesterday's model meter today's. The evidence is on their pricing and limits pages, not in analyst commentary.
OpenAI prices the ladder explicitly. Its current API lineup lists GPT-5.6 Luna ("Fast, affordable model for everyday work") at $0.20 per million input tokens and $1.20 output, GPT-5.6 Terra at $2.00 and $12.00, and GPT-5.6 Sol ("Flagship model for ambitious agentic work") at $5.00 and $30.00. Cached input is one tenth of standard input on all three. The developer pricing table then prices the same model four ways by service level: Batch and Flex at exactly half of Standard (Sol drops to $2.50 and $15.00), "Fast mode" at double (Sol $10.00 and $60.00), and long-context input at double. Flex is described as "lower costs for requests in exchange for slower response times and occasional resource unavailability." Scale Tier and Reserved Capacity are sold by the sales team to "cutting-edge customers running larger workloads." A model with one marginal cost does not need a price list that spans a 4x spread for the same tokens; a scarce input served from constrained capacity does.
ChatGPT's free tier and paid tiers are different products. OpenAI's help centre says "Free users have unlimited everyday text chats, subject to abuse-prevention safeguards," on GPT-5.6 Luna. The same documentation says "Free and Go users receive GPT-5.6 Luna and do not have access to GPT-5.6 Sol." Plus gets Sol's Medium and High reasoning settings; Extra High and Pro need Pro, Business, or Enterprise. When a paid user hits a reasoning limit, "ChatGPT may continue with another available reasoning model." The abundant product is unlimited. The scarce one is metered by plan, by setting, and by allowance.
Anthropic publishes the rationing rules. Its rate-limit documentation states that "Limits are designed to prevent API abuse, while minimizing impact on common customer usage patterns," that "All limits described here represent maximum allowed usage, not guaranteed minimums," and that they exist to "ensure fair distribution of resources among users." Monthly spend caps by usage tier are $500 (Start), $1,000 (Build), and $200,000 (Scale); "Once you reach your tier's spend cap, API usage pauses until the next month unless you request a higher limit."
Google's free tier is paid for in data. The Gemini API pricing page invites developers to "Start building free of charge with generous limits." The free tier's feature list includes "Limited access to certain models" and "Content used to improve our products"; the paid tier's column for the same question reads "No." Gemini 3.7 Flash is "Free of charge" for input, output, and caching on the free tier and $0.75 per million input tokens on the paid tier through December 31, 2026, rising to $1.50 on January 1, 2027. A price increase scheduled a year out on a "Flash" model is its own data point about where the vendor expects demand to sit.
The physical constraint is on the record. Microsoft's chief financial officer told investors on 2026-07-29 that "Customer demand continues to exceed available capacity," that quarterly capital expenditures were $41 billion, and that the company expects "CapEx spend will be over $50 billion" in the coming quarter. Its chief executive said Microsoft "added another gigawatt of capacity this quarter and remain[s] on track to roughly double our overall capacity in just two years." Nine months earlier, the same executives had said "demand again exceeded supply across workloads even as we brought more capacity online." The constraint did not clear in three quarters of record build-out.
NVIDIA's fiscal 2026 results are the upstream mirror image: revenue "was $215.9 billion, up 65% from a year ago," Data Center revenue "rose 68% to a record $193.7 billion," and full-year GAAP gross margin was 71.1%, down from 75.0% the year before. Scarcity rents are large and already being competed down at the same time.
Total spend rises anyway, and "Jevons paradox" is usually the wrong name for why
The standard explanation says cheaper AI means more AI, citing Jevons. The economics literature is more careful. Sorrell, Dimitropoulos, and Sommerville's 2009 review in Energy Policy defines the mechanism plainly: "Improvements in energy efficiency make energy services cheaper, and therefore encourage increased consumption of those services. This so-called direct rebound effect offsets the energy savings that may otherwise be achieved." Their finding for household energy in the OECD is that "the direct rebound effect should generally be less than 30%." A rebound below 100% still saves resources. Jevons paradox, or backfire, is the special case where consumption ends up higher than before the efficiency gain.
AI appears to be in the backfire case, and the IEA's 2026 numbers are the cleanest single statement of it. In the same report that records efficiency per task improving "by at least an order of magnitude annually," the agency writes: "Electricity consumption from AI-focused data centres grew even faster, surging 50% in 2025. While there are no comprehensive statistics on the frequency and depth of AI usage around the world, major model providers reported a threefold increase in active users and a fivefold increase in revenue over the past year." Epoch AI reports the revenue side independently: "Inference revenue at major AI companies such as OpenAI and Anthropic has been growing at a rate of 3x per year or more, even as their models continue to become smaller and cheaper compared to 2023."
The rebound in AI has a second component that energy rebounds lack. Users do not only run more of the same query; they move up a quality ladder. The IEA names it: "new energy-intensive AI applications are increasingly being launched and used, such as those for video generation, reasoning and agentic tasks. These kinds of tasks can consume hundreds or thousands of times more energy per query than simple text generation." Epoch describes the same shift from the serving side: the old benchmark of "human reading speed," about 10 tokens per second, "has become obsolete" because "models are asked to reason at length about complex problems and are placed inside elaborate agentic loops."
Stanford's 2026 AI Index closes the loop on who is paying. "AI company revenue is rising at historically fast rates, but compute costs and infrastructure spending are also reaching record levels," with "Google reporting more than $150 billion in annual capex in 2025." Stanford estimates U.S. consumer surplus from generative AI "reached $172 billion annually by early 2026, up from $112 billion a year earlier," and adds that "Most of these tools remain free or close to it."
The downgrade test
The commodity-or-frontier question is a property of the task, not of the model's name. The test we apply:
Replace the model with one tier cheaper. If the business outcome is unchanged, the task is commodity work and should run on the cheapest model that clears the bar. If the outcome degrades in a way that costs money, the task is frontier work and the question becomes how much verification to add, not whether to pay.
Six questions decide it in practice: What happens to the output if model quality drops one tier? How easily would an error be noticed? What does an unnoticed error cost? Does speed create business value here? Can the work be batched or cached? Would several vendors be acceptable? Low stakes, easy detection, and batchable work route down. High stakes, hard detection, and a decision that moves money route up.
How we route it
Choice OMG runs Anthropic's models at three tiers, with the reasoning effort set per dispatch. The routing rule in our operating documentation, verbatim:
| Task class | Model | Effort |
|---|---|---|
| Mechanical: the plan already contains the code, single-file fix, read-only lookup, mechanical search | Haiku | none |
| Standard: multi-file integration, debugging, moderate research, implementing from a prose spec | Sonnet | medium |
| Hard: architecture and design, whole-branch review, subtle concurrency, broad-codebase reasoning, risky diffs | Opus | high (extra high for hard agentic or coding work) |
The principle written next to that table is "fastest = smallest model plus lowest effort that still succeeds, then escalate on failure," with the caveat that "the cheapest models take 2-3x the turns on multi-step work, costing more overall," so Haiku stays on genuinely mechanical transcription and Sonnet is the floor for anything that produces a report someone will act on.
The commodity tier does real volume. In June 2026 we classified every company name in our opt-in lead database into an 18-label industry taxonomy using Haiku, in chunks of 1,000 to 1,500 rows. It worked, and the failure modes were exactly the kind the downgrade test predicts for cheap tiers: some chunks came back with row IDs renumbered from 1, some were truncated (about 900 of 1,500 lines returned in one case, 680 of 750 in another), and a few contained stray labels outside the taxonomy. All three were caught by mechanical checks (ID reconciliation, line counts, label validation) and re-run on smaller chunks. The model was cheap; the reconciler was the cost. Commodity-tier data work is priced per client in cents, not dollars: a full technical crawl of one client site through our data provider cost $0.54, and the quarterly AI-visibility probe we run for clients outside the weekly programme costs about $0.303 per client per quarter.
The frontier tier is where the spending decision got formalized. On 2026-08-18 an audit of the API key that drafts client strategy documents and monthly report narratives found that a Batch API path built in July, which carries a flat 50% discount, had never been run: every August strategy draft had gone through the synchronous path at full list price, so each client's draft cost exactly double what the same tokens cost batched. The decision that came out of it, recorded the same day:
- Strategy drafts and monthly client summaries stay on the top model (moved to Opus 5 at the same $5 and $25 per million tokens list price). The work is judgment-heavy and client-facing, and the only layer without a review gate had already produced fabricated figures on this same tier (below).
- Critique and condensing layers stay on Sonnet 5 ($2 and $10 per million until 2026-08-31, then $3 and $15).
- Any run covering two or more clients goes through the Batch API by default. The rate is the whole argument: the same strategy draft for the same client, on the same model, costs half when it can wait up to 24 hours for the answer, and nothing about a quarterly strategy refresh needs the answer in seconds.
- Judgment work that a person is doing in a session anyway (discovery interviews, draft review, positioning) runs on staff subscriptions at zero marginal API cost. Headless scheduled work stays on the API because a personal subscription token on a server cron is off-licence and fragile.
Spend did not go to zero because models got cheaper. It got split: commodity work to the cheapest model that passes mechanical checks, frontier work to the best model with a 50% discount for being patient, and human judgment kept off the meter.
Three times a model invented a number
The downgrade test only works if you know what a downgrade actually breaks. In our experience it rarely breaks the artifact. It breaks the evidence trail. Two of the three incidents below were on the cheap tier. The first was on the flagship, which is the point.
1. The monthly summary that lifted the wrong month's spend. In July 2026 the drafter that writes a client's monthly performance summary, running on the flagship model of the day (Opus 4.7), twice pulled "$9,789.30" and once a "$489" cost per acquisition from a strategy document dated July 10 and presented them as the month's campaign figures. The real July numbers in the internal report were $13,002.74 at $260.07. A prompt rule forbidding figures not in the report did not stop it, and neither did the model tier. The fix was mechanical: every dollar amount in a generated summary now has to appear verbatim in the source report (or equal an exact sum of two or three report amounts on a line that says "total"), and the build fails otherwise. Rounded restatements fail on purpose.
2. The client-facing condensed strategy that fabricated figures and leaked ticket IDs. On 2026-08-13 the layer that condenses a full strategy into a short client document was found to be the only layer in the pipeline with no review gate: each regeneration ran a fresh Sonnet call over the source documents with no check on its own output. For one client it produced a document with dollar figures that existed nowhere in the inputs and internal ticket references that should never leave the building. Re-rolling the generation does not fix this class of defect; it reproduces it. The fix added a mechanical check that fails the build on any dollar figure not verbatim in the source documents or any internal ticket pattern, plus a render-only mode so a hand-corrected document can be rebuilt without the model touching it again.
3. The verification output that was reconstructed from expectation. On 2026-08-09 a Haiku worker was given a fully specified nine-line script. It transcribed the script correctly, ran the real verification commands, and then wrote a report whose pasted output line was fabricated: it printed a passing value and invented a sentence explaining it, when the real command had printed a failure. The shipped code was fine. The report was not, and the only way it was caught was a reviewer re-running the command.
The pattern across all three is the same, and the tier did not decide it. Each model handled the mechanical part and failed at synthesizing evidence about its own work, which is the part a reader relies on. The tier changes how often that happens; only a gate changes whether it ships. That is why our routing rule says Sonnet is the floor "whenever the report will be relied on as evidence without independent re-execution," and why every report-generating surface we build now carries a mechanical number check on every tier, not just a prompt instruction. The cost that token prices removed came back as verification.
What this means for a business buying marketing
Agencies are intermediaries in this market, and their clients should know which tier their money is buying.
| Work | Commodity tier | Frontier tier | What the client is actually paying for |
|---|---|---|---|
| Content | Drafts, variations, formatting, tagging | Original research, positioning, editorial judgment | Evidence, brand, and the editor's accountability |
| Paid media | Classification, reporting, bulk analysis | Budget decisions, anomaly diagnosis, new strategy | Conversion history and appointment-level outcomes |
| Reporting | Pulling and formatting the numbers | Explaining what changed and what to do | A number check that fails the build on an invented figure |
| Sales | Research, proposal assembly | Discovery, pricing, negotiation | Reputation and case history |
Two consequences follow. An agency that charges frontier prices for commodity work is capturing the 280-fold price drop for itself. An agency that runs any model, at any tier, without a check against the source numbers is the one whose report will eventually contain a $9,789.30 that never happened. The defensible position is to route honestly and show the gates.
How to read the next "AI costs collapsed" headline
Five questions, in order:
- Constant capability? "GPT-3.5-level performance is 280x cheaper" holds capability fixed. "The newest small model is cheaper than last year's flagship" does not.
- Constant task? A cheaper token is not a cheaper job if the job now runs a reasoning loop, ten candidate answers, browsing, and a verification pass.
- Which tier? Commodity prices can collapse while the frontier is sold by spend cap and waiting list. Both are happening.
- Total system cost? Inference is one line. Integration, evaluation, human review, error correction, and the mechanical checks that catch a fabricated figure are the others, and in our pipeline they are where the money went.
- Who keeps the saving? The customer through lower prices, the application vendor through margin, the model provider through volume, or the chip and power supplier through scarcity pricing. The 280-fold drop and NVIDIA's 71.1% gross margin are the same market seen from opposite ends.
The durable version of the headline reads: a cost decline at fixed capability creates commodity supply; the saving expands usage, raises the expected quality of every task, and funds movement toward the frontier; value moves to whichever complementary input stays scarce. For an agency, the scarce inputs are the client's outcome data, the verification around the model, and the person who signs the report.
Sources and further reading
- Stanford HAI, AI Index 2025: State of AI in 10 Charts (the $20 to $0.07 figure)
- Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks
- Epoch AI, Inference economics of language models
- IEA, Key Questions on Energy and AI: Executive Summary (2026)
- Stanford HAI, AI Index 2026: Economy
- OpenAI, API pricing and developer pricing tables
- OpenAI Help, ChatGPT Free Tier FAQ and GPT-5.6 in ChatGPT
- Anthropic, Rate limits
- Google, Gemini Developer API pricing
- Microsoft, FY2026 Q4 earnings call (2026-07-29) and FY2026 Q1 earnings call (2025-10-29)
- NVIDIA, Financial Results for Fourth Quarter and Fiscal 2026 (2026-02-25)
- Sorrell, Dimitropoulos and Sommerville, "Empirical estimates of the direct rebound effect: A review," Energy Policy 37(4), 2009
All external sources were scraped and checked on 2026-08-20. Choice OMG routing rules, costs, and incidents are taken from the agency's internal operating documentation and incident records dated 2026-06-08 through 2026-08-18; client names are withheld.