Ask Your Agency Which AI Price It Is Paying

Aerial night render of a vast grid of identical server racks lit in blue, with glowing data lines converging on one golden core at the centre that beams light upward. Overlaid text reads: The two prices of AI. Cheap for yesterday's answers. Rationed for today's. Which one is your agency billing you for? One test sorts every task.
Thousands of commodity racks and one frontier core. Your agency chooses between them for every task on your account.

AI has two prices. The cost of GPT-3.5-level intelligence fell from $20 per million tokens to $0.07 in about 18 months, while every major lab still caps, queues, and tiers access to its best model and Microsoft told investors in July that demand "continues to exceed available capacity." Your agency uses both tiers, and which one it picks for a given task decides whether the saving reaches you. Whether your monthly report contains a figure that never happened is a separate question, and the answer is a gate rather than a tier: of the three fabrications we caught, the worst came from the top model. We use one test to sort every task: drop the model one tier and see whether the business outcome changes. None of the three incidents broke the deliverable. All three broke the evidence about the deliverable. Four questions at the end tell you which tier your money buys.

"AI costs are collapsing" and "AI capacity is sold out" appear on the same day, and both are accurate. The first describes the price of reproducing a capability that already exists. The second describes the price of the capability that does not exist yet, plus the chips, electricity, and queue time around it. Our full report on the two prices of AI lays out the source evidence; this post is the operator's version, for anyone paying an agency that uses these tools on their behalf.

The cheap tier is very cheap

The cost of GPT-3.5-level intelligence fell from $20 per million tokens in November 2022 to $0.07 by October 2024, a more than 280-fold reduction in about 18 monthsSource: Stanford HAI, AI Index 2025

Epoch AI tracked the same question across six benchmarks and found prices "declining between 9x per year and 900x per year, with a median of 50x per year." The energy per task is falling at the same pace: the IEA reports it "dropping by at least an order of magnitude annually." Classification, tagging, formatting, pulling numbers into a table, first drafts of variations: all of that can be done by models that cost cents per client per task. A full technical crawl of one client site through our data provider costs $0.54. A quarterly AI-visibility probe costs about $0.303 per client per quarter. Any agency still pricing that work as if it were scarce is keeping the 280-fold drop for itself.

The frontier tier is rationed, in writing

The same vendors that give away last year's model meter this year's, and they say so on their own pricing pages. OpenAI sells the same GPT-5.6 Sol tokens four ways: Batch and Flex at half price for waiting, Standard, and "Fast mode" at double. A product with one marginal cost does not need a 4x price spread for identical output. Anthropic's documentation states that its limits are "maximum allowed usage, not guaranteed minimums," and pauses API usage once a tier's monthly spend cap is reached. Google's free Gemini tier is paid for in data: its feature list reads "Content used to improve our products," and the paid column for the same line reads "No."

"Customer demand continues to exceed available capacity", with quarterly capital expenditure of $41 billion and a guide of "over $50 billion" for the next quarterSource: Microsoft FY2026 Q4 earnings call, 2026-07-29

Nine months and two record build-out quarters earlier, Microsoft had said the same thing. The constraint did not clear. Total AI spend keeps rising while per-task cost collapses because users move up the quality ladder: the IEA notes that video, reasoning, and agentic tasks "can consume hundreds or thousands of times more energy per query than simple text generation," and AI data-centre electricity use rose 50% in 2025.

The downgrade test

The commodity-or-frontier question is a property of the task, not of the model's name. The test we apply to every piece of agency work before it gets a model assigned:

Replace the model with one tier cheaper. If the business outcome is unchanged, the task is commodity work and should run on the cheapest model that clears the bar. If the outcome degrades in a way that costs money, the task is frontier work, and the question becomes how much verification to add, not whether to pay.

Six things decide it in practice. What happens to the output if quality drops one tier? How easily would an error be noticed? What does an unnoticed error cost? Does speed create value here? Can the work be batched or cached? Would several vendors be acceptable? Low stakes, easy detection, and batchable work route down. High stakes, hard detection, and a decision that moves money route up.

Our own routing rule puts it in one line: the fastest path is "the smallest model plus lowest effort that still succeeds, then escalate on failure," with the caveat that "the cheapest models take 2-3x the turns on multi-step work, costing more overall." Cheap is a starting point, and a reconciler has to sit behind it.

What actually breaks, on every tier

In our experience a downgrade rarely breaks the artifact. It breaks the evidence trail. The worst case we caught was on the flagship.

A monthly client summary reported campaign spend of $9,789.30 at a $489 cost per acquisition. The real figures in the source report were $13,002.74 at $260.07. The model had lifted the numbers from a strategy document dated ten days earlierSource: Choice OMG incident record, July 2026; drafter on Opus 4.7, the flagship model at the time

A prompt rule forbidding figures not in the report did not stop it, and neither did the model tier. What stopped it was a mechanical gate: every dollar amount in a generated summary now has to appear verbatim in the source report, or equal an exact sum of report amounts on a line that says "total," and the build fails otherwise. Rounded restatements fail on purpose.

The other two incidents were on the cheap tier and had the same shape. A condensing layer with no review gate produced a client document with dollar figures that existed nowhere in its inputs and internal ticket references that should never leave the building. A cheap worker given a nine-line verification script ran it correctly, saw it fail, and wrote a report saying it passed. In all three, the model handled the mechanical part and failed at synthesizing evidence about its own work, which is the part a reader relies on. The tier changes how often that happens; only a gate changes whether it ships. The cost that token prices removed came back as verification, and in our pipeline that is where the money went.

Four questions to ask the agency you pay

  1. Which tasks on my account go to a cheap model, and what checks them? A good answer names the tasks (classification, reporting pulls, bulk analysis, draft variations) and the mechanical check behind each one. "We review everything" is a person, and a person did not catch the $9,789.30.
  2. Which tasks get the best model, and why those? The answer should be a list of judgment-heavy, client-facing work: budget decisions, anomaly diagnosis, strategy, the narrative in your report. If everything goes to the best model, you are being charged frontier prices for commodity work. If nothing does, the judgment calls on your account are being made by a model chosen for price.
  3. Does any dollar figure in my report get checked against its source by a machine? Yes or no. A prompt instruction does not count, and neither does "we use the best model"; our worst fabrication came from it.
  4. When a model gets cheaper, what happens to my fee? Commodity savings should show up as more work inside the retainer or a lower price, and the agency should be able to say which (flat fees make that answer easy to check). Frontier work will not get cheaper on the same schedule, and an honest agency will say that too.

The saving from the 280-fold drop goes to whoever is closest to the scarce input. For an agency, the scarce inputs are your outcome data, the verification around the model, and the person who signs the report. Ask which of those your retainer is buying, and you will know which price you are paying.

Need Help With Your Digital Marketing?

Book a free discovery call with our team.

Get in Touch