Your AI Cost Model Has an Expiry Date on It: September 1
๐Ÿ“ข
← Back to Blog

Your AI Cost Model Has an Expiry Date on It: September 1

John Aspinall · · 9 min read

If any of your Amazon catalog work runs on Claude's Sonnet tier, the unit cost you validated it against stops existing in 24 days โ€” and I'd bet nothing in your business is going to tell you when it happens.

Anthropic's published pricing page states it directly: Sonnet 5 is on introductory pricing of $2/$10 per million input/output tokens through August 31, 2026, after which standard pricing of $3/$15 takes effect (platform.claude.com pricing docs). Batch moves the same way, $1/$5 to $1.50/$7.50. That's a flat 50% increase on the tier Anthropic's own documentation recommends for "most production workloads."

The increase is a procurement footnote. The pattern underneath it is the story.

Model pricing now behaves like a SaaS trial

Somewhere in the last year, introductory pricing became a normal part of how these products launch. Sonnet 5 shipped at a promotional rate with a published end date, and a lot of operators โ€” me included โ€” did the arithmetic on whether a job penciled catalog-wide using that number.

Here's what makes it different from every other cost in your business. When Amazon raises a fee, you get a notification with an effective date, and a hundred people write about it within the hour. When a model's promotional pricing lapses, nothing happens. No email. No error. No changed behaviour. The job runs exactly as it did the day before and the invoice arrives three weeks later at a number nobody compares to anything.

A price that expires quietly is a price you have to diary, and I don't know a single operator who has "check whether the promo ended" on a calendar.

I'll take my share of this. On July 27 I told you to pin your model version on anything that writes to a live listing, so an upgrade becomes a Tuesday decision instead of a silent behaviour swap. I stand by every word of it. But pinning is a contract on behaviour, not on price. The exact string you pinned to protect yourself got repriced on a published schedule, and the pin did nothing about it. Those are two separate risks and I only wrote about one of them.

Why most people will read this wrong

"The cheap era is over." Eight days ago OpenAI cut a tier by 80%. Today one model on one tier comes off a promotion. Prices in this market move both directions, monthly. If you build a plan around the direction of a single price change you'll be rebuilding it before Q4.

"Just switch to a cheaper model." The migration is a config string. The expensive part is re-validating that the replacement still honours your guardrails โ€” leaves a regulated field blank when it can't source it, respects a character ceiling, refuses to infer a claim from marketing copy. If you don't have a fixed set of real SKUs you can re-run, you can't evaluate the switch at all, and you'd be swapping the model under a live catalog write path to save an amount of money I'm about to quantify. Bad trade, made confidently.

And one comparison to be careful with, because I wrote a whole post about it in May: don't compare cost per token across model generations. The same pricing page notes that Claude 4.7 and later models use a newer tokenizer producing roughly 30% more tokens for the same text, with Sonnet 4.6 and earlier on the previous one. Anthropic hedges that honestly โ€” it depends on your content and workload shape โ€” and you should carry the hedge. I covered the mechanism when Opus 4.7 shipped and won't re-argue it here. The practical version: Sonnet 5 at standard pricing carries the identical posted rate as Sonnet 4.6 and bills meaningfully more tokens for the same block of listing copy. Two different things moved. Only one of them is on the rate card.

What actually changes at $200K a month

Let me be honest about the dollars before anyone gets excited, because I argued eight days ago that token cost was never the agency line item and I'm not quietly reversing that today.

For a brand doing $200K a month, the AI spend attributable to your account inside a $6K retainer is realistically low hundreds of dollars. A 50% increase on part of that is tens of dollars a month. If you reopen a contract over this, you have spent political capital to recover lunch โ€” and you'd rather have that capital in October when peak surcharges, Q4 CPCs and holiday storage all land at once.

Where it genuinely bites is narrower: the catalog-scale tail work that was marginal to begin with.

The jobs that live on the Sonnet tier are the ones you run across everything โ€” bulk copy passes, review mining across 40 ASINs, attribute completeness audits, A+ module copy for the long tail, search-term classification, Q&A drafting. High token volume, low value per call. That's exactly the class of work where somebody ran the numbers this summer, decided it finally penciled catalog-wide, and scheduled it.

A chunk of those decisions un-pencil on September 1, and the ones that un-pencil first are the tail SKUs. Which is the wrong outcome, because tail coverage is what clears the Premium A+ eligibility gate and what keeps a long tail legible to the AI layer deciding whether your products belong in a consideration set. Skipping the boring catalog work because a per-token rate moved by a dollar would be a genuinely stupid way to lose ground.

It doesn't have to happen. The levers are documented and most people still haven't turned them on:

Batch pricing at the new standard rate is $1.50/$7.50 โ€” below the promotional base rate you're paying today. An A+ copy pass across 80 ASINs does not need to come back in 400 milliseconds. Almost none of this work is latency-sensitive and almost none of it is batched. Stack prompt caching on top, where cache hits bill at a tenth of base input, and a catalog job โ€” one fixed instruction block plus a varying SKU record, repeated 200 times โ€” absorbs the September increase with room left over.

I've made the caching-and-routing argument before and I'm not claiming it's new. What's new is the arithmetic: for batchable catalog work, the September increase is fully recoverable inside the pricing structure itself, and the operators who'll feel it are the ones running everything synchronously because that's how the first version worked.

The second place it lands is your vendor stack. Every SaaS tool with an "AI-powered" feature is buying tokens from someone, and any of them who priced a plan off a promotional rate has the same date on their calendar โ€” or doesn't, which is worse.

What I'd do this week

1. Diary the expiry dates, not just the prices. This is the actual lesson. For every model your business depends on, write down the model string, the current rate, and whether that rate is promotional or standard. If you can't tell, that's the finding. A cost model with no expiry column is how you get here again in November.

2. Produce the list of which model string each automation runs on. Not "we use Claude." A model ID per workflow, written down, sorted into pinned versus floating. Most operators can't produce this, and you can't price a change to something you can't name.

3. Measure your actual token counts instead of computing them. Run 20 real SKUs through your real prompt, record actual input and output tokens, do it again in September. A ~30% tokenizer figure is a starting estimate, not your number, and this is a five-minute measurement against the golden set you should already have.

4. Batch what can be batched before you renegotiate anything with anyone. Fix your own cost structure first. You may find there's nothing left to argue about, which is the cheapest possible outcome.

5. Ask your agency and your SaaS vendors one written question: does your pricing change on September 1, and which model and tier are you running? In July I laid out the three pass-through postures โ€” absorbed into the fee, billed separately, or bundled into a tech charge. This is the month you discover which one you signed. A vendor who answers with a model name and a date has thought about it. One who replies with a paragraph about their commitment to AI innovation has not, and that reply is itself the answer.

What I'd ignore

Benchmark tables. Sonnet 5 against anything. Real evals, none of which measure whether a model writes a bullet that won't get your listing suppressed. That eval is the fixed set of your own SKUs, and it's still the only one that bills you when it's wrong.

The "AI bubble pricing" discourse. One promotion ended. This will be a genre for about a week and it contains zero decisions.

Any invoice that shows up with a new "AI surcharge" line and no arithmetic behind it. Ask which model, which tier, what volume. The honest vendors can show you. September is going to be an excellent month for surcharges with very little relationship to anybody's actual token bill.

The urge to switch stacks over this. Migration is cheap and re-validation is not, and the money at stake for a single brand is a rounding error against the cost of shipping an unvalidated model into a live catalog.

Three times in five months the economics under these workflows have moved โ€” a tokenizer in May, a competitor's price cut in July, a promotion lapsing in September. Not one of them arrived as something you'd notice while it was happening. The only durable fix isn't picking the right model. It's owning a list of what you're running, what it costs, and when that number stops being true.

Sources: Anthropic pricing documentation โ€” introductory pricing through August 31, 2026, standard pricing from September 1, 2026, and the tokenizer note for Claude 4.7 and later models. Introducing Claude Sonnet 5.

Put AI to work inside the business you already run.

The Aspi OS Bootcamp is a 4-week live build: second brain, Claude Code workflows, Codex execution — on your real business. The next cohort is forming now.

Get first access →

Not ready? Get the free newsletter — the AI workflows I actually ship, when they're worth your inbox.