🇹🇷 Türkçe: Bu yazının Türkçesini oku →

Here is the awkward thing about paying an agency in 2026: the tools we use to do a decent chunk of the work now come with a meter on the side. Every draft, every rewrite, every product-description refresh runs a small counter, and that counter turns into a bill from Anthropic or OpenAI at the end of the month. On most agency invoices you cannot see any of that. There is a line called “content refresh, 8 hours” and a total. The meter is still running, someone is paying it, and if it is not on your invoice it is being absorbed somewhere silently.

I have spent the last quarter arguing with our own project management tool over what the honest way to show this on a client invoice looks like. Not because we love new line items, but because “we absorb it” is a promise that breaks the moment a new model has a different pricing shape or a client asks us to do three months of what used to be one month of work. The absorption ends. The question is where.

The person who has done the most public thinking about this problem is not an agency owner. It is Simon Willison — independent developer, co-creator of Django, maintainer of the open-source llm tool and the llm-prices.com calculator. His blog has been the running commentary on token pricing for two years. I went and read it in order so you do not have to, and then borrowed his frame to think about what a buyer should be asking their supplier this quarter.

The catch-22 sitting behind every fixed-fee agency contract

Willison put the shape of the problem plainly at the end of May, in a piece called I think Anthropic and OpenAI have found product-market fit: “Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by their staff.” He then did the exercise on himself. Running ccusage against his own laptop for the previous 30 days, he reported “$1,199.79 for Anthropic Claude Code” and “$980.37 for OpenAI Codex” — for one person, one month, if he had been paying API rates rather than a subscription.

That is a single senior developer’s personal usage. Multiply it across a team of ten and you are past $250,000 a year in token spend before anyone has approved a budget line for it. The reason nobody sees this on their invoice yet is that most agencies are still on subscription plans — the $200-a-month Anthropic Max plan, or the equivalent from OpenAI — which cap the visible cost. But the subscription model itself is under pressure. Willison notes that “as of April 2026 the ‘Enterprise’ cost for both OpenAI Codex and Anthropic Claude Code/Cowork is the same as the listed API price.” Translation: the moment your agency’s usage crosses the enterprise threshold, the subscription flat rate goes away and the meter comes back.

Uber has already priced this in. Most agencies have not.

The most visible corporate response so far came on the 2nd of June, when Bloomberg’s Natalie Lung reported that Uber had capped every employee at $1,500 in monthly token spending per AI coding tool, after the company burnt through its full 2026 AI budget in four months. Willison, writing about it the next day, called it “a rational policy response to over-spending” and did the arithmetic: at two tools per engineer, that is a $36,000 annual cap per head. Against a $330,000 median engineer compensation from Levels.fyi, roughly eleven per cent of comp is now formally allocated to a metered software line.

Uber is an internal buyer, not an agency. But the shape of the decision is exactly the one your supplier is quietly making about your account. Somewhere in their finance meeting this month, someone is looking at a graph and deciding either to cap the usage, to raise fees quietly, or to keep absorbing until it hurts. If your fixed-fee contract runs into next year, the third option is the one that puts your supplier under stress. That is not a stability you want.

The three levers a buyer should hear about, and one that isn’t working yet

There are three ways an honest supplier can make the meter visible on your account, and there is a fourth that everyone is still fumbling with. Ask your supplier which of these they are actually doing.

Lever 1 — Show the meter, at the provider’s own unit

Not “AI usage: £400”. That is a fudge. What you want is a line item that carries the provider’s own billing unit — tokens for Anthropic and OpenAI, characters for Google, minutes for a transcription service, images for a generative image model. Cost is still cost, but reading the invoice you can now see whether the number is high because you asked for more work or because a new model is more expensive per unit.

Willison has been running exactly this exercise in public for two years. His llm-prices.com calculator lets anyone put an input-and-output token count in and get a cost back for every current model. That the tool exists, and is used by working developers rather than accountants, tells you the unit itself has stabilised. Ask why your invoice does not use it.

Lever 2 — Separate the labour that stays yours from the meter that is theirs

The most confused invoices I see try to convert AI usage into “AI-assisted hours” and put it in the same time-record column as human hours. This looks tidy and quietly corrupts both numbers. A human hour on a senior consultant is a rate; a token spend on Claude Opus is a wholesale price at $15 and $75 per million tokens in and out, respectively (Anthropic’s public list, as reproduced on llm-prices.com). The two numbers do not share units and they do not move together.

The clean version keeps human hours as human hours, and lists each metered service on its own line at the provider’s own billing unit, at the agency’s own margin. If a supplier tells you they cannot do this, what they mean is their accounting system will not carry a variable-priced non-hour line. That is fixable software; it is not a good reason to hide the meter from you.

Lever 3 — Buy your own cap, at the buyer’s level, before your supplier does it for you

This is the Uber lesson translated for a buyer of agency services. Rather than accept an open-ended “we’ll use AI where it helps”, agree a monthly ceiling for token spend on your account, expressed in your currency, and require the supplier to notify you before it is crossed. Willison, again on the Uber piece: “A $1,500 monthly limit per tool strikes me as a rational policy response to over-spending.” What is rational for Uber’s engineering org is at least as rational for a marketing budget owner whose ad-copy refresh could quietly rack up token spend on a new agent-based tool.

The related warning from Willison’s August piece on GitHub Models is worth reading if you have never priced this yourself. When Microsoft retired its free GitHub Models offering, Willison observed that “coding agent patterns made it prohibitively expensive to offer free or subsidised tokens.” Agent patterns — the ones that let an AI keep working on a task without you standing over it — consume tokens at a rate that broke a subsidised free tier inside months. The same patterns will show up in content and commerce work over the next year, and the same math will happen.

The lever that isn’t working yet — retainer clients on real-time metering

Here is the honest part. Nobody I have read, Willison included, has a clean answer for how to handle metered AI cost inside a monthly retainer. Retainers exist because the buyer wants predictability and the supplier wants a floor. Putting a variable metered line inside a retainer either turns it into consumption billing (defeats the retainer’s purpose) or gets absorbed into a fudge factor at the top (defeats the meter’s purpose). The workable interim, which is what we do, is to fix the retainer against a stated monthly cap of AI units and true-up quarterly against actual usage — with the client seeing both numbers. It is not elegant. It is honest, and every supplier who has moved off “we absorb it” has landed on some variant of the same compromise.

Willison’s most recent piece on the September price war between Anthropic’s Opus 5.5 and OpenAI’s new GPT-6 tier makes the point in one line: “A new model would arrive at the same price (or cheaper!) and paper over most of your problems.” He is only half joking. The cost curve does keep bending down. But bending down is not the same as bending predictably, and predictability is what a retainer is selling. Until models stop shipping every eight weeks with a different pricing shape, the retainer contract cannot pretend the meter isn’t there.

What this looks like inside a real WooCommerce build

Ecommerce is where this hits first, because product content is the piece agencies are most aggressively automating. A 5,000-SKU catalogue with a quarterly refresh cycle — new copy, translated variants, SEO title regeneration — is a real, sizeable, measurable AI cost line, not a rounding error.

The build we run for that: WooCommerce stays the system of record. Product descriptions live on the product object and its post_meta, nothing exotic. A small worker service — a WP-CLI job on the same host, or an external cron on a lightweight VPS — pulls the SKUs that need refreshing, sends the current copy plus the brand style rules to Claude Sonnet or GPT-6 Luna, and writes the returned copy into a revision on the product. Editors approve before publish; nothing goes live unread. Every call logs the input and output token counts to a small `wp_ai_usage` table keyed by product ID and job ID, with the model name and the timestamp. At month end, that table gets aggregated into two figures: cost by client project and cost by model. Those two numbers land on the invoice as their own lines, at cost times our margin, in the model’s own units.

Effort to build the first time: about two weeks for a competent WooCommerce team, most of it in the editorial approval workflow, not the API call. The API call is trivial. The interesting part is the boundary — deciding which SKUs are safe to auto-refresh, which need an editor to look, and which need the copy team to write from scratch. That boundary is what the buyer is really paying for, and it is where the human judgement earns its rate. The meter runs alongside; it does not replace it.

The trap we hit and would flag for anyone building this: the temptation to route every product through the top-tier model because it is “safer” quietly turns a €400 job into a €2,000 job. Route by risk. High-value SKUs, regulated categories, and anything with legal claims go through the top model with human review; tail-of-the-catalogue commodity SKUs go through the cheap model with spot-checked review. The routing decision is the single lever with the biggest effect on the metered bill.

The one question worth asking your current supplier this quarter

Not “are you using AI?” — everyone is, and the honest answer is boring. The question is: “show me what my account has cost you in AI tokens over the past ninety days, by model, at the provider’s own list price.” A supplier who can produce the figure within the day is running the meter properly and has thought about how to bill you for it fairly. A supplier who cannot is absorbing it, which sounds generous until you remember that absorption ends the moment their business math stops working, at which point either your fee goes up quietly or your project quality goes down quietly, and neither of those is a conversation you want to have via a surprise.

If you would like a hand designing the line-item version — the one that shows the meter honestly without turning your invoice into a bill-of-materials — that is the shape of the work we do, and the conversation we would rather have with you now than after your supplier’s fudge factor gives out. Simon Willison’s public working is a good place to start; the version that lands on a real invoice is the piece we would build with you.

Leave a Reply

Your email address will not be published. Required fields are marked *

Close Search Window
↑