An invoice from your agency arrives at the end of the month with a line on it that reads, roughly, Token Usage (Opus): 34 units × €3.00 = €102. The rest of the invoice you know how to read. Human hours have been on invoices for a hundred years. This line is new, and by design it does not correspond to a person sitting at a desk. So the question every buyer is going to ask their agency this quarter is the practical one, not the theoretical one: what does a unit mean, and how do I tell you are counting them honestly.
We started charging clients for AI usage in September. The pricing framework is settled and published: what an invoice line item looks like, why the unit is what it is. What we have not written down yet is the smaller, more concrete question underneath all of that: what is an hour when part of the work is happening in a language model, and how did we manage to publish a broken answer to that question to two real clients before catching it. The answer to the first half is a rough physical calibration you can carry in your head. The answer to the second half is a category error in our accounting system that inflated one client’s reported hours by 183% for four days.
The reason this matters for a buyer is not that our accounting bug is interesting. It is that the same category error will show up in every agency invoice that starts carrying AI cost this year, and the buyer’s leverage is the ability to spot it on the report.
An AI hour is not a person hour, and pretending otherwise breaks the invoice
The unit we bill on is weighted tokens. Input tokens plus five times output tokens, at exactly 100,000 weighted tokens per unit. The five is not arbitrary. Anthropic’s current published rate card prices output at exactly five times input on every current model: Opus 5.5 at $4 and $20 per million, Sonnet 5.5 at $2 and $10, Haiku 4.5 at $1 and $5. That single ratio collapses input cost and output cost into one number a client can be billed on, and if Anthropic ever breaks it we will have to redesign. Until then, one unit at cost on Opus is $0.40, and at our published retail multiplier of two, €0.80.
The problem with all of that, for anyone whose job is to read a report, is that it is not an hour. So when a project management tool asks you to enter effort as hours, and you have some AI cost you also want to show against the same project, the temptation is to do the conversion.
We did the conversion. From our own session logs over two intensively used projects, the empirical calibration works out to be surprisingly stable: our assistant processes roughly 130 weighted tokens per second, roughly 8,000 per minute, roughly 480,000 per hour. So one unit is about twelve to thirteen minutes of active AI work, and one hour of active AI work is about 4.8 units. On sessions with long human pauses the ratio collapses and the measurement gets useless, but on continuous work it is a stable enough number to carry in your head. One AI hour of billable work is about €4 at cost, €8 retail. One human hour on the same projects, at our senior rate, is about €90. Ninety versus eight is the actual gap you are looking at.
That number, twelve minutes per unit, was what we wanted a client to be able to picture when they saw 34 units on a line. What went wrong was not the calibration. It was believing the calibration justified filing the number as an hour.
What we tried, and what each attempt broke
Three attempts to put AI cost on a report that already knew how to display human hours. Each of them worked in isolation. Each of them failed once a real weekly report tried to render both together.
Attempt one: write token units as time-records against a synthetic AI job type. Cheap to build, five lines of change in the tracker adapter, worked immediately. The token-usage entries showed up under the task, and the weekly report summed them along with the human hours. That was the failure. The report summed them. The report said 43.86 hours had been logged against a project whose actual human input for the same period was 15.51 hours. The manager reading it saw a phantom overrun of roughly twenty-eight hours. The overrun was not there. What was there was 23.67 units of Token Usage being counted as if they were hours, plus €4.68 of API expense being counted as if a euro was an hour, plus 15.51 real human hours underneath. The report was not lying about the data. It was doing what its query told it to do, which was to add up a value field across records that had different units in them.
Cost of the attempt: four days of misreporting to two clients. The four days matter because the token records started on 4 September and the next weekly report went out on the 7th. The report that went to the client showed a project running well over its committed hours when it was well inside them. Nobody complained, which is worse than someone complaining.
Attempt two: keep the time-records but switch the reporting endpoint to one that returns a job type. The tracker had a broad tracking endpoint that returned everything logged against a project. We had picked it because it returned everything in one call. It also returned expenses in the same list as time records, and it did not return the job type field, so the code that consumed it could not tell a token-usage entry apart from a senior developer’s hour. Switching to the per-project time-records endpoint gave us the job type back, at the cost of a second call per project and a mild schema divergence. It solved half the problem: human hours were now separable from AI usage. It did not solve the units problem, because we were still writing tokens as fractional hours into the same field.
Cost of the attempt: about a day of work, and a false sense that we had fixed it. The report now correctly said 15.51 human hours. It also, on a separate line, said 28.35 units of Token Usage. Those two numbers together added up to nothing. The client asked, reasonably, whether the sum meant anything, and we did not have a good answer.
Attempt three: stop pretending. Move all resource consumption off the time-record surface and onto the expense surface. This was the fix that stuck. Time on a project means human minutes. Everything else, tokens, external API calls, per-lookup enrichment fees, becomes an expense recorded in the provider’s own billing unit, at its own line, tied to the same task. Twelve expense categories were opened to cover the resource types we already bill on. Fifty-eight records were migrated across two projects, €86.51 total moved from the hours column to the expenses column. The weekly report grew a new table titled AI and resource usage, and the human-hours totals grew a footnote saying that is what they are.
Cost of the attempt: a week’s build and a set of accounting migrations. Every time-record we deleted had to have its replacement expense created first, then verified, then the original removed, so that a report running mid-migration would see the sum once, not twice or zero. One trap worth naming: the tracker’s time-record value field truncates silently to two decimal places, which meant our old fractional-hour token entries were less precise than the underlying unit measurement. The expense surface takes value as currency, which gives us two decimals against a euro instead of two decimals against a rate. That is a real precision gain, not a bookkeeping preference.
The rule the migration produced, and why it holds
Every line on a client’s report answers exactly one question, and no line answers two. Human hours are one question, in one unit, which is a minute of a named person’s attention. Token usage is a different question, in a different unit, which is 100,000 weighted tokens on a named model tier. External API calls are a third question, priced in whatever the provider charges. The moment a report puts two of those questions into one column, someone reading it in a hurry will misread.
The other rule we built in, which we should have built in from the start: an unclassified job type does not get summed into anything. If a new resource type appears on a project and nobody has told the reporting layer whether it is measured in hours or in a resource unit, the report shows the raw number and a warning, not a zero and silence. Loud incomplete beats silent wrong every time. Our first version had it the other way round, and that is exactly the shape of failure that let 23.67 token units and €4.68 of API expense parade themselves as hours for four days.
What we would do differently
Two things, both about sequencing. First, we would have written the client-facing weekly report sentence before we chose the data model. If we had drafted, in plain language, the paragraph a client would read on Monday morning, we would have seen the ambiguity in you used 43.86 hours this week before that sentence appeared under a real project.
Second, we would have separated the two surfaces at the source rather than trying to reunite them in the report. Reusing the time-record surface for machine cost saved a week of build and cost a client conversation that took longer than the week saved. There is a general lesson buried in that: any time you are tempted to make an existing bucket carry a new kind of thing because the tooling already knows how to display it, the display is what will punish you first.
How a buyer reads the number, on a Monday, without a call
The reader test we hold ourselves to now is that a client should be able to open one week’s report and know, without picking up the phone, three things: how many human hours went into the work and by whom, how much AI resource was used and on what task, and how much external API spend was passed through at cost. If any of those three answers requires a follow-up call to disambiguate, the report is broken, not the reader.
Which gives a buyer looking at their own agency’s next invoice four small questions worth asking. What unit is the AI line item in, and is the multiplier stated. Is that unit stable across models or does it silently change when a supplier switches you from one model to another. Are external API fees passed through at the provider’s own unit or blended into the AI line. And, the one we would ask first: could you read one week’s report and tell what your team did versus what your agency’s software did, without a phone call to explain it. If the answer to the last question is no, the framework is not honest yet, whatever it says about being transparent. And if the answer to any of the first three is we bundle it into hours, the invoice you are looking at has the same category error we published to two of our own clients on 7 September.
For an organisation choosing an agency this quarter, the question of how they bill their AI work is not just a procurement detail. It is the shortest test we know of for whether the people you are hiring have thought about the operational half of what they are selling. A supplier who cannot state, on request, what unit an AI hour is in, is not a supplier whose invoice you want to be trying to read in December. If you would like ours to sit next to whichever one you are looking at now, that conversation starts here.
Last modified: September 30, 2026
United States / English
Slovensko / Slovenčina
Canada / Français
Türkiye / Türkçe