The AI cost conversation is happening in the wrong part of the organisation
Most organisations have filed AI spend where they file every other technology cost. It sits with IT, it is reported as a licence line and a usage line, and the questions put to it are procurement questions. How many seats. Which tier. What can we negotiate.
Those are fair questions and none of them will tell you what this is actually costing.
Token pricing does not behave like a licence. A licence is a fixed price for access. Tokens are a variable price for use, and the variable is not how much someone works. It is how well they work. Two people on the same seat, doing the same job, can differ substantially in what they consume and in what they produce. That variation is a distribution of workforce capability, and it has been recorded as an infrastructure cost.
What is actually driving it
A token is a fragment of language, roughly three quarters of a word, and it is the unit everything is counted and charged in.
An assistant carries nothing forward through a conversation by itself. On every turn it reads back everything that came before, and every one of those words is bought again. A long, cluttered conversation is not simply slow to work with. It is repurchased in full at each exchange inside it.
Set against that, someone who has written down who they are once, and saved it, begins every task with the system already briefed. Someone who has not writes that brief again from nothing each time, receives an average draft, and then pays for the rewriting.
That is the whole mechanism, and it explains something that sounds implausible until you see it: the quality and the cost move together rather than against each other. Almost every cost reduction I have run inside an organisation traded something away. This one does not, because the answer being poor and the work being expensive have a single cause.
Why this belongs to People & Culture
The behaviours deciding the bill are ordinary working behaviours. Whether someone briefs properly. Whether they say what good looks like before asking for a draft. Whether they correct early, or let a weak first draft compound through everything built on it. Whether they recognise the point at which a conversation should be closed and a fresh one started.
None of that is configured. It is taught, modelled and reinforced, which is People & Culture work, and it appears nowhere in a procurement discussion.
The trap, and it is worth naming before someone falls into it
The instinctive executive response is to make usage visible per person and compare. It is the wrong move, and it is expensive.
Once people believe their consumption is watched and ranked, they use less, more cautiously, on lower-value work. Under use costs an organisation far more than any wasteful prompt, and it shows up in no report at all, because work that was never attempted leaves no trace.
Measure the capability instead. Whether people have set themselves up. Whether the standard of a first draft is rising. Whether rework is falling.
Where I would start
Give people two habits before you give them anything else. Writing down who they are, once, so that no task starts from nothing. And holding a conversation to one task, so that nothing is carried along and bought again.
Both are free. Both are visible within a day. Both make every later part of an AI programme cheaper to run.
The larger work still has to happen: accountabilities allocated, capability built deliberately, and rhythms that let people learn inside an ordinary week. But the sequence matters. Starting with the habit a person can adopt this afternoon earns the goodwill the harder parts will need.
I keep both habits on a single page. It is the card I use in the organisations I work with, and it is yours to print and pass on.