I recently read a colleague's paper on AI tokenomics, and it was very good. Clear, practical, and grounded in the awkward reality that intelligence may be getting cheaper, but it still arrives with a meter attached.
Naturally, I responded to this thoughtful piece of work in my usual ADHD fashion: by thinking about six adjacent problems at once.
Token cost led me to capability. Capability led me to usability. Usability led me to whether an AI tool is actually fit for the work people are trying to do, rather than merely impressive in a demonstration. Before long, I was wondering how an organisation should choose between frontier models, cheaper 'good enough' tools, and the systems that already live inside its corporate data.
Which led me somewhere less obvious.
Enterprise AI strategy is not decided only in vendor briefings, architecture forums, or carefully formatted recommendation papers. Quite a lot of it is decided at 9 pm, on personal accounts, by the people who will eventually write those papers.
I think of this as the garage hours effect.
The decisions before the decision
A senior engineer, architect, analyst or technically curious executive spends an evening building something on their own time. It might be a side project, a home automation experiment, a small application or a research workflow. In my case, it is often ‘some quick research’ for a book idea. By midnight, I have drafted the book overview, written chapters one to three, added code to my latest Arduino project, and assembled an agentic AI workflow that combines a local LLM with several frontier models to create my own AI council. ‘Quick research’ is another of those dangerous phrases. It has consumed entire weekends.
The practitioner is not running a formal evaluation. There is no weighted scorecard. Nobody has formed a steering committee or booked a room called Endeavour.
They are simply trying to get something done.
That is precisely why the experience matters.
They discover which tool understands a complicated instruction without being coaxed through it six times. They notice which one can work across a large body of material, which one loses the plot after three turns, and which one interrupts a perfectly ordinary task with a warning that makes them feel as though Legal has joined the chat.
They also notice cost. Not in the abstract sense of a procurement spreadsheet, but in the immediate sense of whether the meter is moving faster than the work.
Over time, these small encounters form a view. Six months later, that same person is asked to contribute to an AI strategy, assess a platform, or draft a recommendation for the CTO. The paper may contain sensible headings such as capability, security, integration and commercial fit. Underneath them sits something less tidy and often more influential: lived experience.
The formal decision is still important. So are security, privacy, support, contractual protection and the unglamorous matter of whether the vendor will answer the telephone when something breaks. I am not suggesting that personal experimentation should replace enterprise due diligence. That would be shadow IT wearing a thoughtful expression.
But ignoring garage-hours experience is equally foolish. By the time a preference appears in a recommendation paper, it may have been forming for hundreds of prompts.
The terminal often decides before the boardroom knows there is a decision.
Four questions hiding inside 'Which AI is best?'
Organisations still have a habit of treating AI selection as a horse race. Which model is leading? Which vendor has the largest context window? Who won the latest benchmark? Which logo should appear in the target-state diagram so everybody can go to lunch?
Those questions are not useless. They are simply too compressed.
When somebody asks, 'Which AI tool is best?', I now hear at least four different questions:
- Capability: Can it reason well enough to perform the task?
- Usability: Can a real person get that capability without becoming an amateur prompt engineer?
- Context: Can it reach the information required to do the work?
- Fitness: Is the combination of quality, cost, control and speed appropriate for this particular job?
A tool can be excellent on one dimension and frustrating on another. A very capable model can be awkward enough that people quietly stop using it. A pleasant, well-integrated assistant can produce beautifully formatted shallowness. A cheaper model can be exactly right for a routine task and an expensive mistake for a consequential one.
This is why benchmark arguments tend to become less useful as they approach an actual workplace. Organisations do not buy abstract intelligence. They buy completed work under constraints.
The constraints are where strategy lives.
Three economies of work
The original version of this argument had two categories: the depth economy and the volume economy. That is a useful start, but it misses the part of enterprise work where access to the right information matters more than raw reasoning power.
There are really three territories.
Depth
Consequential work where weak reasoning travels.
Volume
Routine work where throughput and cost matter most.
Context-bound
Work whose answer depends on access to organisational facts.
The depth economy
This is the work where subtle errors travel.
Architecture decisions, regulatory analysis, long-form structured argument, complex design, sensitive risk work and the document whose weakest paragraph will be found by the most senior person in the room.
Here, model quality compounds across the task. Not because a small benchmark lead magically becomes a scientific law, but because sustained work gives shallow reasoning more opportunities to wander. The model must hold constraints, distinguish evidence from inference, notice contradictions and preserve a line of thought across many steps.
This is where frontier capability can earn its premium. The important word is earn.
Paying for the most capable model is sensible when the cost of weak reasoning is larger than the cost of the model. It is less sensible when the job is to turn a meeting transcript into five bullet points and remind everyone that Martin still owes Finance a number.
The volume economy
This is the great mass of routine work that fills a working day: first drafts, summaries, classifications, ticket triage, standard code, extraction, formatting and administrative text that needs to be competent rather than profound.
Quality still matters, but it behaves differently. Once the output clears the bar for the task, further capability may add little practical value. Nobody reads a summary of a routine email thread and wishes it had engaged more deeply with the human condition.
For these tasks, a smaller or cheaper model may be the better engineering decision. Faster can matter more. Predictable can matter more. Cost certainly matters more when the same operation runs thousands of times.
The premium model does not disappear. It stops being the automatic default.
The context-bound economy
Then there is work where the central problem is not reasoning but access.
'What did this project team actually decide?'
'Which version of the policy is current?'
'What commitments have we made to this client across email, meetings and documents?'
A brilliant model without the organisation's information cannot answer those questions. It can produce a polished approximation, which is sometimes worse because approximations arrive wearing a tie.
This is where tools embedded in the corporate estate have a structural advantage. Microsoft 365 Copilot, for example, can ground prompts in information the user is permitted to access through Microsoft Graph. Its value is not simply the model behind the interface. It is proximity to the mail, meetings, files and organisational signals that make the question answerable.
That does not automatically make it the strongest tool for every reasoning task. It makes it fit for a class of work that an isolated frontier model may not be able to perform at all.
Context can beat cognition when cognition is missing the facts.
The good-enough threshold
Premium vendors do not have to be beaten everywhere to lose the default position.
They only have to be matched closely enough on a large number of ordinary tasks while remaining meaningfully more expensive, more constrained, or more awkward to use. Once a cheaper tool crosses the good-enough threshold for a task class, the burden of proof moves.
Why are we paying more here?
There may be a good answer. The model may be more reliable, easier to govern, better at tool use, faster under load, or much less likely to turn a careful analysis into confident porridge. But 'it is the premium model' is not an answer. It is branding leaning on the architecture.
Over time, the threshold tends to move outwards. Not smoothly, and not for every task, but routine work that once required a frontier model is increasingly handled by smaller models. The difficult tail moves too, though not at the same pace. This creates a continuing separation between the capability you want available and the capability you should pay for every time.
That is not a single-vendor strategy.
It is a routing problem.
Tokenomics is more than token price
The paper that started this train of thought was right to focus on tokenomics. But raw price per million tokens is only the beginning of the calculation.
The useful measure is cost per completed, acceptable task.
That includes the input and output tokens, certainly. It may also include retrieval, tool calls, long context, repeated attempts, verification and the human time required to repair an answer that was cheap in exactly the wrong way. The major API providers already price models, caching, tools and service levels differently. Their catalogues are not price lists so much as early architecture diagrams.
A low token price can become expensive if the task needs four attempts and a human rescue. A premium model can be wasteful if a smaller model produces the same acceptable result first time. The cheapest prompt and the cheapest completed task are not necessarily related.
This is where enterprise cost models can become oddly theatrical. We calculate unit prices to four decimal places, then leave 'somebody checks the output' as an uncosted footnote.
Somebody always checks the output. If they do not, somebody eventually checks the consequence.
Tokenomics therefore has to include the whole path from request to accepted result. Once it does, financial architecture and technical architecture become the same conversation.
The router is the architecture
If depth, volume and context-bound work have different needs, then the central design question is not which model wins.
It is who decides where the work goes.
Sometimes that decision will be made by a person choosing a tool. Sometimes it will be policy: sensitive work stays inside an approved environment, routine classification goes to a cheaper model, and consequential analysis is escalated. Increasingly, it will be a routing layer that considers task type, data sensitivity, required context, latency, cost and the standard of output needed.
That layer does not need to become an orchestration cathedral. Enterprise architecture has produced enough ornate structures that nobody can operate without adding one dedicated simply to AI.
Start with a few honest task classes.
- Which work genuinely benefits from the strongest reasoning available?
- Which work only needs to clear a defined quality threshold?
- Which work depends on access to organisational context?
- Which work is sensitive enough that model choice is constrained before capability is considered?
- How will we know when price, performance or user behaviour has changed enough to revisit the route?
The answers will not be permanent. That is the point. A routing posture allows the organisation to change its mind without rebuilding the entire strategy around a new logo.
It also creates a more useful role for evaluation. Instead of asking one model to win an artificial decathlon, test tools against the work they may actually receive. Measure whether the result is acceptable, how often it needs intervention, what context was required, and what the completed task cost.
The model leaderboard can continue without us for an afternoon.
What I would tell the room
If I were writing the recommendation paper today, it would not propose one universal AI tool for the whole organisation. That looks tidy on a slide and becomes untidy everywhere else.
I would recommend a controlled portfolio:
- access to frontier capability for work where reasoning quality repays the premium;
- a lower-cost tier for repeatable, high-volume work with a clear acceptance standard;
- context-rich tools where organisational data is the decisive ingredient; and
- a routing and governance approach that can change as capability, pricing and risk change.
That does not mean buying everything. A portfolio can become a petting zoo surprisingly quickly. Each additional tool brings identity, data, support, procurement and control obligations. Variety is useful only when the difference has a job.
The recommendation should therefore make the trade explicit: add a tool when it serves a distinct class of work better enough to justify the operational burden. Otherwise, consolidate.
I would also ask for one thing most AI strategies still treat as an afterthought: structured feedback from the people using these tools in anger.
Not an annual satisfaction survey asking whether AI has made them feel innovative. Ask what they reach for first, where they abandon a tool, what they use outside the approved stack, which failures waste their time, and which tasks have quietly moved from one model to another.
Those answers are architecture signals.
Watch the garage light
The garage hours effect is not a procurement method. It is an early-warning system.
The people experimenting at 9 pm encounter changes in capability, usability and cost before those changes appear in a formal roadmap. Their behaviour is imperfect evidence, full of personal preference and shaped by whatever they happen to be building. Treating it as objective truth would be a mistake.
Treating it as irrelevant would be a larger one.
When experienced practitioners start drifting from a tool, ask why. Perhaps a competitor is genuinely better. Perhaps the incumbent has become too expensive for routine use. Perhaps guardrails are interrupting legitimate work. Perhaps the preferred tool simply fits the way people think, which sounds soft until you remember that unused enterprise software has perfect compliance and no value.
Capability will keep advancing. Prices will keep moving. Context will remain awkwardly local, because organisations have spent decades putting important information in places that made sense to somebody at the time.
The durable strategy is not to predict one winner. It is to understand the work well enough to route it, measure it and change course when the evidence moves.
And some of the earliest evidence will arrive long before the next vendor briefing.
It will be sitting at a terminal at 9 pm, trying to finish what was supposed to be a quick test.
← Notebook