Squeezed between two pricing models
At the end of June, the Dutch Financial Times (Het Financieele Dagblad) ran a front-page story headlined “Law firms blindsided by sharply rising AI prices”. Since then, I have had law firms on the phone asking how they are supposed to pay for all this. I understand the unease. But “blindsided” is not the word I would choose. Anyone following the underlying dynamics saw this coming. And anyone who only sees the cost increase is missing half the story.
Why the bill is going up
Virtually all legal AI vendors (Legora, Harvey and their competitors) build on the models of OpenAI and Anthropic , and pay per token. Their customers pay a flat price per user per month, based on the assumption that the price per token would keep falling. That price did fall. But consumption rose much faster. Back in May I wrote about why I expect legal AI will not stay priced per seat; Legora is simply the first vendor to acknowledge it openly.
What drives that consumption is not the large documents lawyers feed into these systems (input is relatively cheap), but two developments worth explaining: reasoning and agents.
Reasoning means that modern models think before they answer. They fill a scratchpad with intermediate steps, weigh alternatives, and check their own work. You rarely get to see that scratchpad as a user, but it consists of tokens, and tokens get billed. Take Fable 5, Anthropic’s newest model: through the API it costs $10 per million input tokens and $50 per million output tokens, and all that thinking counts as output. Better answers quite literally come with a token bill.
An agent is AI that does not give you a single answer but carries out a task on its own. It makes a plan, works through it step by step, checks the result, and retries where needed. For hours if it has to, at three in the morning if need be, with nobody logged in. The costs have decoupled from the user behind the desk, while the pricing model still charges per user.
Expensive is a relative term
What the coverage underplays: these newest models are not just a pricier version of what we already had. They are a rare leap in capability. I recently let one work through the night, autonomously, on a task that experts had previously scoped as a project of months and tens of thousands of euros. By the next morning it was done, for roughly eight hundred euros in tokens. And when we had such a model answer legal research questions in a test setup, it produced the best legal memo I have ever seen come out of a system. Cost: around three hundred euros per question.
Together, those two numbers tell the real story. Three hundred euros per question does not fit inside any subscription of three hundred euros per user per month; that is why the vendors’ pricing model is cracking. But eight hundred euros for work budgeted at tens of thousands is not a cost item, it is a bargain. Expensive and cheap are never absolute terms. They are always relative to the value on the other side of the equation.
The squeeze
For law firms this arrives at an awkward moment, because the pressure comes from two directions at once. Clients are asking for different (read: lower) prices, because billing by the hour while AI makes you ever more efficient no longer holds. The vendors are asking for different (read: usage-based) prices, because their own costs are climbing. That double pressure has to land somewhere, and it lands in the firm’s own pricing model.
Which is exactly where the opportunity is. If you sell by the hour and buy by the seat, all you see is costs rising on both sides. If you think per task, you see something else: what does this result cost me, and what is it worth to my client? A three-hundred-euro memo is absurd inside hourly billing, but can work out very well inside a price tied to the value of the answer.
That does not have to mean an abrupt farewell to the billable hour. Pim Betist recently described a middle road: keep billing by the hour if you must, but differentiate what an hour is worth (judgment, experience and responsibility deserve a premium, production work does not) and pass token usage through transparently as a separate line, rather than hiding it in the rate. Whether that is the answer, I do not know. But it is at least a starting point for thinking about pricing with more nuance than pressing everything into a single hourly rate.
What comes next
This does require direction. The most expensive models are not something you roll out firm-wide (start without a plan and you will burn through a monthly budget in a matter of minutes). It takes people who can judge which tasks are worth the investment, and it takes visibility into your own consumption, something almost nobody had as long as the flat fee covered everything.
My expectation is otherwise unchanged: virtually every vendor will introduce some form of usage-based pricing this year, phased in the way Legora is doing now. The first contract periods are expiring. And there is something else at play. The subscriptions of OpenAI and Anthropic themselves (ChatGPT’s twenty dollars, Claude’s two hundred) currently buy you a multiple of their price in compute: the labs deliberately lose money on them to bring in as many users as possible in the run-up to their IPOs. Once those IPOs are behind them and profitability starts to outweigh growth, that subsidy will likely end. And with it goes the skewed comparison every legal AI vendor now has to live with: why is ChatGPT twenty dollars a month when your tool costs two hundred?
So the question for law firms is not how to keep the AI bill down. The question is whether you dare to treat both sides of the squeeze as a single problem: what does (artificial) intelligence cost us, and what is our output worth to our clients? Firms that have that conversation now will not be blindsided by anything.


