.png)
The original article was published on Law.com, here.
I have a stake in this, so let me say it before I say anything else. I sell an AI-assisted legal research product that charges by the job instead of by the seat. When I argue that consumption pricing is coming to legal AI, I am arguing for my own book. Read accordingly. The arithmetic is still the arithmetic.
Here is what most firms have not priced. You buy legal AI the way you buy every other tool in the building. Harvey, CoCounsel, whatever your firm standardized on, it lands in the technology budget as a per-seat license with unlimited use. That model works for traditional software because the marginal cost of one more query is close to zero. A lawyer who creates a thousand documents in the document management systems or runs a thousand searches in the discovery database costs almost the same as the lawyer who runs ten.
Legal AI is a different business. Every run costs the vendor real money in tokens, paid to a model provider, every single time. (Or maybe, maybe, money paid to a GPU rental, but still—not a fixed cost.)
As this years’ renewals come up, the question everyone needs to be asking is: how much? How much is your vendor paying in AI tokens because, even if they are subsidizing that cost to your firm now, they won’t be going forward. The honeymoon is ending.
Model cost is quoted per million tokens, which means nothing to a lawyer, so here is a translation. As a very rough rule of thumb, think of a million tokens as about a hundred documents. You could quibble with these exact round numbers, but the order of magnitude is correct, which is what I’m focused on. In litigation the cost sits mostly on the input side, meaning the model reading your documents, rather than the output side. Flagship models currently list around $5 per million input tokens. Output runs about five times that.
Now the routine case. You are getting ready for a deposition, or answering a research question across a body of cases, or building a statement of facts out of discovery pleadings. You load a thousand documents into a workspace and ask the system to look at each one. A thousand documents is roughly ten million tokens. Ten million tokens at $5 per million is $50. That is the cost of a single query.
Say a lawyer runs ten of those a week, fifty weeks a year. That comes to about $25,000 a year in raw model cost for one lawyer. Not licensing, not support, not margin. Tokens.
Now push it. Some platforms will hold ten thousand documents in a single workspace. A lawyer who ran ten queries a day against a workspace that size would generate model costs approaching seven figures over a year. Nobody works that way and I am not suggesting anyone does. But take a tenth of it and you are still at roughly $175,000 for one lawyer.
My assumptions are simplified and they cut both directions. Retrieval and search strategies mean the system often reads a fraction of what you loaded. Caching and model routing bring the blended cost down further. Plenty of seats go barely used, which is the oldest subsidy in enterprise software.
On the other side, I left out output tokens entirely, and those are roughly five times the price of input and are what a reasoning model burns through. Several of the cost-reduction techniques also trade accuracy for price, and a wrong citation costs a lawyer more than it costs a software company. I am not after precision here. I am after the order of magnitude, and the order of magnitude is tens of thousands of dollars per lawyer per year, running to six figures for a heavy user.
By the way, if I’m wrong on this, it’s something your vendors should be happy to tell you. If someone has figured out how to reliably ingest thousands of documents for cents rather tens of dollars, that says a lot about the quality of their product—and the sustainability of their pricing. The reverse is also true, if you are getting charged $250 a seat for a product that costs the vendor $2,000 a seat, you should be worried—even if the results of that product are good!
Put another way, subsidies can’t last, and the sooner firms realize it, the better they will be able to make rational budgeting decisions and product choices. To date, vendors across AI, not just in legal, have been pricing to build user bases and habits rather than to cover cost. That is rational in a land-grab market, and I would probably do it too if I had raised the money for it. What it is not is a permanent condition. It can’t last; it’s already starting to change; and firms have to be clear-eyed about that reality.
Three ways this ends
There are only so many exits, and every firm should know which ones its current contract permits.
• The vendor raises the subscription price at renewal. The cleanest option and the easiest to plan for.
• The vendor throttles usage. Fair-use language, rate limits, and excessive-use clauses that do not bind anybody today are sitting in agreements right now and will start binding the day someone decides to enforce them.
• The vendor meters it, and charges you for what you use, i.e., “consumption” pricing.
I suspect option 3 is where we end up, but it doesn’t really matter. None of these are comfortable choices if you haven’t planned for them. If a firm has built a workflow that only makes economic sense at zero marginal cost, that’s going to be a problem when either the cost of that workflow goes up (options 1 and 3 above), or the quality or reliability goes down (option 2). Firms are standardizing on document-heavy AI workflows at exactly the moment those workflows are cheapest to run. If the terms change underneath a workflow that partners have already been trained on and clients have already been promised, the firm does not get to unwind it quietly.
It also matters for clients. Firms are used to internalizing fixed subscription costs. It is forseeable (fixed), and it is easier to reflect it in increased rates rather than trying the difficult (and sometimes ethically complex) business of passing it through. But it doesn’t work for a consumption pricing model. Usage-based fees can, and probably should be passed on to clients; at a minimum, that discussion needs to start happening now.
A cost you can attach to a matter is a cost you can explain, allocate, and in some cases pass through. That is the whole argument for consumption pricing, and it has nothing to do with whether metered pricing is philosophically superior. It is that a flat fee for a defined task sits next to what the same task used to cost in associate hours, and a client can see the difference. A per-seat license sits next to nothing.
I will speak from my own book, since I flagged it at the top. My company is bootstrapped. I pay token costs on every job that runs, which means I have never had the option of pretending those costs are not there. So I price per job. That is not a strategic insight, it is a constraint. But it does mean I know what the work costs. A firm buying by the seat has no equivalent visibility.
Here’s what I would do before your next renewal with any legal vendor:
The answers to these questions will allow you to plan realistically. If your firm’s users spent $500,000 in tokens and are being offered an annual renewal at $100,000, you want to know that. Not because you shouldn’t take the renewal; you probably should. But because you shouldn’t build your practice around the premise that those costs will stay the same.
Similarly, if you built out skills or workflows on the premise that you will be able to use Claude Opus 5 or ChatGPT Sol, you’ll want to know if you are going to be moved onto a proprietary, or open-weight, or simply lower-cost model. It will tell you that at a minimum you need to test if your workflows work on those models, and perhaps that you need to consider going directly to the model provider.
All of this, of course, will help you plan better not only for your firms internal budgeting and AI strategy, but importantly, for your discussions with your clients about the cost of their matters.
In short, in the world of AI-based software there is not only no “free lunch” there isn’t even a fixed price buffet. Token costs are real, and you need to get smart about them now. If you aren’t paying for them already, you will be.
Sam Davidoff is a trial lawyer and the founder of Align, which builds litigation software including Align Research.