Thomson Reuters Just Answered Your Build vs. Buy Question

Thomson Reuters spent $40M fine-tuning an open-source model on Westlaw data to match frontier performance. Compare that build vs. buy math against your own.

Scott Armbruster
11 min read
Thomson Reuters Just Answered Your Build vs. Buy Question

On August 24, Thomson Reuters announced Thomson, its first proprietary large language model. They didn’t train it from scratch. They took an open-source base model, fine-tuned it on decades of Westlaw, Practical Law, Checkpoint, and Reuters content, and spent about $40 million getting there.

For two years, every build-versus-buy conversation I’ve had with a client ended the same way. Someone asks what it would cost to train their own model. Nobody has a number. The conversation dies and they sign the API contract.

Now there’s a number.

Quick Verdict

QuestionThe Answer
What launched?Thomson, Thomson Reuters’ first proprietary LLM, announced August 24, 2026.
Built from scratch?No. Fine-tuned from an open-source base, most recently Qwen 3.5, per The Decoder.
Total cost?Roughly $40 million in compute and talent over two-plus years.
Cost of the final training run?About $450,000, per LawSites. Those two numbers are the whole story.
Trained on how much data?Less than 10% of Thomson Reuters’ available proprietary content.
Does it beat frontier models?On its own domain benchmark with proprietary data access, yes. On general legal benchmarks, no.
Where does it ship first?Tabular Analysis inside CoCounsel Legal, a high-volume document review feature.
Did they drop OpenAI and Anthropic?No. Legal IT Insider reports they still run frontier models where those win.
Who should copy this?Companies with 15+ years of structured, expert-reviewed proprietary content in one domain.
Who should not?Almost everyone else, and I’ll be specific about why.

What is a domain-specific LLM?

A domain-specific LLM is a language model that starts from a general-purpose open-source base and then gets additionally trained on one industry’s proprietary content until it outperforms general models at that industry’s tasks. You inherit the base model’s reasoning and language ability for free. You add the expertise nobody else has.

That’s the entire strategy in two sentences. The hard part isn’t the concept.

The Two Numbers That Reframe Everything

Forty million dollars. Four hundred fifty thousand dollars.

The $40M is what Thomson Reuters spent over two-plus years on staff, compute, experimentation, and roughly six swaps of the underlying open-source base model as better ones shipped. The $450K is what the final training run cost for the version they launched.

Read that gap again, because it’s the most useful thing in this announcement. The compute was never the expensive part. The expensive part was two years of legal experts, data engineers, and evaluation infrastructure figuring out what to train on and how to measure whether it worked.

CTO Joel Hron said it directly to Legal IT Insider: “The open-source foundation is the starting point, not the moat. The difficult part is everything that happens after that: the data, the training methodology, the domain expertise.”

I’ve been making a version of this argument for a while. When I wrote that your data moat matters more than your model tier, the counterargument was always that nobody outside a hyperscaler could actually act on it. Thomson Reuters just acted on it for the price of a mid-sized office building.

The Benchmark Result Everyone Should Read Carefully

Here’s where the press release and the reporting diverge, and the reporting is more useful.

Thomson Reuters says the model performs “competitively with the latest frontier models across a range of tasks.” True, with an asterisk. According to The Decoder’s benchmark reporting, Thomson scores 0.823 on Stanford’s LegalBench, which puts it behind Gemini 3.1 Pro and GPT-5.5. It also trails Claude Opus 4.8 on the Harvey Legal Agent Benchmark.

Then look at Thomson Reuters’ internal Deep Research benchmark:

ConfigurationScore
Thomson with proprietary content access0.83
Thomson without proprietary content access0.53

A 30-point swing based purely on whether the model can reach the Westlaw corpus.

That’s not a weakness in the result. That is the result. The model isn’t smarter than GPT-5.5 in general. It’s better at Thomson Reuters’ specific job, on Thomson Reuters’ specific data, at a fraction of the inference cost. LawSites reported that Thomson outperformed GPT-5.4 and Claude Sonnet 5 when tested with proprietary content in play, and performed comparably on web-only tasks.

So the honest summary is this: fine-tuning didn’t buy them general intelligence. It bought them a cheaper, controllable model that matches frontier performance on the narrow slice of work they run millions of times a month. That’s a very different value proposition, and a much more copyable one.

The Math You Should Actually Run

Most people will read “$40 million” and stop. Wrong reflex. Run the comparison at your scale.

PathUpfront costPer-query costTime to valueYou own it?
Frontier API, prompted~$0HighestDaysNo
Frontier API + RAG on your data$20K-$150KHigh4-12 weeksPartly
Managed fine-tuning on a vendor model$50K-$300KMedium8-16 weeksNo, weights stay with vendor
Fine-tune an open-weight base yourself$250K-$40MLowest6 months to 2 yearsYes

The ranges on the bottom two rows are wide because the variable isn’t compute. It’s how much curated, labeled, expert-reviewed content you already have. Thomson Reuters had a century of it. If you have three years of unstructured Slack messages, no amount of GPU time fixes that.

The threshold question is inference volume. Fine-tuning a small specialized model pays off when you run the same task type at enormous scale, because you’re trading a large fixed cost for a much lower marginal cost. Thomson Reuters is deploying Thomson on tabular document review, which is exactly that shape: high volume, structured, repetitive, and expensive at frontier-model token prices.

If your AI spend is $4,000 a month across nine different task types, there is no version of this math that works for you. Keep buying. I’ve argued before that licensing beats building for most agent work, and this announcement doesn’t change that for the median company. It changes it for a specific profile.

How do you know if fine-tuning your own model is worth it?

Six checks. A technical lead and a finance lead can work through this in an afternoon, and five of the six are questions you can answer without writing any code.

  1. Count your monthly spend on your single highest-volume AI task. Not total AI spend. One task type, one number. If it’s under $25,000 a month, stop here and keep buying.
  2. Measure how much curated, expert-reviewed proprietary content you own. Structured, labeled, and reviewed by people who know the domain. Raw document dumps don’t count.
  3. Test whether a frontier model plus retrieval already closes the gap. Run your task with RAG against your own corpus first. If accuracy hits your bar, you just saved yourself two years.
  4. Price the talent, not the GPUs. Thomson Reuters’ final run was $450K of compute against $40M total. Budget the same ratio: roughly 1% hardware, 99% people and time.
  5. Define the benchmark before you start. Thomson Reuters built an internal Deep Research eval that measures completeness and citation accuracy on real legal queries. Without your own eval, you cannot tell whether training worked.
  6. Confirm the task is stable for 24 months. If the workflow you’re specializing on gets redesigned next year, the fine-tune dies with it.

Fail any of the first three and the answer is buy. Pass all six and you have a real conversation on your hands.

Where This Goes Wrong

Three things I’d flag before anyone takes this to a board.

Open-weight base models are a moving target, and you don’t control them. Thomson Reuters swapped its foundation roughly six times during development as better open models shipped. That’s a feature of their discipline and a warning about yours. If Alibaba changes Qwen’s licensing, or the open-weight frontier stalls, your two-year program inherits that risk. Meta already demonstrated how fast this shifts when it pulled back from open-sourcing its best models.

“We own it” is worth less than people think until you can serve it. Owning weights means owning inference infrastructure, evaluation pipelines, safety testing, versioning, and rollback. Thomson Reuters retrained its base model with Imperial College for safety and neutrality before it ever touched legal content, producing an intermediate version internally called Snowdon. That’s a full research workstream most companies don’t budget for and can’t staff.

A domain model doesn’t replace your frontier contracts. This is the part the headlines flatten. Thomson Reuters is running a multi-model strategy. Thomson handles the tasks where specialization wins. OpenAI and Anthropic still handle the rest. Anyone reading this as “build your own model, cancel your API contracts” got it backwards.

My Read

Three things I think are true.

The build-versus-buy debate just got a floor price, and the floor is lower than the headline. Forget $40 million for a second. The reproducible unit here is $450,000 for a final training run on top of a free open-weight base. The two years and the other $39.5 million bought institutional capability that Thomson Reuters now owns permanently, so version 2.0 doesn’t cost $40 million again. For a company with real domain data and a $30M+ annual AI budget, that’s suddenly a line item rather than a moonshot.

Open-weight models crossed a competitiveness threshold and most enterprises haven’t updated their assumptions. Two years ago, starting from an open base meant accepting a large capability gap you’d spend the whole project trying to close. Qwen 3.5 gave Thomson Reuters a starting point good enough that domain data became the deciding variable. That shift is why this is viable outside a hyperscaler, and it’s the same dynamic making small local models genuinely useful for production work.

The less-than-10% figure is the most strategically loaded number in the announcement. Thomson Reuters trained a frontier-competitive domain model on under a tenth of its content. The remaining 90% is a runway no competitor can buy, borrow, or scrape. Every future version gets better on inputs that are structurally unavailable to anyone else. That’s what a data moat looks like when someone finally builds on it, and it’s a more durable position than any model architecture.

Here’s what I’d tell a business owner who reads this and feels behind. You are not behind. Thomson Reuters spent a century assembling the asset that made this possible, and the model is the last step, not the first. The transferable lesson is about the asset, not the training run. Whatever your company knows that competitors don’t, it’s currently sitting in unstructured PDFs, closed tickets, and one senior person’s head. Structuring that is the work, and it’s worth doing whether or not you ever train anything.

The Bottom Line

Thomson Reuters proved that a non-AI company with deep proprietary content can fine-tune an open-source base model into something competitive with frontier models on its own domain tasks, for about $40 million over two years, using less than 10% of its available data. The final training run cost $450,000.

That’s the precedent. It doesn’t mean you should build. For most companies reading this, the correct answer is still buy, and it will stay that way. The same logic I applied to bundled versus standalone agent components holds here: the question isn’t which option is better in the abstract, it’s whether you already own the thing that makes the expensive option pay off.

What changed is that “we can’t afford to build” is no longer a complete answer. There’s a number now, and someone on your board is going to ask you to compare it to yours.

Your Next Step: This week, pull your AI spend broken out by task type instead of by vendor. Find your single highest-volume task and write down what it costs per month. Then ask one question about it: do we own content that would make a specialized model better at this than a general one? If the answer is no, you’ve just confirmed your buy decision with evidence instead of instinct, and you can stop having the argument. If the answer is yes, you’ve found the one workflow worth pricing a fine-tune against, and that’s a two-week evaluation rather than a two-year program.


Related Reading:

TAGS

build vs buy AIproprietary LLM enterprisefine-tune open-source modelThomson Reuters Thomson LLMdomain-specific AI model

SHARE THIS ARTICLE

What is this worth in your business?

The free Build Audit is 30 minutes. You leave with a ranked list of the automations worth doing in your business, whether or not we build them.