Thomson Reuters Invests $40 Million to Train Legal Model, Large Models Compete on Computing Power, Industry Models Compete on Data Rights
Thomson Reuters has released its self-developed large language model, Thomson, based on an open-source foundational model, with an investment of approximately $40 million for two years of training and development; the company states that the cost of each training run is about $450,000, with funds primarily used for talent and computing resources.
The model is trained using years of professional content from Westlaw, Practical Law, Checkpoint, and Reuters, with hundreds of experts in law, tax, and regulation involved in setting training goals and evaluation standards. Westlaw covers over 40,000 databases, accumulating more than 150 years of legal publishing and editorial content; however, Thomson Reuters claims that the proprietary content used for training is less than 10% of its total content library.
The company states that Thomson scored 0.352 on the PrBench Legal Hard benchmark, surpassing other leading models tested. This result primarily comes from the company's publicly available evaluations, and Thomson Reuters has opened the model for external testing by legal and AI scholars, releasing a small open-weight version for academic and non-commercial use verification; thus, its true production performance remains to be independently assessed.
The first production scenario for Thomson is Tabular Analysis in CoCounsel Legal, used for high-volume, structured document review. This tool can answer up to 100 questions simultaneously for as many as 10,000 documents, tracing each answer back to the original materials; it is currently not a general legal agent, nor is it a system for automatically drafting pleadings, representing clients in court, or independently bearing legal opinion liability.
Thomson Reuters explicitly states that the model does not train on clients' proprietary documents and is not sold as an independent model. CoCounsel Legal maintains a multi-model architecture: using its self-developed model for workflows where Thomson excels, while continuing to call on external leading models in other scenarios.
In terms of market mechanisms, Thomson Reuters transforms its long-accumulated editorial copyright content, expert annotations, and legal workflows into model capabilities, monetizing through CoCounsel subscriptions and the Westlaw ecosystem. Large law firms and corporate legal departments may complete due diligence, contract reviews, and evidence sorting at lower costs; businesses relying on junior lawyers for hourly document screening will be pressured, but complex legal judgments, client communications, strategy formulation, court representation, and final professional liability will still be borne by lawyers.
Source: Public Information
ABAB AI Insight
Thomson Reuters' advantage lies not in training a larger model from scratch, but in having legally accessible professional data assets. Westlaw has long established a legal research entry through case law, regulations, citation references, editorial notes, and retrieval systems; Practical Law and Checkpoint cover practical and tax workflows, respectively. It leverages open-source foundational models to inherit general language capabilities, supplemented by proprietary content, expert feedback, and tool usage training to enhance professional reasoning, avoiding a capital-intensive head-on competition with OpenAI and Anthropic in general computing power.
In terms of capital pathways, the $40 million is not about "replacing the entire legal industry with the salary of a junior lawyer," but rather a fixed investment to upgrade the content library to the AI product layer. The model is first embedded into the existing CoCounsel and Westlaw distribution systems, cross-selling to existing law firms and corporate legal clients; the marginal cost of adding each user primarily involves reasoning, retrieval, and support, rather than re-hiring document review teams. The first areas to be compressed are standardized, traceable, high-volume, and repetitive e-discovery, contract extraction, and due diligence tasks.
Historically, Westlaw and LexisNexis have reduced the time lawyers spend manually reviewing paper cases through electronic retrieval; electronic evidence platforms have also software-ified the process of screening massive emails and documents. Thomson's distinction is that it integrates retrieval, reading, comparison, classification, and tabular output into an agent-style workflow. Each upgrade of efficiency tools in the legal industry has not eliminated lawyers but has increased the volume of documents a single lawyer can cover, shifting junior roles from mechanical retrieval to review, fact judgment, and client delivery.
Essentially, this is about technological substitution. The first roles to be substituted are not the identity of lawyers, but the billable hours associated with document review, information location, clause comparison, and tabular organization that can be clearly accepted. The mechanism of change is that professional data owners simultaneously control corpus copyright, expert validation, product distribution, and user workflows; general models can replicate language capabilities but cannot easily replicate 150 years of editorial systems, institutional client trust, and legal content authorization. The next batch of industries with similar conditions may come from medical publishing, accounting and tax, engineering standards, and financial data services.
ABAB News · Cognitive Law
Large models compete on computing power, industry models compete on data rights
Once professional data is trainable, it will reassess labor pricing
AI first replaces billable hours, then rewrites organizational structures