Thomson Reuters has built its own artificial intelligence model for legal, tax and compliance professionals, spending roughly $40 million on training it with proprietary content from decades of company archives, according to a report published by The New Stack. The model, called Thomson, wasn't created from scratch but instead started with an existing open-source foundation that the company then trained on its specialized material. Early testing shows Thomson competing with models from OpenAI, Anthropic and Google on several professional and general-purpose assessments, offering a middle path for companies with deep data libraries who want control without the multi-billion-dollar price tag of building from zero.
The model draws on content from Westlaw, Practical Law, Checkpoint and Reuters, with hundreds of subject-matter experts involved in reviewing outputs and identifying failures. Thomson Reuters tells The New Stack it has used less than 10% of the content available to it for training so far. Most of the $40 million investment went toward further training on proprietary material and expert evaluation rather than pre-training a foundation model from the ground up. In published comparisons with Gemini 3.1 Pro, Claude Opus 4.8 and GPT-5.5, Thomson led in three of seven reported categories, including PRBench Legal Hard, an instruction-following composite and long context, while the competing models scored higher in other areas such as Stanford LegalBench and coding. Thomson currently powers Tabular Analysis in CoCounsel Legal, a feature that can process up to 10,000 documents and answer up to 100 questions about them, and a smaller open-weight version is available to researchers on Hugging Face.
The company says Thomson is specifically trained to flag uncertainty rather than produce a confident-sounding answer simply because a user asked for one. "That's a deliberate choice: a model that always tries to please the user is a model more prone to hallucination, sycophancy, and other failure modes that matter far more in professional work than in casual use," Thomson Reuters tells The New Stack. Even with that training, the system still relies on retrieval, pulling from sources like Westlaw and Practical Law so responses can be grounded in material a legal professional can verify. The company acknowledges it can't guarantee every output can be traced line by line to an exact statute or ruling, but says Thomson is designed to make its reasoning checkable where possible while leaving final verification to the professional using it.
Thomson Reuters argues that owning the training process gives it more control over how the model handles tradeoffs that general-purpose models don't optimize for in the same way, including how Thomson weighs helpfulness against accuracy. Because legal work involves high-stakes consequences for errors—a fabricated citation or wrong answer can escalate quickly once it enters actual case preparation—the company chose to train its model for the domain and then give it access to authoritative information when it needs to answer a question, rather than relying entirely on outside vendors. Thomson Reuters continues to use frontier models elsewhere in its products, with the new CoCounsel Legal built on Anthropic's Claude Agent SDK, but Thomson gives the company its own model for jobs where specialized training makes sense. The question the company now faces is whether the specialized slice of the business where Thomson excels is large enough to justify the investment over simply licensing third-party models for everything. For firms weighing whether to build or buy in the AI stack, the calculus turns less on technical capability and more on whether proprietary data and workflows create enough differentiation to defend the cost of ownership over time.

