Thomson Reuters launched Thomson, the company's first proprietary large language model, developed in-house. Frontier labs have typically spent billions of dollars on compute and years of infrastructure investment to reach the frontier. Thomson Reuters took a different path: starting from a strong open-source foundation and investing $40 million to train Thomson into the right intelligence for the jobs that matter most, covering talent and compute. The result is a model Thomson Reuters fully controls, without the heavy inference costs of typical frontier models.
Thomson's first deployment is inside Tabular Analysis in CoCounsel Legal, exactly the kind of high-volume, structured document review where a purpose-built model's advantage shows up immediately.
Thomson Reuters built Thomson on decades of proprietary content, technology, and domain expertise. The model is built to Fiduciary-Grade™ standards, at a fraction of the typical cost. Thomson starts from a strong open-source foundation. What makes it different is what happens next: state-of-the-art mid-training and post-training techniques, drawing on decades of authoritative content from Westlaw, Practical Law, Checkpoint, and Reuters, with hundreds of subject matter experts integrated from the design of training objectives through to the final evaluations.
The model has been trained on less than 10% of Thomson Reuters content so far, and what comes next is not simply feeding it more data. It is continued discovery of new kinds of specialization and understanding, made possible only by building on decades of proprietary content and editorial expertise.
Thomson marks a shift for Thomson Reuters into a world where questions about AI sovereignty: how a model is trained, what behaviors and biases live inside it, where it runs, and how the privacy of their information is protected are answered directly, not left to third parties.
Thomson shows a meaningful uplift from its base model in instruction following, the ability to execute complex, multi-part professional instructions precisely. The domain-specific gain challenges a common assumption, that the most capable general-purpose models only need access to the right content to perform at an expert level. Thomson Reuters’ early results suggest otherwise. Proprietary training and human subject matter expertise, applied to a strong foundation, produces gains that content access alone does not.
A full (and fascinating) description of how Thomson Reuters built Thomson, written by Alexander Kardos-Nyheim Senior Director, Thomson Reuters Foundational Research (former Founder & CEO of Safe Sign Technologies, acquired by Thomson Reuters). Is here. The report is very detailed and includes sources for all its external claims.
In a Rundown AI newsletter article, it opined that Thomson, as a homegrown frontier model, matters because "$40M might sound like a lot, but to a giant company previously on expensive models, it's closer to a year of API invoices than a moonshot budget. Expect more orgs with deep content and knowledge bases to run the same math, especially as training costs become even more accessible on top of strong open base models."
It will be interesting to see if other companies follow Thomson Reuters lead and build their own frontier models.
The full Thomson Reuters press release is here.