Thomson Reuters Launches $40M Proprietary AI Model Built on Alibaba's Open-Source Foundation

3 Sources

Share

Thomson Reuters unveiled its first proprietary large language model after investing $40 million in training costs. Built on Alibaba's Qwen open-source foundation, the model powers the company's CoCounsel Legal AI assistant for lawyers. The move aims to reduce reliance on external AI providers while maintaining intellectual property control over legal and tax workflows.

Thomson Reuters Invests $40 Million in Proprietary AI Model

Thomson Reuters launched its first proprietary large language model called Thomson on Monday, marking a strategic shift toward owning its AI infrastructure rather than solely relying on external providers

1

. The company invested $40 million over two years covering talent and compute costs, though economies of scale reduced the final training run to approximately $450,000

2

. This investment positions Thomson Reuters to control its generative AI capabilities while building equity in its own intellectual property rather than continuously paying licensing fees to frontier AI labs.

Source: SiliconANGLE

Source: SiliconANGLE

Built on Alibaba Qwen Open-Source Foundation

Joel Hron, chief technology officer at Thomson Reuters, revealed to Business Insider that the large language model starts from a base called Snowdon, which emerged from "realigning" Alibaba Qwen—specifically identified as Qwen3.5

1

. A joint team from Thomson Reuters and Imperial College in the UK adapted the open-source model over several months to ensure it was "ethically and politically de-biased and safe to use"

1

. The company's official announcement described only "a strong open-source foundation" without naming Alibaba Qwen directly. Hron emphasized that "there's nothing that necessarily ties us to Qwen," acknowledging Alibaba's recent announcement about charging its biggest users

1

.

Deployment in CoCounsel Legal AI Assistant for Lawyers

The proprietary AI model for legal work first deploys inside Tabular Analysis within CoCounsel Legal, Thomson Reuters' AI assistant for lawyers focused on high-volume document review

1

2

. CoCounsel Legal remains multi-model by design, using Thomson for tasks where domain-specific expertise provides advantages and third-party frontier models for other functions

2

. Thomson will serve as the default model for Tabular Analysis, though administrators retain the ability to select alternative models. The capability reaches law firms and corporate legal departments in the next release, with plans to extend Thomson across the company's legal and tax portfolio

1

.

Training on Westlaw and Professional Content

The training process incorporated reinforcement learning that taught the model to work with Thomson Reuters tools including Westlaw and Practical Law

2

. Westlaw encompasses over 40,000 individual databases and more than 150 years of legal publishing and editorial curation. Hundreds of subject-matter experts helped define training objectives, create examples of legal questions, and judge responses in blind comparisons. Thomson Reuters has trained the model on less than 10% of its total content so far, drawing from Westlaw, Practical Law, Checkpoint, and Reuters

1

. Jonathan Schwartz, head of foundational research, noted the team focused on continual learning to add domain skills without erasing existing capabilities

2

.

Continued Partnership with Anthropic and Multi-Model Strategy

Despite launching its proprietary AI model, CoCounsel still relies mostly on Claude from Anthropic

1

. Thomson Reuters expanded its partnership with Anthropic in May for the same product. Hron stated that "our main objective is to make Thomson the model that powers more and more of CoCounsel's capabilities over time," while clarifying the new model does not replace collaboration with Anthropic and other labs

1

. This multi-model approach allows the company to leverage specialized models for specific tasks while maintaining flexibility. Notably, Anthropic has accused Chinese labs including Alibaba of illicitly distilling Claude's outputs to train their own models, calling it the largest distillation campaign yet run against Claude in June

1

.

Expanded iManage Partnership and Model Context Protocol

Thomson Reuters and document-management company iManage announced an expanded partnership on August 20, pushing CoCounsel Legal deeper into the iManage platform alongside HighQ, Noetica, and Legal Tracker

1

. The companies will add support for Model Context Protocol, enabling approved Thomson Reuters tools to reason over iManage content while preserving access controls, ethical walls, and privilege boundaries. Rawia Ashraf, co-head of CoCounsel Legal, emphasized that legal work "lives in too many places," highlighting the need for integrated workflows across platforms

1

.

Intellectual Property Control and Sovereign AI Strategy

Hron cited cost and intellectual property control as primary drivers behind building a proprietary large language model. He compared the decision to buying versus renting a house: "Renting a house, you still have a roof over your head, and somebody's taking care of it, and it's great. But you're not building any equity that compounds into something valuable for you long term"

1

. Ownership gives Thomson Reuters authority over deployment, governance, and future development while reducing heavy inference costs typical of frontier models

1

. The sovereign AI approach ensures customer data is not used to train the model, addressing data privacy concerns critical to legal and tax professionals

2

.

Source: The Next Web

Source: The Next Web

Performance Testing and Academic Validation

Internal tests showed Thomson performs competitively with leading models when all have access only to the web, moving to roughly equal or slightly better performance when connected to Thomson Reuters content

2

. Andrew Bean, senior research scientist, noted tests assessed both answer completeness and whether citations supported claims. Thomson Reuters opened the model to legal and AI academics before launch, with Jonathan H. Choi of Washington University School of Law testing Thomson against ChatGPT and Claude using Corporate Tax class questions

1

. The company plans to release a smaller open-weight version on Hugging Face under a noncommercial academic license and is developing a portal for outside developers to request API keys for direct testing

2

.

Future Expansion Across Legal and Tax Portfolio

With only 10% of Thomson Reuters' total information base used so far, the company sees substantial room for growth

2

. Bean indicated the next step involves converting the most useful content and product activity into better training signals rather than simply adding more material. Thomson Reuters is discussing direct model access with large law firms and corporations, remaining open to customers adapting Thomson to their own knowledge and workflows. Hron acknowledged questions about whether the company can maintain pace with faster-moving AI laboratories, arguing that improvements in open-source models will provide stronger foundations for later versions while Thomson Reuters concentrates investment on professional work. "I don't see owning an AI model that embodies the knowledge and expertise that TR possesses as something that's non-core to what we do," Hron stated. "AI is a new mechanism for expertise delivery"

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved