汤森路透构建的自有人工智能模型现已跻身一流水平。
Thomson Reuters built its own AI model that now ranks among the best

原始链接: https://www.thomsonreuters.com/en-us/posts/innovation/thomson-reuters-built-its-own-ai-model-that-now-ranks-among-the-worlds-best/

汤森路透(Thomson Reuters)发布了名为“Thomson”的专用人工智能模型,旨在满足法律及专业领域严苛的需求。与通用型前沿模型不同,Thomson 基于开源基础构建,并利用公司数十年来积累的权威专有内容(包括 Westlaw、Practical Law 和路透社等来源)进行了微调。 该模型在数百名主题专家的参与下开发,旨在实现专业级的逻辑推理,从而确保高准确性和可靠性。基准测试显示,Thomson 的性能不仅能与行业领先的 AI 模型相媲美,在某些情况下甚至更胜一筹,同时其规模更小,运行成本也更具优势。 作为汤森路透“受托级 AI™”(Fiduciary-Grade AI™)战略的核心部分,该模型优先考虑安全性(从不使用客户数据进行训练),并通过可信、可验证的引文确保事实的完整性。Thomson 将于今年 8 月在 CoCounsel Legal 中率先推出,并最终整合到公司整个专业产品组合中。通过将深厚的领域专业知识与高性能模型开发相结合,汤森路透旨在为通用型 AI 提供一种专业级的替代方案,专为那些对准确性要求极高的行业量身打造。

```Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 Thomson Reuters 构建了自己的 AI 模型,目前排名跻身前列 (thomsonreuters.com) 17 分 | Cynddl 发布于 9 小时前 | 隐藏 | 过往 | 收藏 | 3 条评论 帮助 rabidheuristics 2 小时前 | 下一条 [-] 在旧模型上进行测试并大肆宣传。非常酷。 回复 blitzar 8 小时前 | 上一条 [-] > 目前排名跻身前列 如果他们自己这么说的话。 回复 _--__--__ 8 小时前 | 父评论 [-] 他们需要说服的是 65 岁的律师事务所合伙人,而不是你或者我。 回复 考虑申请 YC 2026 年秋季批次!申请开放至 7 月 27 日。 准则 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:```
相关文章

原文

The most capable AI models no longer come only from frontier AI labs. One now comes from Thomson Reuters. 

Today, we are sharing early benchmarking results for Thomson, a first of its kind AI model. Across a range of benchmarks assessing legal and general capabilities, Thomson performed competitively with the strongest frontier models on the market, including Claude Opus 4.8, and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro.  

Why? Because it knows the work. 

Launching later this summer, Thomson is the newest layer of the Thomson Reuters AI strategy, and a demonstration of what becomes possible when authoritative content, expert judgment, professional tools, and model development come together. 

In 2024, Thomson Reuters acquired Safe Sign Technologies, an AI research company. At the time, the market was betting that access to increasingly powerful general-purpose models would be enough. 

We made a different bet. We believed the future of professional AI would require more than general-purpose intelligence. It would require models built specifically for the domains, standards, and consequences of professional work. We believed that the distinct advantages of Thomson Reuters decades of world-class content and expertise could be best expressed in a model that we ourselves crafted.  

Thomson is the result of that bet. And it is why Thomson Reuters will continue to set the standard for Fiduciary-Grade AI™.  

Meet Thomson 

Thomson starts from a strong open-source foundation, so it performs general-purpose work just as effectively as the frontier models. It then goes further: trained using state of the art mid-training and post-training techniques on decades of authoritative content from Westlaw, Practical Law, Checkpoint, and Reuters, content professionals have staked their reputations on for generations. 

That training was shaped by hundreds of subject matter experts who evaluated outputs, identified failure modes, and validated that the model reasons the way legal professionals actually work. The same professional standard governs how Thomson is deployed. Customer data is never used to train the model. 

The result is a model that thinks and reasons like a lawyer while outperforming models multiple times larger on the work that matters.  

Thomson Matches the Best. And Beats the Rest.  

We evaluated Thomson against the leading general-purpose models on the market for general professional work and categories spanning: 

Legal  Coding  
Tax   Math  
Accounting   Multilingualism  
Journalism   Agentic tasks  
Safety  Long context  
Reasoning   Following instruction  


Thomson is competitive with the world’s leading frontier models despite being a fraction of their size and cost to train and operate. Thomson Reuters has achieved that performance  by combining exceptional AI talent with authoritative proprietary content and deep domain expertise. And with less than 10% of Thomson Reuters content used in its training so far, there remains significant opportunity to expand its capabilities. 


Table titled 'Model Benchmark Comparison' comparing Thomson Reuters (Thomson-1-Large), Google DeepMind (Gemini 3.1 Pro), Anthropic (Opus 4.8), and OpenAI (GPT-5.5) across legal and general domain benchmarks, with the best score in each row highlighted in green. Legal Domain: Stanford LegalBench — Thomson Reuters 0.823, Google DeepMind 0.843 (best), Anthropic 0.818, OpenAI 0.832. PrBench Legal Hard — Thomson Reuters 0.352 (best), Google DeepMind 0.293, Anthropic 0.315, OpenAI 0.333. Harvey Legal Agent Benchmark — Thomson Reuters 0.857, Google DeepMind 0.555, Anthropic 0.869 (best), OpenAI 0.781. General Domain: Instruction Following — Thomson Reuters 0.914 (best), Google DeepMind 0.848, Anthropic 0.861, OpenAI 0.885. Reasoning — Thomson Reuters 0.684, Google DeepMind 0.748 (best), Anthropic 0.737, OpenAI 0.589. Coding — Thomson Reuters 0.399, Google DeepMind 0.500, Anthropic 0.598 (best), OpenAI 0.414. Long Context — Thomson Reuters 0.753 (best), Google DeepMind 0.750, Anthropic 0.741, OpenAI 0.703. Notes: Thomson-1-Large used test-time scaling; Gemini 3.1 Pro and Opus 4.8 used reasoning mode; GPT-5.5 used non-reasoning mode. Caption below reads: 'Head-to-head comparison across Legal Domain and General Domain benchmarks. The best score in each row is highlighted in green

Instruction Following is a composite average of the IFEval and FollowBench benchmarks.  
Reasoning is a composite average of the GPQA Diamond, HLE, and MMLU-Pro benchmarks.   
Coding is a composite average of the SWE-Bench Pro and Terminal-Bench 2.1.  
Long Context is a composite average of the Infinity Bench as well as some internal benchmarks developed by Thomson Reuters. 


These evaluations show Thomson’s competitiveness with industry recognized benchmarks. Our internal evaluation and training cover a wide range of scenarios, including carefully designed agentic use cases optimized for real-world professional work, tens of thousands of real-world queries written by experts, end-to-end deep research training with human-calibrated judges, and data-centric mid-training on a large amount of our content. To increase safety and robustness, we conduct training and evaluations consistent with Thomson Reuters values, and stress-test the models through human and automated red-teaming. Thomson is still early in its development. To date, less than 10% of Thomson Reuters content has been used in its training, leaving significant opportunity to expand its domain knowledge and capabilities through additional training, rigorous evaluation, and expert validation. 

Thomson also showcases its strength when a native integration with Thomson Reuters content is added. When compared with leading frontier models given unrestricted access to the web, Thomson’s access to proprietary data sources such as Westlaw, Practical Law, and Reuters news ensures both superior completeness and factuality (i.e. the ability to back up claims through accurate citations to trusted sources).


This evaluation covers 53 legal research queries written by our internal Subject Matter Experts to represent real world questions. LLM’s are connected via an in-house agentic harness to Westlaw/Practical Law for TR content and Brave search engine for web search. Completeness and Factuality are scored using LLM’s as a judge. Completeness is scored based on SME-written rubrics that list every element that would be required for a good answer to the question. Factuality is based on extracting the claims made in each report and checking whether the cited sources provide evidence for each claim. These metrics were developed and calibrated against SME scoring.


The First Deployment. Not the Last. 

Thomson’s first integration will launch in August inside Tabular Analysis in CoCounsel Legal, and that choice was deliberate. Tabular Analysis performs high-volume, structured document review against a clear, measurable accuracy standard. It is exactly the kind of work where a purpose-built model has a demonstrable advantage over a general-purpose alternative, and where that advantage is immediately visible to the professionals relying on the output. Thomson will become the default model powering Tabular Analysis, and over the next year we will continue to integrate it across the Thomson Reuters product portfolio in legal and tax.  

Content. Expertise. Tools. Now the Model. 

Thomson Reuters has always brought together authoritative content, deep domain expertise, and the tools professionals rely on every day. Thomson adds the fourth element: a model purpose-built to power it all, trained on content competitors cannot access and validated to the standard professional’s demand. It delivers frontier-level performance at a fraction of the size and operating cost of many general-purpose models. A model only Thomson Reuters could build. 

This is our commitment to Fiduciary-Grade AI™ in action: AI designed for professionals with duties of care and accountability, where almost right is not good enough.  

General-purpose AI is built for everyone. Thomson is built for the professionals who cannot afford to be wrong.

联系我们 contact @ memedata.com