微软高管称人工智能抓取数据是“人类历史上最大规模的劳动剽窃”。
Microsoft exec called AI scraping 'the largest theft of labor in human history'

原始链接: https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/

《纽约时报》针对 OpenAI 和微软的版权诉讼中新解密的文件显示,这些公司内部曾承认其人工智能训练行为可能构成“盗窃”。文件详细说明了这些公司如何通过绕过付费墙、删除版权声明以及大规模抓取数百万篇文章来构建海量数据集。 至关重要的是,这些内部通讯与其公司辩称的“合理使用”法律抗辩相矛盾。根据该抗辩,人工智能产品不得直接替代原始材料的市场地位。微软高管将其产品描述为一种威胁新闻出版商“经济基础”的“毁灭循环”(doom loop);OpenAI 领导层内部也承认,其模型对所依赖的新闻工作者构成“生存威胁”。微软首席执行官萨提亚·纳德拉也作证称,付费内容应获得授权;他指出,如果早知 OpenAI 在抓取此类材料,他会要求重新训练模型。 这些文件揭示了公司将受版权保护的数据货币化的蓄意行为,内部备忘录甚至将其做法称为“人类历史上最大规模的劳动窃取”。这些披露显著升级了这场为期三年的法律斗争,突显了他们内部清楚地意识到,其人工智能开发是在损害而非转换提供训练数据的出版商。

这篇 Hacker News 帖子讨论了一篇援引微软高管言论的报道,该高管称人工智能的数据抓取是“人类历史上最大规模的劳动力窃取”。 参与者对人工智能训练的伦理问题看法截然不同。批评者认为,企业正在系统性地掠夺全球文化,并将其高价转卖给用户,他们警告称这种不受限制的数据收集会导致反乌托邦式的未来,即人类劳动力遭到前所未有的剥削。一些人指出,随着人工智能破坏了盈利模式,原创内容正从网络上消失,这进一步加剧了“窃取”这一说法。 相反,一些评论者质疑“窃取”是否是准确的术语,并将人工智能训练比作人类通过书籍学习。另一些人指出了监管方面的现实困境,认为严格的监管可能只会让财大气粗的大公司受益,同时阻碍开源进程。讨论还涉及系统性不平等的更广泛议题,一些人认为这种剥削不过是传统资本主义劳工实践的最新、最极端的演变。最终,舆论反映出人们对人工智能发展的社会成本,以及公众在阻止这一范式转移时缺乏话语权而感到的深层且未解决的焦虑。
相关文章

原文

New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications.

Per the lawsuit, a top Microsoft executive privately described the companies’ AI training practices as “theft,” and OpenAI’s own leadership said its AI models posed an “existential threat” to the publishers and journalists whose work trained them. 

The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data. 

It’s worth noting that much of the new information comes from The Times’ own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context.

The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content. 

The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. 

Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule’s requirement that use doesn’t substitute or harm the market for the original work.

For example, Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” reads the Microsoft document, as quoted in the filing. 

Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and made clear that, if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.” 

Other admissions cut against different pillars of the fair-use test: OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”

OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

That kind of language speaks to how the technology could directly compete with, rather than transform, the original work. 

A Microsoft document states that there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” 

The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. 

In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”

The filing lays out in new detail how OpenAI and Microsoft went about acquiring the plaintiffs’ content, including scraping it from the Bing Index. 

“OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing reads. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”

The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers. 

In order to get the most out of their scraping, OpenAI employees allegedly came up with a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.” 

OpenAI employees also allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers “wouldn’t want model outputting” “copyright notices” to users.

“The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,” Steven Lieberman, counsel for the New York Daily News, said in a statement shared with TechCrunch.

OpenAI and Microsoft did not return requests for comment.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

联系我们 contact @ memedata.com