近期一篇 Hacker News 的讨论聚焦于微软一位总监的言论,他将 AI 数据抓取称为“人类历史上最大的劳动窃取行为”。
社区对此反响多持怀疑态度。许多用户指出了由微软高管发表此类言论的讽刺意味,并援引了该公司自身激进的数据操作历史,以及对微软“Copilot”可能在暗中收集用户专有数据以获取竞争优势的担忧。
评论者还对该总监的夸张言辞提出了质疑,指出将 AI 抓取与历史上真实的奴隶制相提并论是不恰当的。另有人推测,这位高管的公开立场可能是一种战略举措,旨在让微软与 OpenAI 保持距离,或转移公众对其自身在 AI 竞赛中困境的关注。总体而言,该讨论反映了业界对大型科技公司的一种深刻犬儒主义态度:即这些公司一方面大肆构建依赖数据抓取的巨型 AI 模型,另一方面却突然标榜数据保护,其背后的伦理动机备受质疑。
相关文章
原文
The New York Times sued OpenAI and Microsoft for copyright infringement in late 2023, with the case apparently still ongoing almost three years later. Now, the publication’s legal team has asked the court for a summary judgment after it filed a revealing legal brief based on statements and documents from the defendants. According to 404 Media, these documents remain sealed or redacted at the request of both companies, with the revelations showing potentially damaging statements from their leadership, including claims AI scraping is the biggest theft of labor in human history and an existential threat to publishers.
Go deeper with TH Premium: AI and data centers
(Image credit: Microsoft)
The brief cited an internal memo dated January 2023 by Microsoft director of Applied Science Brent Hecht, where he allegedly said, “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and also called it “the largest theft of labor in human history.” Another Microsoft document was cited saying, “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”
As ChatGPT surged in popularity throughout 2023, the software giant’s own data revealed that Copilot dropped click-through rates for The New York Times by as much as 93% compared to Bing search. Another memo by the Applied Science director called it a “doom loop” and said it would “hurt the performance of our models and the entire web at the same time.” The NYT brief quoted Hecht from the document, saying, “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”
Latest Videos FromTom's Hardware
OpenAI Head of ChatGPT Nick Turley said in internal communications that the AI chatbot is an “existential threat” to publishers as they are “largely substitutive” and “will get more and more substitutive as they get better,” while another OpenAI engineer testified that “no matter how prominently we show the links, users won’t click.” Nick Ryder, another OpenAI researcher, told company president Greg Brockman about a “hack to get around nytimes paywall,” to which he replied, “ah nice.”
AI companies argue that scraping the internet for data to feed to their models is “fair use,” with one court agreeing that Anthropic’s use of published material falls under this category. The law defines this as “criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research.” Some of the factors that determine whether a particular use falls under “fair use” include “(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.”
However, all these revelations in NYT’s brief could complicate OpenAI’s fair use defense, especially as it shows that the leadership of both companies are aware of the possible market repercussions of AI scraping. Microsoft CEO Satya Nadella said in a deposition from earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training” and that if he “had been made aware that OpenAI has scraped and trained on information that was behind a paywall,” the company would have required OpenAI “to retrain its models.”