亚马逊得以将此称为合理使用,正是当今时代的写照。
It is a sign of the times that Amazon gets to call this fair use

原始链接: http://observationalepidemiology.blogspot.com/2026/08/it-is-sign-of-times-that-amazon-gets-to.html

近期,404 Media 的一项调查揭露了亚马逊专门收购稀有书籍以供其人工智能模型训练的行径。调查人员通过追踪一本稀有书籍发现,它被送往了亚马逊的一个仓库,员工在那里通过切断书脊的方式系统性地销毁书籍,以加快扫描速度。这种做法将速度和效率置于实体书所承载的历史、知识及情感价值之上。 该报道将亚马逊的破坏性手段与互联网档案馆进行了对比。后者在进行数字化时会逐页扫描书籍,以保留其物理完整性。尽管谷歌在 21 世纪初率先采用了快速破坏性扫描技术,但互联网档案馆对“让所有人都能获取知识”的承诺证明,大规模数字化并不意味着必须永久消灭珍稀文献。 文章作者最终指出,亚马逊的行为——销毁不可替代的书籍以及潜在的知识产权窃取——属于刻意的公司决策。尽管亚马逊资源雄厚,却选择了以廉价、快速且具破坏性的模式来训练人工智能,这进一步凸显了该公司为了盲目的技术扩张而无视文化保护的态度。

近期的一场 Hacker News 讨论凸显了围绕各大人工智能公司(特别是亚马逊)利用实体书训练 AI 模型所引发的日益激烈的争议。据报道,这些公司正在获取稀有书籍,将其数字化以提取数据,随后销毁这些实体副本。 用户对这种做法的道德性看法分歧严重。批评者认为,大规模销毁文献是“越界之举”,并指出企业一方面声称“合理使用”以支持 AI 发展,另一方面却利用法律手段针对互联网档案馆(Internet Archive)等非营利项目,这种做法十分虚伪。相反,一些评论者为这种做法辩护,称这是一种务实的保存方法,认为这些书本注定要被丢弃,与其让它们在倒闭的书店中腐烂,数字化是一种更有成效的结果。 归根结底,这一讨论反映了人们对企业权力、知识产权未来,以及追求 AI 发展是否足以证明抹除实体媒介的正当性等问题的更广泛焦虑。
相关文章

原文
This 404 report justly been getting considerable coverage.

Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process. 

A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination. 

That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.

...

We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak. As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable. 

“There are different types of value,” the bookseller said. “There's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together.”

As mentioned elsewhere, this is also IP theft on a massive scale. 
There is, however, one aspect which hasn't gotten to play it deserves, namely how a genuinely ethical and public spirited company handles the same problem.
Scanning all the Books: The Work of Scribes for the Internet Archive
Anne-Laure Freant

In 1996, computer engineer Brewster Kahle founded the Internet Archive with the mission to provide "universal access to all knowledge." Today, that vision drives the methodical work happening in scanning centers where operators carefully digitize books one page at a time, preserving both the content and the physical integrity of centuries-old volumes.

The Internet Archive's approach stems from a fundamental disagreement with the digitization methods that emerged in the early 2000s. When Google launched its Books project in 2004, it revolutionized the scale of digital libraries but introduced a troubling trade-off: speed versus preservation. Google's industrial approach often involved destructive scanning—cutting book spines and dismantling bindings to facilitate rapid automated processing.

The Internet Archive chose a different path. "At the Internet Archive, we never destroy a book by cutting off its binding. Instead, we digitize it the hard way, one page at a time". That led to the adaptation of machines and software to fit the very specific purpose of the Internet Archive, and to a job: book scanner, or scribe operator.

 

Just to be clear, the destructive method is faster and cheaper but given the tremendous resources of Amazon and the spectacular amount of money that has been spent and in many cases demonstrably wasted pumping up the AI bubble, that's not much of an excuse. Arguably even worse when you remember this is all going to train the latest of Amazon's crappy Nova series. 

This was a choice they made, just like setting up massive fossil fuel power plants now rather than taking the time to increase nuclear and renewable capacity was a choice, just like stealing the intellectual property of countless writers and artists was a choice, just like rolling out products that weren't ready for prime time was a choice, just like setting up ridiculously at complex and deceptive financing schemes rather than growing the industry in a sustainable way was a choice.

 

联系我们 contact @ memedata.com