伊莱亚斯·索恩案:AI 聊天机器人痴迷的虚构人物
The Case of Elias Thorne, Imaginary Man AI Chatbots Are Obsessed With

原始链接: https://www.vice.com/en/article/the-strange-case-of-elias-thorne-the-imaginary-man-ai-chatbots-are-obsessed-with/

AI 聊天机器人经常创作以虚构人物“埃利亚斯·索恩”(Elias Thorne)为主角的故事,他通常被塑造成灯塔看守人或钟表匠。康奈尔大学的一项研究发现,埃利亚斯这类比喻出现在近 88% 的 AI 生成故事中,这表明 AI 在原创性上存在系统性的缺失。 研究人员认为,这种现象源于 AI 的安全与对齐训练。为了避免版权侵权和风险内容,开发人员限制了模型可访问的数据集,导致其“创作”素材库非常浅薄。此外,由于现代 AI 越来越多地使用早期 AI 系统生成的内容进行训练,这些模型缺乏数据多样性。因此,一旦像埃利亚斯·索恩这样的比喻被创造出来,它就会在各个平台(从亚马逊书籍到健康指南)上被反复循环使用。 归根结底,埃利亚斯·索恩象征着生成式 AI 的局限性。这些模型远非无所不能,它们往往依赖于陈旧且重复的反馈循环,这证明了如果没有新鲜的人类输入,AI 的输出终将空洞且缺乏原创性。

Hacker News 最新 | 往期 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 Elias Thorne 之谜:AI 聊天机器人为何对这个虚构人物如此痴迷 ( vice.com ) 13 分 作者 celadonuproot 1 小时前 | 隐藏 | 往期 | 收藏 | 1 条评论 help Mistletoe 16 分钟前 [–] 我本以为会对这个故事进行更深入的挖掘,但它写得有点散乱,然后就草草收场了,就像是人工智能写的。 回复 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

No matter the company, AI chatbots were raving about the same guy named Elias Thorne. He must be pretty fascinating. And he is, at least on paper. Depending on the AI, he’s a lighthouse keeper, a clockmaker, a librarian, an explorer, and the star of countless stories. He’s appeared in books, music listings, YouTube videos, and even health guides. You’d think he was one of the most influential men on the planet.

But he doesn’t exist.

According to reporting by fine folks at 404 Media, researchers at Cornell University may have figured out why large language models invent and keep telling tales of the same fictional man. In a study examining roughly 20,000 AI-generated stories from all the big LLM models, including OpenAI, Anthropic, and Google, the research team found that the same handful of names and occupations kept cropping up. Specifically, names and words like Elias, Mara, Elara, lighthouse keeper, clockmaker, and librarian showed up in 88 percent of stories. Elias the lighthouse keeper appeared in nearly two-thirds of them.

The obvious explanation is that AI models learned the name from some book or from somewhere in the tangled web of Internet culture. But the researchers couldn’t find evidence for that. So this theory shifted, this time to a side effect of AI safety and alignment training.

AI companies don’t want to run afoul of gigantic corporations that are notoriously litigious, like your Nintendos or your Disneys, so they run it through training that steers it away from copyrighted material. Same goes for any risky adult material. All that training creates a shallower pool of resources AI models can draw from when generating a story.

It May Be Impossible to Trace Elias Thorne’s Exact Origins

Add to that the fact that modern AI models are often trained on datasets built from earlier AI systems, essentially just rehashing the same old ideas again and again, with very little diversification in its gene pool, and it starts to make sense why once Elias Thorne was invented, he just kept getting dragged along from one iteration of an LLM model to another.

It may be impossible to trace where Elias Thorne came from, but with so much cross-pollination going on between chatbots, it’s no wonder that 404 media was able to find the name having broken containment, with the name cropping up in AI slop books and music available on Amazon, YouTube videos, and, as spotted by software engineer Daniel May, in some questionable health guides.

The pool of information AI chatbots pull from seems like it should be vast, limitless. It’s actually quite shallow, and by this point, they seemingly consume every scrap of data, every page of literature, they possibly can. Without fresh input, the output goes stale, and fast.

Elias Thorne is a weird quirk of the system, but also a symbol of just how hollow and deeply unoriginal chatbots can be.

联系我们 contact @ memedata.com