HN 上有多少内容是 AI 生成的?
How much of HN is AI?

原始链接: https://blog.coredump.cx/p/how-much-of-hn-is-ai

作者审视了自己与 Hacker News (HN) 之间爱恨交织的关系,指出它既是重要的流量来源,也是有毒言论的集散地。然而,作者最主要的担忧是该网站充斥着越来越多的 AI 相关内容。 通过对 2026 年初该网站头条新闻的系统性抽样,作者记录了 AI 内容主导地位的急剧上升。二月份,大多数日子的头五篇推文中都有 AI 主题;到了六月,每日列表中约有 50% 到 60% 的内容要么聚焦于 AI,要么可能是由 AI 生成的。 为了识别 AI 撰写的内容,作者使用了“Pangram”——一种旨在识别大语言模型(LLM)常见准确定性风格模式的检测工具。在通过人工审核验证结果后,作者得出结论:该平台正在经历深刻的转变,“AI 自我审视”和机器生成的文本已成为其每日信息流的标准特征。

近期的一场 Hacker News 讨论反映了社区对于平台上充斥 AI 生成内容和虚假民意操纵(astroturfing)的深切焦虑。许多用户表达了对互联网“手工时代”的怀念,感叹科技讨论正变得愈发愤世嫉俗、重复乏味,且充斥着 AI 炒作。 这场辩论的核心在于:AI 的盛行究竟是因为它是几十年来最重大的技术变革而产生的自然结果,还是对人类间高质量交流构成了一种生存威胁。关于 AI 检测工具的有效性,观点依然分歧严重,一些用户主张采取更严格的审核机制,或建立私密的邀请制社区,以维护高信任度的环境。 归根结底,这一讨论凸显了拥抱 AI 的必然进步派与担忧社区失去自我身份的保守派之间的张力。虽然有人认为用户只需顺应变化,或协助版主标记可疑行为,但另一些人则认为,这种由机器人驱动的“永恒九月”(Eternal September)标志着一个时代的终结,让许多曾因其独特、极客且以人为本的视角而珍视该平台的用户感到疏离。
相关文章

原文

I have a complicated relationship with Hacker News. The site is the most important aggregator of geek news and a major source of traffic to this blog. At the same time, it has a fair number of toxic commenters, making it a dependable source of insults hurled in my general direction; if you want a taste, this article has been called “watered-down” and “slop”.

The site is run by geeks and for geeks, so it’s not immune to tech trends; for example, around 2018, it had a fair number of stories focused on cryptocurrencies and NFTs. That said, the recent shift feels more profound: almost every day, it feels that the lineup is dominated by stories focused on AI, written by AI, or commented on by AI.

That images shows a particularly bad day, so to give a more honest assessment, I also performed a more systematic survey in February 2026, and again in June of the same year.

To get a sense of how much of the feed is occupied by AI-related topics, I took a sampling of the daily top #5 for all of February:

AI took four out of five spots on Feb 4 and Feb 12, plus arguably the entire line-up on Feb 5 (story #3 was submarine marketing for an AI vendor). The only days without LLM news in the top 5 were February 1 (with the first AI story at #7, then #9), February 9 (first at #8), and February 25 (with AI at #6, #9, #10).

For the second part of the experiment — figuring out which stories were likely AI-written — I tapped into Pangram. Pangram is a remarkably good, conservative model for detecting LLM-generated text. These detectors have bad rap among techies, but the objections are often based on outdated assumptions or outright misconceptions. For the tools to work, AI writing doesn’t need to be in any way “inhuman”. It’s enough that the default voice of the current crop of LLMs is quasi-deterministic: ask for the same essay twice and you’ll get a stylistically similar result. The individual mannerisms are human-like, but it’s very unlikely that your writing combines the exact same set. I write about it a bit more here.

To validate the results, I also reviewed all the flagged stories and I think the findings make sense; if anything, Pangram had a couple of false negatives. To give you a sense of what was flagged, have a look at the #3 story on February 19 (“AI is not a coworker, it’s an exoskeleton”). It had 500+ upvotes and 500+ comments. In my opinion, it has a wide range of red flags.

In June, to capture more detail, I used solid black for pure-play AI navel-gazing (vendor announcements, op-eds about the benefits or drawbacks of the technology, etc) and hatched shapes for stories that lean heavily into AI, but have broader ramifications (e.g., the Instagram AI support agent account hack). As before, stories that are only tangentially related to AI (e.g., reports of RAM price hikes) are not flagged.

In the first half of the month, roughly 60% of the daily HN lineup was AI-related or AI-generated, tapering off to ~50% as we approached the end of the month. This is up from 40% in February.

联系我们 contact @ memedata.com