旧时代的互联网去哪了?我们追踪了 657,607 个链接一探究竟
Where did the old web go? We followed 657,607 links to find out

原始链接: https://0.mk/blog/link-rot

在恢复了来自已停用的马其顿短网址服务 0.mk 的 657,607 条 2009-2014 年间的链接数据库后,研究人员分析了那个时代的“小型网络”状况。到 2026 年 8 月,他们抓取了这些历史链接以测试链接失效情况。 研究结果凸显了早期网络的脆弱性:大约 76.7% 的链接已无法正常访问。即便这一数字还算乐观,因为它包含了正在“加载”的停放域名、登录墙以及错误页面。虽然谷歌和维基百科等大型平台表现出较高的留存率,但去中心化的网络——包括个人博客、论坛和马其顿当地新闻——已基本消失。 该数据集相当于 2010 年代的数字博物馆,记录了在线共享的演变,从测试 URL 长度限制到早期对 Rapidshare 和 Megaupload 等服务的依赖。原团队曾因维护和垃圾邮件成本于 2014 年关闭了该服务,但最近又利用现代人工智能实现了审核和管理的自动化,从而重启了该服务。归根结底,该项目强调了一个严酷的现实:虽然短网址旨在长期有效,但它们所指向的内容却很少能持久存在。

已停用的短网址服务 0.mk(活跃于 2009 年至 2014 年)的创作者近期恢复了其历史数据库,旨在分析早期互联网内容的持久性。通过对该项目存续期间存储的 657,607 条链接进行爬取,他们发现约 78.7% 可爬取的独立网址已无法加载。 这项分析凸显了数字内容的脆弱性,并指出即使链接在技术上能够“加载”,也往往会导向停放域名或受限页面等死胡同。作者还发现了一些异常情况,例如旧版 Facebook CDN 照片完全消失,以及包含 `localhost` 等特殊的测试链接。 作者强调,该数据集代表的是特定社区的分享习惯,而非整个互联网。他们将 0.mk 作为一项实验重新上线,试图探究现代人工智能工具能否自动化处理曾经导致该服务难以为继的维护、垃圾信息过滤及审核工作。该项目深刻地提醒了人们“链接失效”这一困扰现代互联网的问题,并引发了关于我们存储和保存数字历史方式中固有设计缺陷的广泛讨论。
相关文章

原文

An old 0.mk database backup held 657,958 links created between 2009 and 2014, along with their click counts. We restored 657,607 of those records as pre-2015 links and followed every destination in August 2026. Of 655,178 safe, crawlable link records, 76.7% no longer returned a loading page.

Most 0.mk users were in Macedonia, so this is not a census of the entire web. It is a large surviving record of what one online community shared during that period, including local news, personal blogs, photo hosts, forums, and the major platforms of the time.

When 0.mk started in 2009, it was a passion project built by a team of three. We worked on it when we could, usually for a few hours a week around our regular jobs. Seventeen years later, one of us found an old database backup on a disk and decided to bring it back.

Here is what those six years of link creation look like, with the long silence after them:

2009: 3,668 links20094k2010: 20,283 links201020k2011: 103,053 links2011103k2012: 23,148 links201223k2013: 224,931 links2013225k2014: 282,524 links2014283k2026: 404 links2026404
Raw link records, not users. The 2011 spike includes one 83,398-link batch; 97.8% of 2013 records and 99.9% of 2014 records are not attached to a recovered account. The green sliver is the 2026 relaunch.

The survival test

The crawl covers all 657,607 restored link records dated through December 2014. We excluded 2,429 records whose targets were malformed, internal, credentialed, or policy-blocked, leaving 655,178 crawlable historical links:

Could not connect: 51.24%HTTP error: 25.44%Loaded: 23.32%

51.24% could not connect (DNS, timeout, TLS)25.44% http error (4xx / 5xx)23.32% loaded (2xx / 3xx response)

Even that 23.3% overstates how much survived. A login wall, a parked domain full of ads, or a "this content is no longer available" notice all count as loading. A working page does not mean the original content is still there.

Why 657,607 links but 494,781 URLs? Multiple short links sometimes point to the exact same destination. There are 162,826 such repeat records. Counting each destination once leaves 494,781 distinct URLs, of which 492,620 were crawlable. Only 21.3% of those loaded. The percentage barely moves when repeated destinations are removed: 78.7% still did not load.

At the unique-URL level, 55.0% failed at the network layer after retrying uncertain results from a second network, and 23.7% returned an HTTP error. The most common HTTP result was 404, across 76,403 distinct URLs. Another 29,663 returned 403 or 429; those pages did not load for the crawler, but may be blocking automated requests rather than missing. A 403 or 429 can mean the site blocked our crawler, so "did not load" is more honest than saying every one of those pages is gone.

The same pattern appears at the domain level. Of 133,605 crawlable hostnames, only 34,827 had even one URL load. The other 98,778 had none.

Share with no loading page in the complete crawl. URL-level results count every distinct path; host-level results count each hostname once.

The 2011 split explains the strange annual totals. One account created 83,398 distinct links to pelaphptutorials.com. At URL level, 92.5% of 2011 destinations did not load. Count that host once and the figure is 61.7%, almost identical to 2010 and 2012.

The annual totals do not show a collapse in ordinary usage during 2012. Remove that one batch and 2011 falls from 103,053 records to 19,655; 2012 had 23,148. Almost all records from 2013 and 2014 are anonymous in the recovered data, and three quarters of their hostnames have no loading URL. Raw link volume is not a user-growth curve.

Many of the recognizable survivors are giants: YouTube, Wikipedia, and Google properties. Personal blogs, forums, local news sites, and photo hosts appear throughout the unavailable set. The centralized web has generally held up better than the small web.

A walk through the graveyard

The database reads like a museum of the 2010s internet. Some residents, with the number of links pointing at them:

Facebook photo CDN (fbcdn.net), none loaded835 links

Google Code, now redirects many URLs to its archive803 links

PureVolume, the old service is gone but its domain responds796 links

Rapidshare, no URL loaded139 links

Megaupload, no URL loaded71 links

Picasa Web Albums, no URL loaded69 links

People shared Facebook photos as direct CDN links; none of the 789 distinct fbcdn.net URLs behind those 835 records loaded. Yet PureVolume now returns pages for 633 of 653 distinct URLs, and Google Code loads or redirects 628 of 754. The original services are gone, but their domains respond. An HTTP response is not the same as preserved content.

The Macedonian layer

0.mk was the first Macedonian URL shortener, so the data is also a record of a national web that partly no longer exists. The links point to A1 Television (shut down 2011), and to the newspapers Utrinski Vesnik, Dnevnik, and Vest, all of which stopped publishing in 2017. Hundreds of links to local news that can no longer be read anywhere except, sometimes, the Internet Archive. The short links outlived the newsrooms.

The gems

Seventeen years of other people's bookmarks contain some treasures:

The first link ever shortened (July 14, 2009, 3:52 AM) was not a manifesto or a launch post. It was a CSS stylesheet on someone's WordPress blog. Two clicks, ever. Empires begin humbly.

On day two, someone shortened localhost. 0.mk/localhost pointed at http://127.0.0.1/. It got two clicks, each of which sent the visitor to their own machine. The shortest URL for the loneliest destination.

The shortest link points at the longest domain. 0.mk/1 has recorded 10,415 clicks while pointing to thelongestlistofthelongeststuffatthelongestdomainnameatlonglast.com, a 2000s curiosity that is itself now gone. Four characters pointing at sixty-three, for 17 years.

The longest URL we ever shortened is 38,753 characters, a 2012 CodePen link whose query string literally repeats TRYING_THE_MAXIMUM_URL. Someone was testing us. We passed, and we still have their test.

4,478 of our links point at other URL shorteners: bit.ly, TinyURL, goo.gl. A short link to a short link, twice the fragility. Google shut goo.gl down in 2025, so every one of those is now a chain with a missing middle: our half still works, and points at a service that no longer resolves. Link rot squared.

And the immortal one: 0.mk/7, created July 15, 2009, points at google.com. 95,999 clicks and counting, and it still points there today, seventeen years and one resurrection later.

Why 0.mk came back

By 2014, 0.mk's revenue did not cover hosting or the work required to keep it running. Spam was constant. Filtering it meant more engineering, and abuse reports needed someone to review them. The original team closed the service.

Seventeen years later, AI has changed that equation. It now handles much of the development, spam detection, abuse review, support, and monitoring that the old project could not afford. That made bringing 0.mk back realistic. The recovered links now run from edge machines in more than 300 cities. More about how that works is here.

How we checked

We tested every restored link dated before January 1, 2015. The crawler followed up to five redirects and retried connection failures from a second network. We counted HTTP 2xx and 3xx responses as loading, reported HTTP errors separately, and never contacted local or unsafe addresses.

联系我们 contact @ memedata.com