Chess.com 数据泄露影响 730 万用户,证据指向网络爬取
Chess.com Leak Exposes 7.3M Users, Evidence Points to Scraping

原始链接: https://securityaffairs.com/197174/breaking-news/chess-com-leak-exposes-7-3-million-users-evidence-points-to-scraping.html

近期,730 万条 Chess.com 用户记录在数据泄露论坛上被免费公开,其中包括姓名、电子邮箱、地理位置及平台活动信息。技术分析证实数据真实,但表明这些信息很可能是通过大规模抓取获取的,而非服务器遭到了入侵。 证据显示,这些数据在九天内被收集,且存在重复条目,这与 2023 年攻击者滥用“查找好友”功能的事件如出一辙。虽然泄露的数据中不包含密码和支付信息,但其中含有通过公共 API 无法获取的特定内部营销标签,这表明其提取手段比简单的公共抓取更为复杂。 尽管 Chess.com 的核心基础设施和用户密码仍然安全,但泄露的数据足以被用于实施复杂的网络钓鱼攻击。建议用户对看起来来自 Chess.com 的邮件保持警惕,并防范潜在的社会工程学攻击。泄露者看起来更倾向于数据采集者而非勒索者,但个人标识符的暴露对于所有受影响的成员来说,依然是一个隐私隐患。

一份最新报告显示,Chess.com 有 730 万条用户记录遭到泄露,目前关于这些数据是通过未经授权的抓取获取,还是平台直接遭到入侵,各方仍存在争议。 虽然最初的报道指向抓取,但 Hacker News 上的评论者对此表示怀疑,理由是数据中包含了详细的谷歌广告管理器(Google Ad Manager)受众群体细分信息(如实验组和试用资格),而这些信息通常无法通过简单的抓取工具获取。批评者认为,此次事件很可能是一起严重的安全漏洞,并指出该平台的“查找好友”功能此前曾被利用来收集用户数据。许多用户对 Chess.com 未能实施必要的防护措施表示不满,并指出此前已有 70 万用户受到过类似事件的影响。此次讨论凸显了人们对该平台数据保护措施的持续担忧,以及对公共数据抓取与内部数据泄露之间区别的关注。
相关文章

原文

Chess.com Leak Exposes 7.3 Million Users – Evidence Points to Scraping

7.3 million Chess.com profiles leaked online: the data is genuine, but evidence points to large-scale scraping, not a server breach.

Free is a strange price for stolen data, and that’s exactly what makes this listing worth a second look. A 15.5 GB file containing over 7.3 million chess.com user records showed up on two data-leak forums this week, no cost, no ransom demand, just handed out. Ransomnews’s technical analysis confirms the data is real and recent. What it isn’t, on the evidence, is a hack.

“The archive is a single 744 MB 7-Zip file that expands to a 15.5 GB tab-separated table: one header row and 7,337,395 records, each with 38 fields. The schema is chess.com-specific throughout. Alongside the obvious identifiers, email, partial email, username, user ID, UUID, first and last name, country, location and locale, it carries platform state: chess title, points, skill level, premium status and label, verification and activation flags, best rating and rating type, official rating, member-since and last-login timestamps.” reads the report published by Ransomnew. “Two fields at the end are the interesting ones. Every record has gam_audiences and audiences_member_of populated, Google Ad Manager audience segments, with values like coach-nudge experiment groups, trial eligibility, lapsed-user cohorts and rating-band targeting. Those are marketing-stack fields, not profile data. They do not appear in chess.com’s public API.”

The file carries email addresses, usernames, real names, countries, chess ratings, subscription tiers, and something odder: internal Google Ad Manager audience tags, the kind of marketing segmentation data that never shows up in chess.com’s public API. Roughly three-quarters of records include an email address. There are no passwords, no password hashes, and no payment data anywhere in the file, which matters a lot for how seriously affected users need to react.

Proving this data is genuine didn’t require touching chess.com’s servers at all. Every account UUID in the file is a version-1 identifier, the kind that embeds the exact timestamp it was generated, and researchers decoded that hidden timestamp across 200,000 sample records to compare it against each account’s registration date. The match rate came back at 100%, which isn’t something anyone could fake without possessing actual chess.com-issued identifiers down to the millisecond.

Three separate details point toward scraping rather than an actual system breach. The data wasn’t captured in one moment, it was stamped across nine consecutive days in daily batches, the pattern of a scheduled collection job rather than a single database dump. About 7.4% of user records appear twice, the same accounts revisited on different days, something that simply doesn’t happen inside a genuine database export.

This has happened to chess.com before, and the company was blunt about it at the time. Back in 2023, a similar leak of 828,000 records surfaced with a nearly identical field structure, and chess.com stated plainly,

“In November 2023 a threat actor published 828,000 chess.com records with a near-identical field set. Chess.com’s response then was unambiguous: as it told Hackread, “This was NOT a data breach.” continues the report. “Our infrastructure, member accounts, and data such as passwords are secure.” The data had been pulled by abusing the platform’s find-friends feature, feeding in externally sourced email addresses to resolve them against accounts. A second scrape affecting roughly 476,000 users followed. This 2026 file is the same technique at roughly nine times the scale.”

That earlier incident came from abusing the platform’s find-friends feature to resolve external email lists against real accounts; this new file looks like the same technique running at roughly nine times the scale.

One detail doesn’t fit a purely public-facing scrape, though. Advertising-audience segment data isn’t something chess.com’s open API exposes, and it appears on every single row in this file, which suggests whoever built this had access to an authenticated or internal-facing endpoint rather than just the public developer tools. That’s the specific question chess.com is best positioned to answer, and it’s the one that actually matters for understanding how this happened.

The account distributing the file, going by V0idix, isn’t monetizing anything here. The same handle has posted dozens of free database dumps across other unrelated companies, building reputation through volume rather than through sales, which fits a collector who harvests and republishes data rather than someone selling access to a fresh intrusion.

None of this means chess.com users should shrug it off just because passwords weren’t exposed. A verified email sitting next to a real name, country, skill rating, and subscription tier is more than enough raw material for a convincing phishing message about a membership renewal or a fair-play dispute. The right response isn’t panicking about a hacked account, it’s treating unexpected chess.com emails with more suspicion than usual and checking whether that same email address has turned up anywhere else, since reused credentials remain the far more dangerous exposure than anything sitting in this particular file.

Follow me on Twitter: @securityaffairs and Facebook and Mastodon

Pierluigi Paganini

(SecurityAffairs – hacking, Chess.com)



联系我们 contact @ memedata.com