如何通过高维统计赢得一场啤酒赌局
How to win a beer with high-dimensional statistics

原始链接: https://jamiesimon.io/blog/how-to-win-a-beer-with-high-dimensional-statistics/

Dhruva Karkada 的一项病毒式传播的研究近期表明,大型语言模型(LLM)对于结构化序列(如月份)的嵌入向量,在 PCA 图中呈现出完美的圆形,并产生了循环格拉姆矩阵。这暗示了数据统计与表征几何之间存在着深刻而优雅的联系。 作者对此持怀疑态度,并打赌称可以利用看似无关的词汇来“伪造”这种几何模式。通过使用迭代搜索算法从庞大的词汇库中筛选,作者成功识别出一组随机的十个词,它们同样构成了清晰的圆形 PCA 图和循环格拉姆矩阵。 虽然作者承认这些“虚假圆环”缺乏原始月份研究中的统计稳健性(尤其是在集合规模增加时),但该实验对研究人员而言是一个警示。它证明了只要有足够的搜索能力,人们纯粹凭偶然就能在高维嵌入空间中发现具有误导性的几何结构。因此,这挑战了某些自动特征发现和可解释性方法的可靠性,证明了 LLM 中“漂亮”的几何模式有时可能是统计筛选的产物,而非内在的语义属性。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 如何用高维统计学赢得啤酒 ( jamiesimon.io ) 22 积分 作者 jamie-simon 3 小时前 | 隐藏 | 过往 | 收藏 | 2 条评论 帮助 mdritch 1 小时前 | 下一条 [–] 我最近做了一些分析,研究了诸如“简洁明了”和“避免矫揉造作的文风”等不同风格提示词对大语言模型(LLM)输出的影响,我还在相似度矩阵的二维降维(MDS)中看到了一个近似圆的关系。我在想这是否是 PCA/MDS + Gemma 输出分布产生的人为现象。 回复 ViscountPenguin 1 小时前 | 上一条 | 下一条 [–] 一个有趣的进阶做法是,试着通过一组点来重复这个过程,这些点在现有圆圈上距离除一点之外的所有点都尽可能远,看看能不能找到一个无穷大(∞)的形状。 回复 准则 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

My longtime labmate-turned-student/friend Dhruva Karkada recently wrote a sick paper on data statistics which deservedly went viral on Twitter, in part because it has one of the prettiest scientific figures I have ever seen:

Say you’ve got a bunch of words that live in a vocabulary $\mathcal{V}$. We’re here studying models $f$ that map $f: \mathcal{V} \rightarrow \mathbb{R}^d$: that is, they map every word to a $d$-dimensional vector. We’re letting $\{v_i\}_{i=1}^{12} = \{\texttt{January}, \texttt{February}, \ldots\}$ be the months of the year, taking the 12 associated embedding vectors $\mathbf{w}_i = f(v_i)$, and computing two things:

  • a projection onto the top two PCA directions of $\{ \mathbf{w}_i \}$ (left column), and
  • the Gram matrix $\mathbf{M} \in \mathbb{R}^{12 \times 12}$ such that $M_{ij} = \mathbf{w}_i^\top \mathbf{w}_j$ (right column).

Reading the rows of this figure from top to bottom, they find that:

  1. LLM embeddings project down to a circle (as Engels et al (2024) also saw), and the Gram matrix is approximately a circulant matrix;
  2. these findings are decently approximated even with primitive word2vec embeddings; and
  3. an analytical theory of the circulant Gram matrix gives a very compelling-looking match.

This is a big deal because it connects data statistics to representational geometry with a really simple mathematical theory.

Finding a needle

After seeing this a bunch of times and staring at it for a while, I was feeling in the mood to poke a hole in this beautiful result, and so I bet Dhruva a beer that I could find a collection of other, seemingly-unrelated words that form a circle + circulant matrix in the same way. He (and most others I told) thought this was crazy, since the circle clearly comes from the special relationship between the words. We settled on the terms of the bet: I had to find ten random-seeming words whose word2vec embeddings, when plotted as the above, made a clear and compelling circle.

Why’d I think this was possible? Well, we have vocabulary of $25000$ words to choose from. That gives you $N = \binom{25000}{10} \approx 3 \times 10^{37}$ sets to select from. I figured that if you threw ten darts at a board that many times, you’d definitely make a circle at least once. Info-theoretically speaking, you have $\log_2 N \approx 124$ bits of information, and surely you can make a decent 10-point circle with that amount of resolving power. The question’s just how you find a set of ten good words in the haystack.

Here’s how I did it:

  • From looking at PCA plots of random sets, I guess you’d get a decent circle from a random selection with probability maybe $3^{-10}$, so random guessing could plausibly work.
  • I wrote a “looks circular” objective function, drew tens of thousands of random sets, and chose the best one. It was borderline, but not good enough to utterly obliterate Dhruva.
  • I upgraded it to an iterative search, where at every step, we drop the worst point and choose the best replacement from the vocabulary. That worked pretty well.
  • I also changed the objective from “looks circular on a PCA plot” to “matches a target circulant Gram matrix.” That worked really damn well.

Note that the embedding size $d = 10000$ never entered into this. This all took an afternoon with a coding agent.

Here’s what I got:

That’s circular. You can just find other sets of random-looking words that form circles!

Does this have any actual significance?

This raises certain open questions, including “how can one man be so wrong?”, which I am not qualified to answer.

But seriously: clearly we can find spurious geometric patterns. Should this change our understanding of representation geometry? I’d note a few caveats first:

  1. While the Gram matrix of my spurious-circle-set is indeed beautifully circulant, the amplitude of the (sinusoidal) off-diagonals is less than with the months. I couldn’t get em up to match the months’ Gram matrix, even to within a factor of two.
  2. This works damn well with a set of ten, but I doubt it’d work with a set of, say, 50 (though admittedly I didn’t try very hard), so Dhruva’s other geometric findings (about e.g. all the years from 1700-2020) couldn’t be spoofed in this way.

Nonetheless, it does show that doing a kind of pursuit-matching-style search for a certain low-dim PCA’d geometry will trick you unless you’ve got enough statistical constraints on your target that it won’t happen by random chance! This does rule out certain automatic-feature-finding algorithms, which has implications for research agendas like scalable interpretability.


联系我们 contact @ memedata.com