RIP,向量数据库
RIP, vector database

原始链接: https://turbopuffer.com/blog/rip-vector-database

turbopuffer 正在重新设计其存储架构(“v3”),以支持更多查询方案、更大的规模和更好的性能。 其原有架构将 ANN 向量索引作为主键:文档存储在基于聚类的 ANN 地址下,属性和全文检索倒排项则引用这些地址。这种方式对向量搜索非常有效,能够支持超过 1000 亿个向量的索引,在每秒超过 1000 次查询(QPS)的情况下实现 200 毫秒的 p99 读取延迟,但也带来了三个主要限制: - **存储放大:** 多向量文档会重复存储非向量数据。 - **写入放大:** ANN 负载均衡会迫使文档数据和二级索引一起迁移。 - **向量化受限:** 查询引擎受 ANN 聚类大小限制,无法使用更大、更高效的块。 turbopuffer v3 不再将 ANN 地址作为主键。文档改为独立存储,ANN 则像属性和全文搜索一样,成为二级索引。这样,每种查询方案都可以采用针对其执行模式优化的布局。 这次重新设计并不简单。尽管所有持续集成(CI)测试目前都已通过,但新架构最初导致生产环境性能明显下降,因为团队优先保证正确性和基础设计,而非性能调优。现在,团队已经开始进行性能优化,并计划在改进过程中持续发布基准测试结果。

Turbopuffer v3 放弃了以向量为主键的存储方式,转而采用更传统的设计:文档使用稳定的内部 ID,ANN 则作为二级索引。这类似于 PostgreSQL 与 InnoDB 之间的取舍:PostgreSQL 通过直接指向行来换取更快的读取速度,而 MySQL 则让二级索引指向稳定的主键,以提高写入效率。 支持者认为,这一改动减少了 ANN 重新平衡导致的高成本索引重写;批评者则质疑额外的一层查找是否会增加延迟,尤其是在冷存储的 p99 延迟方面。开发者表示,各集群可以在本地保留向量,因此只有获取结果时才会产生这层间接访问。 许多评论者认为,“向量数据库已死”只是夸大的营销说法。向量搜索正在成为更广泛检索系统中的一种查询原语,而传统数据库和搜索引擎也越来越多地支持它。独立的向量存储还会带来同步和运维复杂度。 另一个讨论分支则质疑,AI 生成的评论是否正在破坏 Hacker News 的讨论质量,以及采用邀请制准入的社区是否更能遏制垃圾信息。
相关文章

原文

September 30, 2026Dan Harrison (Engineer)

We are changing turbopuffer's storage architecture. The new engine, informally called "turbopuffer v3", changes how documents and indexes are laid out, written, compacted, and queried in turbopuffer. It will allow us to serve many more query plans, at much greater scale.

turbopuffer launched as a serverless vector database, highly specialized to the task of serving extremely cheap and reasonably fast vector searches. Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance. The value of these particular tradeoffs was validated by our earliest customers, including Cursor and Notion.

Over time, turbopuffer has become a generalized database for search (and other things). The query engine has evolved along the way to support all of these query plans, but the storage architecture has remained largely unchanged: the ANN vector index was and still is the primary index around which all other indexes and query plans revolve.

We've pushed the primary ANN index as far as we can, but it's time to move on. We're in the process of moving to a new primary index, and making ANN "just another" secondary index. We thought it might be fun to open up the doors and let you follow along.

For this first update, we'll set the stage with why we're doing this in the first place. Walk with me on a short journey from tpuf v1 to today.

v1: an ID and a vector

In the first version of turbopuffer, documents consisted of nothing but an ID and a vector. The prevailing wisdom at the time was graph-based vector indexes, but a hierarchical clustering index plays better with object storage. We started with SPANN, and eventually migrated to SPFresh to support incremental indexing. Vectors are clustered into groups, whose centroids are clustered in turn, repeated to form a tree with a single root.

      ┌───────────────┐
      │ root centroid │
      └───────────────┘
        ╱     │     ╲
       ╱      │      ╲
┌────────┐┌────────┐┌────────┐
│  leaf  ││  leaf  ││  leaf  │
│centroid││centroid││centroid│
└────────┘└────────┘└────────┘
   ╱  ╲      ╱  ╲      ╱  ╲
┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐
│vec││vec││vec││vec││vec││vec│
└───┘└───┘└───┘└───┘└───┘└───┘

We implemented this on top of a storage layer presenting as a key-value map, with sorted and unique keys. Each cluster is given a ClusterId, and vectors within each cluster are given a dense LocalId.


K::Vector(C0L0) = vec![0.45, 0.32, ...]
K::Id(C0L0) = 7
K::Vector(C0L1) = vec![-0.28, 0.96, ...]
K::Id(C0L1) = 13


K::Vector(C1L4) = vec![0.64, -0.48, ...]
K::Id(C1L4) = C0

As you can see above, everything is keyed by ClusterId and LocalId (e.g. C0L1), which together we call the ANN address. This is what we mean when we say the ANN index is the primary index.

Two new query plans marked the informal transition from turbopuffer v1 → v2: attribute filtering and full-text search.

Attribute filtering

Naturally, customers wanted to be able to add attribute values and filter vector searches on them. To make filtering fast and high-recall, we modeled these as an inverted index that maps an attribute value to the ANN address of the documents that contain it.

K::AttrIndex("family", "Alcidae") -> vec![C0L3, C1L2, C1L3, ...]
K::AttrIndex("genus", "Fratercula") -> vec![C0L3, C1L2, C1L9, ...]

For projections (include_attributes), we also stored the document attributes alongside the ID and the vector.

K::Vector(C0L0) = vec![0.45, 0.32, ...]
K::Id(C0L0) = 7
K::Attr(C0L0, "family") = "Alcidae"
K::Attr(C0L0, "genus") = "Fratercula"

BM25 full-text search was another obvious and much-demanded query plan. Similar to attribute search, full-text search works by first finding the documents that have the query term present (commonly called "postings"). For an FTS index, we also include the (term count, document length) metadata necessary for BM25 scoring:

K::FTS("description", "Atlantic") -> vec![(C0L0, 2, 37), (C9L4, 1, 42), ...]
K::Attr(C0L0, "description") -> "A sharply dressed black-and-white seabird with a \
huge, multicolored bill, the Atlantic Puffin is often \
called the clown of the sea. It breeds in burrows on \
islands in the North Atlantic, and winters at sea."

Over time, we've shipped several other index structures and query engines: aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering — all built around the same vector-primary storage layout.

The problem with a vector primary index

The ANN primary index has largely remained intact until today for one simple reason: it works really, really well for ANN search on object storage. On top of this architecture, we've pushed vector search to single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS. Any significant change here risks introducing regressions in ANN performance.

However, this layout holds us back from being state-of-the-art for the non-vector query shapes we support, in three main ways: storage amplification, write amplification, and limited vectorization.

Storage amplification

As described above, turbopuffer currently puts the full contents of each document under its ANN address. When there is only one vector, the non-vector data is stored alongside the vector only once.

However, for multi-vector representations of a document, such as document nesting or late interaction, this means we have to duplicate the contents for each vector. This is the reason for some of our more unfortunate limits.

Write amplification

Any time a document is inserted, updated, or deleted, SPFresh may rebalance the vectors to ensure they remain well clustered (otherwise recall may suffer). Because everything in a document is stored keyed by the ANN address of the document's vector, this rebalancing cascades to moving the full document contents, as well as any inverted (attribute and FTS) indexes that reference it. Updating just one vector can move hundreds of attributes and their indexes.

This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

Limited vectorization

Modern query engines are vectorized: they run tight loops over blocks of values, which amortizes fixed per-block costs, compresses better, keeps the CPU pipeline full, and unlocks SIMD. DuckDB, for example, works in batches of 2,048 rows, ClickHouse up to ~65k, Lucene's posting blocks are 256 docs, and our ANN index works best with clusters of around 100–200 documents. Every query plan has an optimal block size, but today they are all constrained by the ANN primary index. A plan that wants blocks of thousands of documents to keep the CPU saturated is still stuck at 100–200.

We've already documented how much this matters in turbopuffer. Our first version of full-text search partitioned posting lists along ANN cluster boundaries, and the median block held just ~1.5 postings. FTS v2 reworked postings into fixed blocks of ~256, and the index got 10x smaller and queries got up to 20x faster. Posting lists could do that because they're stored separately and point at documents, so their layout doesn't have to follow the clusters. Aggregations and other scans read the documents themselves, and those are stored one block per cluster. As long as the ANN address is the primary key, their block size is constrained to the cluster size, even if they'd prefer something larger.

RIP, primary vector index

The solution to these problems is simple: don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

We recently hit a major milestone: As of earlier this month, 100% of CI passes on turbopuffer v3. However, it represented a significant performance regression from production turbopuffer. That's not a surprise! So far, we've focused primarily on the foundational design and correctness, and we have just started tuning it. Watching the numbers go down is great fun, so we wanted to get you in at day zero of perf grinding.

We'll add results to the benchmarks as we write up what changed. Along the way, we'll dive into detail about the new architecture and the optimizations that will ultimately make turbopuffer faster for more query plans in prod.

turbopuffer

turbopuffer is a fast search engine that hosts 1T+ documents, handles 10M+ writes/s, and serves 25k+ queries/s. We are ready for far more. We hope you'll trust us with your queries.

Get started
联系我们 contact @ memedata.com