turbopuffer is killing the vector-primary index. In a post published September 30, engineer Dan Harrison laid out the company’s plan to rebuild its storage layer so that the approximate nearest neighbor (ANN) index becomes “just another” secondary index, rather than the key around which every other index and query plan revolves. The headline is architural. The subtext is that the vector database category, as it has been sold for the past three years, was a transitional shape.

The company says all of its CI now passes on the new architecture, which it calls v3. It has not shipped to production. It is starting a public performance-grinding phase and will publish benchmarks over the coming weeks, targeting parity and beyond before rollout.

What turbopuffer actually built

turbopuffer’s v1 was a serverless vector database. Documents were an ID and a vector. The index was hierarchical clustering, chosen because it plays better with object storage than the graph-based indexes that were conventional wisdom at the time. It started with SPANN and migrated to SPFresh for incremental indexing. Every key was addressed by ClusterId and LocalId, which the company calls the ANN address.

v2 added attribute filtering and BM25 full-text search, both modeled as inverted indexes that map values to ANN addresses. Then aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering. All of it built around the same vector-primary layout.

That layout worked. turbopuffer says it has pushed vector search to single indexes of 100B+ vectors, serving 200 ms p99 reads at 1k+ QPS. Cursor and Notion were early customers. Linear uses it as a syncing engine. The company now claims 1T+ documents hosted, 10M+ writes/s, and 25k+ queries/s.

Three reasons the primary index had to go

Harrison names three failure modes, and they are worth reading as a general diagnosis of vector-first infrastructure.

Storage amplification. The full contents of each document sit under its ANN address. For multi-vector representations, like document nesting or late interaction, the contents get duplicated once per vector. That is the source of some of turbopuffer’s documented limits.

Write amplification. When a document is inserted, updated, or deleted, SPFresh may rebalance vectors to keep clusters tight. Because everything is keyed by the ANN address, that rebalancing drags the full document and any inverted indexes that reference it. Updating one vector can move hundreds of attributes and their indexes. The company says indexing-throughput tuning has started hitting diminishing returns.

Limited vectorization. Query engines run tight loops over blocks of values. DuckDB batches 2,048 rows. ClickHouse goes up to roughly 65k. Lucene’s posting blocks are 256 docs. turbopuffer’s ANN index clusters best at 100–200 documents. Every query plan inherits that block size, even plans that would prefer thousands.

The full-text search history is the proof. In v1, postings were partitioned along ANN cluster boundaries, and the median block held about 1.5 postings. In FTS v2, postings moved to fixed blocks of roughly 256, stored separately and pointing at documents. The index got 10x smaller and queries got up to 20x faster. Posting lists could escape the cluster layout. Aggregations and scans could not, because they read documents stored one block per cluster.

The fix is one sentence: stop keying on the ANN address.

The category label was always a bet on one query shape

The interesting thing here is not that turbopuffer is changing its storage. It is what the change implies about the label. “Vector database” described a product built around a single access pattern, and it made sense when the dominant AI workload was embedding lookup for retrieval-augmented generation. That workload is now one of several. Agents issue regex and fuzzy matches over tool outputs. Coding assistants run hybrid search over repositories. Applications want filters, aggregations, and ordering over the same corpus they embed.

When the vector index is primary, every other query plan pays a tax. turbopuffer quantified the tax and decided to stop paying it. The company is not abandoning vector search. It is demoting it, which is a different and more consequential move.

When the vector index is primary, every other query plan pays a tax. turbopuffer quantified the tax and decided to stop paying it.

There is a read here that the standalone vector database was a category error, a point solution that should have been a feature of a general search or analytics engine. The evidence is not only turbopuffer’s post. The major cloud warehouses have all added vector types and ANN indexes to existing engines over the past two years. Postgres extensions do the same. If the vector index is just another secondary index, the argument for a dedicated system narrows to economics and operational simplicity, which is precisely what turbopuffer’s object-storage-first design was selling in the first place.

What it means for AI builders

Two practical consequences.

First, if you are choosing retrieval infrastructure now, ask what happens when your query mix changes. A system that only does ANN well will force you to bolt on a second store for filters, aggregations, or lexical search, and then you own the consistency problem between them. A system where ANN is one index among several can absorb the new query shape without a migration.

Second, watch the benchmark numbers turbopuffer says it will publish. The claim on the table is that removing the ANN address as primary key unlocks “significant performance improvement on all query plans,” including vector search itself. That last part is the hard part. The company admits any change to the layout risks regressing ANN performance, which is why it kept the design for so long. If v3 matches v1 on vector search while beating it on aggregations and scans, the vector-primary era is over on the merits, not just on principle.

The open question is whether the rest of the category follows. Turbopuffer is a small company with an unusual architecture, and its competitors have their own storage bets. But the diagnosis in the post is not specific to turbopuffer. It is a description of what happens when you optimize a system around one workload and then the workload multiplies. Every team that shipped a vector-first retrieval stack in the last three years is now running that same experiment, whether they know it or not.