Simon Eskildsen, turbopuffer: a search engine that keeps its indexes in object storage
Vector and full-text search live in RAM by tradition, so the bill grows with the corpus whether anyone searches it or not. turbopuffer keeps the durable copy in object storage, caches on SSD, and prints both latencies: a p90 of 444 ms cold and 10 ms warm over a million vectors. Profiled from Simon Hørup Eskildsen's own post.
Photo: graph8 In late 2022 Simon Hørup Eskildsen was helping Readwise scale its Reader app, and the team priced vector search across its 100 million-plus documents. It would have cost $20,000 a month or more, against $5,000 a month for the relational database that ran the rest of the product. The feature was shelved, and turbopuffer, the company he went on to co-found, was built to bring that first number down (“turbopuffer: fast search on object storage”, turbopuffer.com, updated 2026-03-05).
Eskildsen co-founded turbopuffer and runs it as CEO, and the post this profile draws on is his, with the talk he gave to the CMU Database Group, “Object storage-native database for search”, as a secondary source. Every turbopuffer number below comes from the post. The numbers about our own node come from the earlier articles this week.
Why search indexes live in memory
A search index has to answer in milliseconds, so the tradition is to keep all of it in memory. A vector index holds every embedding and the structure that connects them, and a full-text index holds the posting lists. Both are sized by the corpus and priced at what memory costs, on as many machines as the durability requirement demands. The bill rises with the document count whether or not anyone searches most of those documents.
The design in the post bets that only a share of them is hot at any time. A reader app with a hundred million documents runs its queries against a small hot set while the rest sit in RAM at the same price. Where the durable copy also has to sit on replicated SSDs, the same corpus is paid for again on every replica.
An LSM tree over object storage
The post describes the design one layer at a time. Object storage, S3 or GCS, is the source of truth, written through an LSM tree. SSD and memory on the query nodes are caches for hot queries. Any node can compact data and serve any namespace, because no node holds state that object storage doesn’t also hold. There are no triply replicated disks to keep in step, because the bucket holds the copies.
An LSM tree is a way of writing data that never goes back to edit a file. Every new batch of writes lands in a fresh file, and a background job merges the small files into bigger ones over time. That fits object storage, where a file is written once and cannot be changed afterwards, only replaced.
Eskildsen tabulates what a terabyte of index costs to hold for a month at each point on the ladder, from everything in memory to everything in the bucket.
| Where the index lives | Per TB-month, as the post tabulates |
|---|---|
| RAM plus 3x SSD | $3,600 |
| RAM cache plus 3x SSD | $1,600 |
| 3x SSD | $600 |
| S3 plus SSD cache | $70 |
| S3 alone | $20 |
turbopuffer sits on the fourth row, and the top row costs about fifty times as much, $3,600 against $70 on the post’s table.
The price of the fourth row is paid on the cold read, and the post prints that number too. Over one million 768-dimension vectors, a query whose data is not in cache answers at a p90 of 444 milliseconds, and the same query warm, with its data already on the SSD cache, answers at a p90 of 10 milliseconds. Full-text search with BM25 over one million documents runs at 285 milliseconds cold and 18 milliseconds warm at p90. Dividing each pair, cold is about 44 times slower for vectors and about 16 times slower for text.
By his account the design now runs at more than a trillion documents, more than ten million writes a second, and more than 25,000 queries a second. The aim he states there is to let customers “search every byte they have” (“turbopuffer: fast search on object storage”, updated 2026-03-05).
What transfers to ClickHouse on R2
Two things carry over to our work. The first is the cold-against-warm pair, which we can’t produce for our node yet. Our object-storage node keeps every part of every intent table in a Cloudflare R2 bucket behind a 2.5 TB NVMe cache, as described in the architecture article. We don’t know our cache hit ratio yet, because the counters sit behind a grant our reporting user doesn’t have, and the time of a first uncached scan is one of the stages in the restore drill, which hasn’t run. When the drill has run, turbopuffer’s 444 milliseconds against 10 is the pair to print next to that scan. We should expect a worse number, because our node is a general-purpose database on a 1 Gbit port reading a bucket an ocean away from the compute, and only the drill will say how much worse.
The second is the boundary itself. turbopuffer’s design draws a line through the index. Above the line are the slabs the cache holds and answers from in milliseconds. Below it is everything else, in the bucket, at the fourth row of the table. Our cache disk draws the same line through the set of parts. Caching on write puts a fresh part above the line, and eviction moves it below. The keyword resolver’s hot tables are not on that line at all. They sit on local_fast, the plain NVMe policy, outside both the cache and the bucket, because a 50 millisecond lookup cannot absorb a round trip to the bucket on a miss.
What does not transfer
turbopuffer owns its file format and its compaction. The LSM tree, the size of a slab, when a node folds small objects into large ones, and what a query fetches on a miss are all the company’s own code, written for search and for the bucket. When a cold read is too slow, the fix can go into the engine.
We run ClickHouse on the path its vendor documents for open source. MergeTree decides what a part is and when it merges, and we set the levers the settings expose: cap the number of concurrent merges, fix every non-trailing multipart upload part at 128 MiB because R2 rejects an upload whose non-trailing parts differ in size, cache a part as it is written, and accept the read cost that comes with each choice. If a cold read is too slow for us, the fix is another setting, because the engine isn’t ours to change. The vendor’s cloud-only engine, SharedMergeTree, keeps the map from part to objects in Keeper instead of on each server. The open-source engine we run keeps that map on the machine’s local disk, so we back it up every night.
Our extension: the boundary for revenue search
A revenue platform searches three kinds of record: contacts, companies, and the signals that tie them to a moment in time. The searches we expect to see land on a small share of each: the accounts in an active sequence, the contacts a workflow touched this week, the signals from the last few days. The long tail, years of history on companies nobody is working right now, is what makes an in-memory index expensive, and it is exactly what a bucket holds well. The boundary we intend to draw runs by time and by activity: recent signals and active accounts on the cache, everything else in the bucket, and the query layer told which side a read will land on so that a workflow trigger never waits for a cold read. We have not built any of it yet. The intent tables are on the object-storage layout today, search over contacts and companies is not, and the boundary is a design note.
Still open
Two questions go into the courtesy note to Eskildsen. One is the consistency model, meaning what a reader should expect between a write and its visibility to a query on another node. Any team pricing the design needs that before it reads the storage table, and the post does not state it (revision of 2026-03-05, read 2026-09-08). The other is distance. Our bucket sits an ocean from the node that queries it, so every cache miss crosses that distance, and the post doesn’t say whether turbopuffer has run into the same problem.
Builder Spotlight profiles an engineer outside graph8 from their own published work only, and a courtesy note goes to the subject before publication.
Further reading
- Simon Hørup Eskildsen, turbopuffer: fast search on object storage, turbopuffer.com, updated 2026-03-05: the post every turbopuffer number above comes from.
- Simon Hørup Eskildsen, Object storage-native database for search, CMU Database Group talk: the secondary source.
Keep cold data where storage is cheapest and hot data where reads are fastest, and decide deliberately which is which.