InfoQ 中文 · 10/8/2026, 2:18:35 AM
Perplexity Replaces DynamoDB with Rust-based CobbleDB for LLM Search
Perplexity migrated its core search service from Amazon DynamoDB to CobbleDB, a custom Rust-based distributed key-value store, to address latency and cost bottlenecks in high-throughput LLM retrieval. By decoupling persistent storage from the hot data layer, the team achieved a fivefold reduction in batch read latency and cut storage costs by at least 20%. The system was built in two months by two engineers working with AI coding agents and is planned for open-source release.
SOURCE COVERAGEOriginal coverage
Perplexity has migrated its core search service from DynamoDB to CobbleDB, a self-developed distributed KV store written in Rust, to address latency and cost challenges posed by high throughput and large payloads in LLM retrieval scenarios.
Decoupling the persistence layer from the hot data layer reduced batch read latency to one-fifth of its original value; storage costs dropped by at least 20%; and it mitigated the issue of uncontrollable tail latency caused by DynamoDB’s black-box architecture.
This article is suitable for search architects, distributed systems engineers, and AI infrastructure engineers.

Perplexity has migrated its core search service layer from Amazon DynamoDB to CobbleDB, an internally developed distributed key-value store written in Rust. This migration aims to resolve severe latency and cost bottlenecks arising from delivering batches of multi-kilobyte documents to Large Language Models (LLMs) under high query volumes. By decoupling persistent document storage from hot data layer retrieval, the engineering team reduced batch read latency by fivefold while lowering overall storage costs by at least 20%.
The read patterns of AI answer engines differ significantly from traditional document search. Each query sent to Perplexity generates 100 to 120 target page keys, which the retrieval service splits into multiple batches of 10 to 20 keys processed in parallel. While traditional search engines return short metadata snippets, retrieval for language models requires extracting full chunked paragraphs and dense vector embeddings, resulting in an average record payload of approximately 50 KB.
At production traffic scales exceeding 200,000 requests per second, DynamoDB’s pay-per-use billing model became economically unsustainable, as AWS charges for every byte transferred. Furthermore, DynamoDB operates as a black box, with internal partition placement, memory caching strategies, and replica routing remaining opaque. Engineers could not mitigate tail latency spikes caused by uncached reads, cross-Availability Zone network hops, or replica lag. Reprocessing jobs triggered by chunking algorithm updates or the adoption of new embedding models also pushed massive write loads directly into DynamoDB, creating noisy neighbor contention with real-time user requests.
To address these limitations, Perplexity split its storage architecture into three specialized systems: Pillar for persistent state management, Lorry for batch aggregation, and CobbleDB for low-latency serving.

Pillar is built on YTsaurus running on high-capacity mechanical hard drives, maintaining versioned table families for web page metadata, paragraphs, and vector representations. YTsaurus atomic transactions ensure that crawl updates, state changes, and export queues are committed together. Lorry acts as a stateless queue consumer, combining Pillar’s exports into partition-aligned batch files, storing payloads in Amazon S3, and sending metadata notifications to CobbleDB. CobbleDB workers independently pull and ingest these S3 batches, completely isolating hot data serving nodes from the write-intensive crawling pipeline.
CobbleDB is a distributed key-value store specifically optimized for batch lookups. Each partition maintains three replicas across independent compute nodes. The core daemon uses RocksDB as the embedded storage engine, combined with a memory-mapped cache and local NVMe SSDs.
A stateless query router maps hashed page identifiers to partitions and coordinates read execution. To minimize network overhead, the router sends requests to node replicas located within the same Availability Zone. If the response time of the target replica increases, the router performs speculative hedging, simultaneously initiating read requests to backup replicas on other nodes. Within each node, CobbleDB retrieves multiple keys concurrently via RocksDB’s batch MultiGet interface, eliminating round-trip overhead.
pub struct BatchedPageRequest { pub keys: Vec<PageKey>, pub zone_affinity: AvailabilityZone,}impl StorageEngine { pub fn multi_get_pages(&self, keys: &[PageKey]) -> Result<Vec<Option<PageRecord>>, Error> { let rocksdb_keys: Vec<&[u8]> = keys.iter().map(|k| k.as_bytes()).collect(); self.rocksdb.batched_multi_get(&rocksdb_keys) }}Copy Code
The database discards standard distributed transaction protocols and synchronous consensus algorithms. Since the search service can tolerate slight replication lag, individual replicas apply updates asynchronously at their own pace, significantly reducing operational overhead.
In actual production measurements, CobbleDB reduced median batch read latency from 31.4 ms to 5.60 ms, p90 latency from 56.7 ms to 9.77 ms, and p99 tail latency from 123 ms to 24.2 ms. Synthetic benchmarks handling payloads up to 100 KB demonstrated stable throughput even at scales of up to 500,000 requests per second.
![](/api/v1/file/hot_img_01e63542c1386f0740c111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c50111a69c5
CEO Aravind Srinivas stated that CobbleDB consists of approximately 40,000 lines of Rust code and was built in two months by two systems engineers working alongside an autonomous AI coding agent cluster. These AI agents handled integration testing, build monitoring, and operational runbooks. Perplexity announced plans to open-source the CobbleDB codebase in upcoming releases.
Original article: https://www.infoq.com/news/2026/09/cobbledb-perplexity/