Keys and Values Don't Always Belong Together
published on October 9th, 2026
Let’s talk keys and values. Most all databases in the world are built on this foundation, storing raw bytes and indexing them in the form of a key and its value once called upon.
There are many data structures to achieve this, today we will focus on a normally persisted one called the log-structured-merge-tree.
In an LSM tree engine, writes are buffered at Level 0, flushing to sorted files at Level 1 then compacting up levels. Those sorted files usually have K, V, K, V, side by side often described as inline. That can be fine for small values, but for large values, this can cause severe write overhead.
An LSM tree keeps its levels sorted by merging them as it works, rewriting whatever each merge reads. Doing that for 16 byte keys is not a big deal, but when the values riding along are 4 KB, every merge writes them out again even though they never change. That phenomenon is called write amplification, the bytes the engine writes to disk for every byte you wrote, and it’s what separating values is meant to cut.
The fix, laid out by the WiscKey paper (Lu et al., FAST ‘16), is to keep only keys in the tree and put values somewhere they’re written once and left alone. In TidesDB a value at or above value_separation_threshold goes to a value log shared by the whole engine, and the key log(klog) stores a small logical id in its place. Merges then move keys and ids, never values.
In TidesDB, a value is appended to the value log once — at commit — and from there onward, every layer just carries the key and the id, so the write-ahead log record, the memtable entry and the key log never copy the value.
In TidesDB an SSTable (sorted string table) is essentially a klog which in itself is a B+tree. A key log node is 4 KB by default, set with btree_klog_block_size, and the default large value threshold is a quarter of it (1 KB) on purpose — since a 4 KB value kept inline fills a node by itself.
I thought I’d measure inline with 64 KB nodes to see how that plays out! A program against TidesDB 10 loads 250k keys with 4 KB values in random order, overwrites them all once more, waits for compaction to finish, then measures 100,000 point reads and a full scan to see how they perform.
Separated writes run almost three times faster, and the compaction panel shows they rewrote nothing.
The 64 KB node run still rewrote 5.1 GB, because compaction pays for the bytes it moves, however large nodes do make scans the fastest of the three, and point reads three times slower, since every lookup now decodes a 64 KB node.
There is a cost of separating, of course, a scan pays a second read per row, and old values stay on disk until compaction drops the keys pointing at them. Reclaiming that is cheap in TidesDB, since every key log records which value log segments its values live in, thus a segment nothing references is deleted whole and a half empty one is drained by the next compaction that was going to run anyway.
So for large values written often and read by key, separating them is a big win, and for tables you mostly scan, keep_values_inline with bigger nodes is the better fit.
If you are curious about how TidesDB compares to RocksDB utilizing a tool I wrote called keybench, especially comparing against the BlobDB configuration you can find that here.
Thanks for reading!
—
Thank you to Amar Sood (@tekacs) for proofreading this article and putting his own twist on it.