Skip to content

Testing and Tools

Testing and Tools

Building covers running these to check a build. This chapter is about what they actually verify, and is written for someone changing the engine rather than someone installing it.

The test suite

Around fifty test binaries, one per module, built by default and run with ctest. They are unit tests in the strict sense: each constructs its module directly, with no database anywhere. That is possible because no module below the engine takes a tidesdb_t, and it is the main practical payoff of that rule.

Two suites are integration-level: engine_tests drives the public API end to end, and crash_recovery_tests covers torn-write repair. That one arms a write to tear at a chosen point, crashes a child the instant it lands, and reopens what was left — swept across every crash point in three phases: a flush, the writes a commit itself issues, and a merge. The child is the same binary run again rather than a fork, so it works where there is no fork.

Two are unusual. doc_samples compiles every C sample in this manual against the real public header, so a sample that drifts from the API it documents fails the build. It also checks that every backticked identifier in the prose exists somewhere in the sources, which is what catches a chapter still naming a function that has been renamed away.

portability_create and portability_verify are a pair. The first writes a database, the second reopens it and reads every key back. Locally they run back to back as a round trip; in CI the artifact produced on one architecture is verified on the others, which is what catches a format or endianness divergence. They are built as ordinary targets rather than embedded in the workflow, so a public API change breaks the build: a program that lives inside a CI file is seen by no compiler until the job runs it, and can stop compiling against the API while the jobs keep reporting success. They are also denied the internal include path, so they prove the public header alone is sufficient to use the database.

The fuzzers

Four harnesses, none enabled by default:

Terminal window
cmake -S . -B build-fuzz -DCMAKE_BUILD_TYPE=RelWithDebInfo -DTIDESDB_BUILD_FUZZERS=ON
cmake --build build-fuzz -j

Each is model-based: the harness maintains an independent model of what the database should contain and compares the engine against it. That is what makes them able to find wrong answers rather than only crashes.

fuzz_standalone — the logical model

Generates random operation sequences against both the engine and a versioned in-memory model, then compares. Catches wrong values, lost writes, resurrected deletes, and visibility errors.

The model is version-aware — it tracks sequences and tombstones, because a model that compares only latest values reports failures constantly, all of them its own fault rather than the engine’s. When a model disagrees with the engine, suspect the model first.

It also carries each entry’s expiry, converted from the lifetime at the put exactly as the engine converts it, and hides a version that has outlived it. The lifetimes the harness generates are only ever never or a whole day, deliberately: expiry is decided by comparing against a second-granular published clock, and a lifetime that elapsed part-way through a run would leave the model and the engine on opposite sides of that boundary, making the oracle the thing that failed. What these lifetimes do catch is the direction that matters — an expiry the engine loses or brings forward turns into a key the model still expects and the database no longer has.

The harness now also generates empty values, roughly one in sixteen. An empty value is a value, and the engine must not confuse it with an absent key; this is the oracle that says so under crash and concurrency rather than only in a unit test.

It generates interval deletes too, both the two-bound and the prefix form, with bounds drawn from the same small alphabet the keys use so an interval actually meets them. Modelling one is what forces the transaction buffer to be an ordered log rather than a map: an interval hides what was buffered before it and leaves what was buffered after, so position in the buffer is what decides which of the two is newer, and a read resolves a key by finding the last entry that either names it or covers it.

The conflict oracle for an interval is deliberately not symmetric, because the engine’s is not. An interval is claimed against the other intervals only — there is no one key to hash it under — while a point write is checked against both the intervals and the point reservations. So a point meeting either kind must lose, and an interval meeting another interval must lose, but an interval meeting a held point is allowed through. Mirroring that asymmetry is the difference between an oracle and a source of false failures.

Alongside the value comparison it carries a structural oracle that owes nothing to the model. A closed database owns no sstables and no level-set layouts, so both counts have to come back to what they stood at when it opened — and that is checked at every close the run performs, not once at the end. The granularity is the point: a balance checked once per run says only that something between the first operation and the last leaked, while one checked per close names the twenty-odd operations since the previous one.

The two counts are worth reading together when it fires. A leaked handle with the layouts balanced is a reference some reader took and never gave back; a leaked layout means the sstables it lists are pinned by a structure nothing will reclaim. They fail in the same way and have entirely different causes, and separating them is what stops an investigation committing to the wrong one.

fuzz_crash — durability

Forks a child that commits under a sync barrier, kills it, reopens the database, and checks what survived.

Its oracle is the sharpest of the four: the recovered state must be an exact durable prefix of what was acknowledged. Not a superset — a torn final record must be discarded, not half-applied — and not a subset. It fingerprints the model at every plausible prefix and requires the recovered state to match one of them.

This is the harness that guards the ordering rules in the write path.

fuzz_conc — concurrency

Worker threads over disjoint key partitions, each with its own exact oracle, while the shared memtable, write-ahead log, flush, compaction, cache, and value log all race underneath.

Partitioning the keyspace is what keeps the oracle exact under concurrency: each worker knows precisely what it wrote, so any disagreement is a real engine fault rather than a race in the checker.

fuzz_decode — untrusted input

Feeds mutated bytes to the decoders — sstable footers, column-family configurations, WAL records, partition range filters, btree nodes, and manifest batches.

These are the one place untrusted input enters the engine. Everything else operates on data the engine itself wrote; a decoder operates on whatever is on disk, which after a crash or a bad disk may be anything. A torn or corrupt file must be rejected, never over-read.

Several of them sit behind the block manager’s checksum, which is why the bytes are handed to them directly rather than written to a file first. Going through the file would only ever exercise the checksum: random damage is caught before the decoder runs. What is left to get wrong is the part a checksum says nothing about — a record loop walking offsets and lengths that are internally consistent and still wrong.

Controlling a run

VariableEffectHonoured by
TIDESDB_FUZZ_ITERSIterations to runall four
TIDESDB_FUZZ_SEEDSeed, for reproducing a failureall four
TIDESDB_FUZZ_DIRWhere databases are createdfuzz_standalone, fuzz_crash, fuzz_concfuzz_decode opens no database
TIDESDB_FUZZ_DUMPDump the failing inputfuzz_standalone, fuzz_crash
TIDESDB_FUZZ_DUMP_LASTWrite the last iteration’s input to a file and run only that one, for cutting a single-iteration reprofuzz_standalone
TIDESDB_FUZZ_VERBOSEPer-operation tracingfuzz_standalone
TIDESDB_FUZZ_LOG, TIDESDB_FUZZ_LOG_FILEEngine log level and destinationfuzz_standalone

The last three live in the model harness, which only fuzz_standalone and the coverage-guided fuzz_tidesdb link; fuzz_crash and fuzz_conc build against the model alone and ignore them.

Terminal window
TIDESDB_FUZZ_DIR=/fast/disk TIDESDB_FUZZ_ITERS=200 ./build-fuzz/fuzz_crash
TIDESDB_FUZZ_SEED=519270 TIDESDB_FUZZ_VERBOSE=1 ./build-fuzz/fuzz_standalone # reproduce

A failure prints its seed. Re-running with that seed reproduces it deterministically, which is the first step of every investigation.

Coverage-guided runs

Under Clang, two additional targets instrument the whole library rather than just the harness, so a guided run explores the engine itself: fuzz_tidesdb for the logical model and fuzz_decode_guided for the decoders. The decoders benefit most — they are small and branch heavily on the input bytes, which is what coverage guidance is good at reaching.

The standalone runners need no instrumentation and build with any compiler, which is why CI uses them.

The allocation-failure sweep

An allocation being refused is the one failure a caller cannot arrange from outside, which makes the code that runs afterwards — the half-built structure freed, the lock dropped, the handle never published — the least travelled in the engine. TIDESDB_WITH_ALLOC_FAULT builds an injector that counts the library’s allocations and refuses a chosen one, and alloc_fault_tests measures how many a small workload makes and then runs it once per allocation, refusing a different one each time.

Only the armed allocation fails, never the ones after it, so what each round exercises is one site’s unwind rather than a cascade in which the cleanup paths are also failing. What a round must show is that the database still opens and still holds every commit it acknowledged — a refusal may cost a write that never returned, and may cost the whole round, but never one the caller was told had landed. Run it under the address sanitizer, which is what turns an abandoned partial structure into a failure rather than a leak nobody sees.

The workload writes to two column families rather than one, which is not incidental. A flush demuxes the shared memtable into one output per family, so a single-family workload produces a single output and can never reach the unwind that returns an already-installed output’s build reference when a later one cannot be installed. The reachable set of a sweep like this is bounded by what the workload actually does, and a path the workload cannot enter is not covered no matter how many allocations are refused.

It needs a linker that can rewrite a symbol, so it is a build option rather than something always present; see Building.

What CI enforces

Beyond the platform matrix, four gates are worth knowing about because they fail builds that otherwise look fine.

Zero warnings. -Wall -Wextra are always on but nothing makes them fatal, so a warning can accumulate unnoticed. One lane compiles the library, the tests, the fuzz harnesses and the benchmark driver with warnings as errors. It is the only job that fails on one, and it is the only job that compiles the bench driver at all.

Formatting. Every first-party source directory is checked, not just src — the public header and the test, fuzz and bench trees are maintained code too. Only vendored sources are exempt.

Fuzzing. All four harnesses run bounded, with the seed drawn from the run number so each build explores new ground while any failure reproduces from the log.

Thread sanitizer. It cannot be combined with the address sanitizer, so without its own lane the data races it finds have no coverage anywhere.

The benchmark driver

Terminal window
cmake -S . -B build-bench -DCMAKE_BUILD_TYPE=Release -DTIDESDB_BUILD_BENCH=ON
cmake --build build-bench -j
./build-bench/tidesdb_bench --benchmarks=fillrandom --threads=8 --num=1000000 --dir=/fast/disk

Workloads

WorkloadWhat it does
fillseqSequential inserts
fillrandomRandom inserts
readrandomPoint reads of present keys
readmissingPoint reads of absent keys — the filter path
readwriteMixed, split by --read_ratio
deleterandomRandom deletes
deletewhilereadDeletes concurrent with reads
scanRange scans of --scan_length
scanrangeThe same scans told their range up front, so the two forms can be compared on one dataset
scanemptyBounded scans of a range holding no keys, in the gap between two written keys — what it costs to prove a range empty
fillindexBuilds a table family and an index family together, one row and the index entry pointing at it per transaction. Needs --cfs=2; it is what indexscan and indexmixed read
indexscanA secondary-index ref join: seek the index family, then a point get on the table family per matching entry, which is what the MariaDB plugin’s index_read_secondary issues. One op per lookup, since the lookup is the unit being measured. --index_fanout sets rows per indexed value
indexmixedThe same lookups running concurrently with writes, split by --read_ratio. The one to use for an index measurement that should mean anything — see the caution below

Several may be given at once: --benchmarks=fillrandom,readrandom,scan.

Options that change what you are measuring

OptionEffect
--threads, --numConcurrency and operation count
--key_size, --value_sizeRecord shape. The single most consequential setting — a value either side of value_separation_threshold exercises entirely different paths
--dist, --zipf_thetaKey distribution: uniform, zipfian, sequential
--syncDurability mode; the largest single lever on write throughput
--isolation0–4; snapshot and above add conflict detection
--cfs, --txn_opsColumn families, and operations per transaction
--memtable_size, --cache_size, --l0_stallThe engine knobs most worth sweeping
--flush_threads, --compaction_threadsBackground pool sizes
--bloom, --bloom_fprFilter on or off, and its false-positive target
--secondsRun each workload for a wall-clock duration instead of to its operation count. Prefer it: a run bounded by operations takes whatever time it takes, and at high thread counts that is often a fraction of a second
--compressionThe column families’ encoding pipeline — none, snappy, lz4, lz4fast, zstd. A store that compresses is a different engine from one that does not: every key-log node is encoded on the way out and decoded on every read that misses the cache, so a benchmark left at none does not describe a deployment that sets it
--index_fanoutRows sharing one indexed value, for fillindex and the index workloads
--rateHold total throughput to N operations per second. Required for any question about the store’s shape — see below
--seedReproducibility

Options that change the tree’s shape

These drive the compaction structure itself, which is what a sweep of the layout varies. Each one left unset keeps the engine’s own default, so an unswept run stays comparable to a swept one.

OptionEffect
--level_size_ratioT, the ratio between successive level capacities. Smaller deepens the tree sooner and raises write amplification
--dividing_level_offsetHow far above the largest level the dividing level sits. 1 means X = L - 2
--min_levelsFloor on the level count
--l1_triggerFlush-tier runs that make a compaction due
--tombstone_density, --tombstone_min_entriesThe density trigger and the entry count below which it is ignored
--value_separation_thresholdWhere values spill to the value log. Pairs with --value_size
--btree_block_sizeTarget klog node size
--idle_flush_secondsTimer that rotates an idle non-empty memtable
--skip_list_max_level, --skip_list_probabilityMemtable skip-list shape

A worked contrast, on the same 2M-key load: the defaults (T=10) settle at two levels with read amplification 2.0 and write amplification near 1.5, while --level_size_ratio=5 --l1_trigger=2 reaches three levels but pays read amplification 6-7 and write amplification 4.4. Deepening the tree is not free, and at a given data size it may buy nothing.

Time series output

--sample_interval_ms=N starts a background sampler writing tab-separated series into the data directory, one file per stream:

FileContents
throughput.tsvOperations per second over time
resource.tsvProcess CPU, resident memory, and live heap
dbstats.tsvDatabase statistics, including the backpressure counters
cfstats.tsvPer-column-family statistics
cachestats.tsvBlock cache hits, misses, residency

Every row is flushed as it is written. A run that is killed — by the OOM killer, by a watchdog, by an operator who has seen enough — is exactly the run whose series matters, and a buffered stream loses the tail of it, which on a short run is the whole file.

Each plots directly. The series matter more than the summary for anything involving flush or compaction, because the interesting behaviour is a stall partway through a run and a single average hides it.

Reading the results

Three rules, all learned the hard way:

Interleave your A/B. Run baseline and variant alternately in one session. Comparing against a number from last week measures the machine, not your change.

Take enough samples for the metric. Throughput stabilises within a few runs; tail latency does not. Repeated runs of one unchanged build can spread tail percentiles wide enough that a two-sample comparison tells you nothing, so take several and compare medians.

Run long enough to leave the drive’s write cache. A short run on a consumer SSD measures its write cache rather than the drive, and the two differ by orders of magnitude. Decide which regime you care about before choosing a size, and before concluding anything about the engine from a sustained run, write the same volume with dd to the same filesystem — that control tells you whether the device was the limit rather than the engine.

Also: benchmark a Release build, never a sanitizer build, and keep the data directory off any disk something else is using.