InventDB
All articles Reliability

What happens to your data when the power goes out

When an InventDB instance stops without warning, every write it acknowledged is still there when it starts again. This article follows that next start step by step: how the engine knows the last stop was a crash, what it replays, what it re-indexes, and what it deliberately leaves for later.

A database process can stop at any instant. The host loses power, the kernel kills the process to reclaim memory, or a disk write is half finished when everything halts. Two questions matter afterwards: which writes survived, and whether the database opens again quickly and correctly.

InventDB answers the first question with a rule and the second with a recovery pass whose cost is bounded.

What an acknowledged write means

Every insert, update, delete and transaction commit is written to the write-ahead log and fsynced before the call returns. A transaction's entries are followed by a commit marker in the same batch, and recovery applies a transaction only when it finds that marker.

The rule follows from that: a write that was acknowledged survives a crash, and a transaction is restored whole or not at all. Our transactions article covers the commit path in detail.

Sources of truth and derived structures

Each type is stored as a series of segments, and a segment holds two kinds of data. The record file, an append-only stream of encrypted JSON records, and the id index that maps each id to its place in that file are the sources of truth, and they are fsynced as they are written. Everything else is derived from them: the B-link trees that index each field, the count trees behind aggregates, the zone maps and the Bloom filters. Derived structures are updated without a disk barrier on every change, and the checkpoint that runs every five seconds makes them durable in batches.

Sources of truth, fsynced on every write, beside derived structures that are made durable at each checkpoint and can be rebuilt from the sources SOURCES OF TRUTH fsynced before a write is acknowledged Write-ahead log Record file, append-only Id index DERIVED, REBUILDABLE made durable at each checkpoint Index trees, one set per field Count trees and zone maps Bloom filters over ids
A crash can leave the structures on the right behind the records on the left. Recovery rebuilds what is missing from the sources.

That split keeps writes fast, and it decides what a crash can do. After a kill, the records are all present, but some index entries for the newest of them may not have reached the disk. Recovery is therefore not about salvaging data. It makes every derived structure agree again with records that were never in doubt, in time proportional to the writes since the last checkpoint rather than to the size of the database.

Knowing that the last stop was a crash

On an orderly stop, whether an operator stops the instance, a new version is deployed or an idle instance goes to sleep, the server catches the stop signal, runs a final checkpoint and syncs every store. Only if that sync succeeds does it write a small marker file. The next start reads the marker and removes it. A missing marker means the last process did not finish in order, and so does a failed final sync, because the marker is never written after one.

A marker is a direct signal. Inferring a crash from the state of the log is not reliable, because the log and the segment files reach the disk at different moments, and a kill can leave a log that looks clean while index entries still lag the records.

Two checks come first. A lock on the data directory, stamped with its owner's host and process id, refuses a second InventDB process; a lock whose owner is gone, or that names another host as a restored backup would, is taken over. Then a volume in a newer on-disk format than the running platform understands is refused rather than read, and an older one is converted.

The first start after a crash

Every start validates segments and replays the log; after a crash the engine also re-indexes the rows at risk and re-derives counts and zone maps before serving Every start take the lock · check the format version · validate active segments · replay the log Did the last process stop in order? yes no Serve requests nothing more to do AFTER A CRASH, ALSO 1 Re-index the rows the log still names, by id 2 Re-index each record file past its watermark 3 Widen zone maps, rebuild counts from the trees then serve requests
The work after a crash is limited to what a crash can leave behind: the rows the log names and the tail of each record file.
  1. Validate each active segment. Compare the id index's entry count with the record file's, and read ten sample records through the id index, checking each one's CRC. An id index that counts more entries than there are records, or points past the end of the record file, is rebuilt from the record file.
  2. Note the rows at risk. Read the log entries above the last checkpoint position and collect the namespace, type and id of every row they name. These are the rows whose index entries may not have reached the disk, because the checkpoint position advances only after the index pages are fsynced.
  3. Replay. Apply the entries in order, skipping any transaction without a commit marker. A row that already exists only gets its missing index entries, and the delete of a row that is already gone is skipped. Each entry is applied in isolation, so a damaged entry is logged and skipped instead of stopping the start.
  4. Re-index by id. Repair the rows from step 2 by id, with one batched lookup per segment instead of one probe per id per segment.
  5. Re-index the tail. Scan each segment's record file from its watermark, described below, to the end, and add any index entries that are missing.
  6. Re-derive. For types the log named, recompute counts, the reverse index and zone maps from the main trees. For every type, compare each zone map with the first and last key of its tree, and widen any map that is narrower.

Then the server starts accepting requests. Steps 4 to 6 run only after a crash, and after an orderly stop the log is normally empty, so a clean start does little more than validate.

Why the work is bounded

Nothing on the crash path may take time proportional to the size of the store. Every segment is visited, but a segment with nothing to repair answers at once, and two mechanisms bound the rest. The log names the rows at risk exactly. The watermark covers the record file: each time a checkpoint fsyncs a segment's index trees, it records the position in the record file up to which every record is known to be indexed, in a 12-byte file with its own CRC32 that is itself fsynced. Recovery scans only from that position to the end. In our crash suite those tails measure 81 to 600 records, or 120 KB to 895 KB.

Re-indexing a row that is already indexed costs one tree probe, because indexing stops early when the entry exists. A pass that covers too much is therefore harmless, and only one that covers too little could lose a row, so both mechanisms err towards covering more. If a watermark file is missing or fails its checksum, the engine adopts the end of the record file as the new watermark instead of rescanning the segment, because the log already names the rows at risk.

A full check of every field against every record reads every row, so it is deliberately not a startup step. It runs on demand through the repair endpoint, against an instance that is already serving.

Checksums, torn records and deletes

Every durable structure carries a CRC32: each 4 KB index page, each record as a 4-byte trailer over its header and payload, the superblocks and the segment manifest. Checksums are verified on every read, so bit rot or a torn write surfaces as an error instead of a wrong answer. Our design measurements put the cost at about 50 nanoseconds per 4 KB page and 15 nanoseconds per 500-byte record, under 1% of the page and record I/O. CRC32 detects accidents; tampering is caught by the authentication tag of AES-256-GCM, which encrypts every page and record.

A crash usually tears the end of a file, so the record format is built to survive a torn record. Each record header carries a two-byte marker. When a scan meets a header that makes no sense, it searches forward in 1 MiB windows for the next marker and accepts a candidate only if that record's CRC verifies. A chance match inside encrypted data costs one failed CRC, and the scan continues, so one torn record does not hide the intact records after it.

Deletes leave a trace as well. A delete appends a tombstone to the record file before the id index entry is removed, so rebuilding the id index from the record file keeps deleted rows deleted. Pruning structures fail safe: a zone map that has recorded nothing answers that a value might be present, so it never causes a segment to be skipped that it cannot vouch for.

What recovery does not do

  • It does not invent data. Recovery replays the log and restores agreement between indexes and records. Bytes that reached neither the log nor a record file belonged to writes that were never acknowledged, and index entries pointing at them are pruned.
  • It is not replication. Durability means fsynced to the instance's own volume. Losing the volume itself is what backups are for.
  • Damage inside a file is a different event. If the active segment's records fail their checksums and rebuilding the id index does not clear the fault, the engine moves that segment's directory aside under a timestamped name for inspection, starts the segment empty and replays the log into it. Rows already checkpointed into that segment are then not served until they are restored from a backup.
  • A torn index page is rebuilt rather than patched. Pages are encrypted whole, so a torn page cannot be partly read. The next insert into that type rebuilds that one field's index from the records; until then, a query that reaches the page returns an error rather than a partial answer.
  • Replay has a budget. Log replay stops after 120 seconds so that a pathological log cannot keep the instance from starting. Entries past that point stay in the log and the next checkpoint writes them, but their index entries need the repair endpoint below.

How we test it

The main crash suite runs a live multi-user import, kills the process at six different moments, and checks between every restart that filtered queries and counts agree with a full scan of the records. The current run passes all six cycles, with 2,186 rows verified. Each recovery mechanism also has a test that we run against a build with that mechanism removed, to prove the test can fail: on the same damaged record file, the torn-record scan returns 799 of 800 rows with resynchronisation and 266 without it.

We measure startup cost as well. On a 1.4 GB store with 78 segments, 76 types and 145,331 rows, an open that runs the recovery path takes 22.4 to 23.1 seconds, and the first query after it returns in 0.1 seconds or less. Repairing 200 rows named by the log takes 2.94 seconds, and repairing 2,000 takes 9.53 seconds.

Checking an instance yourself

For administrators, GET /api/recovery/summary reports the on-disk format version against the one the platform expects, how many id-index entries have been found pointing past the end of a record file, and the status of every type. GET /api/recovery/{namespace}/{type} gives one type's status, record count and segment count. If a filtered query on a type ever looks short, the repair endpoint rebuilds that type's id index from its records, verifies every field index against them, and reports what it changed:

POST /api/recovery/pms/leases/repair
Authorization: Bearer <admin token>

All of this works the same way on InventDB Serverless and InventDB SOAR. Because a sleeping instance is stopped in order, with a final checkpoint and the marker written, waking it does not pay for crash recovery; only a real crash does.