A sale touches several records at once: the order, its lines and the stock count of each item. If the process dies halfway through, or two clerks sell the last unit at the same moment, the application needs a clear answer about what was saved.
InventDB's transactions give three answers. A commit is applied whole or not at all. A commit that would overwrite a change it never saw fails instead of succeeding. And a commit that has been acknowledged is on disk, so it survives the process dying a moment later. This article explains how each of those is achieved, then states the anomalies the design does not prevent and how to guard against them.
A transaction over HTTP
A transaction is opened with POST /tx, which returns its id. Each operation is a POST to /tx/<txId>/execute that names the operation: get, insert, update or delete. The transaction ends with /commit or /rollback. Here is the clerk's sale of the last unit of an item:
POST /tx
201 { "ok": true, "txId": "<txId>" }
POST /tx/<txId>/execute
{ "operation": "get", "namespace": "shop", "type": "stock", "id": "sku-118" }
POST /tx/<txId>/execute
{ "operation": "update", "namespace": "shop", "type": "stock", "id": "sku-118",
"document": { "sku": "SKU-118", "on_hand": 0 } }
POST /tx/<txId>/execute
{ "operation": "insert", "namespace": "shop", "type": "orders",
"document": { "customer_id": "c-77", "sku": "SKU-118", "qty": 1 } }
POST /tx/<txId>/commit
200 { "ok": true }
Grants and foreign keys are checked when each operation is executed, so a write the caller may not make returns 403, and one that breaks a foreign key returns 409, before anything is staged. Row rules are checked against each record as it is staged. An update replaces the record's fields with the document sent. A transaction has five minutes to commit; a later commit fails and the transaction is rolled back.
Until commit, writes stay private
Nothing a transaction writes touches shared storage before it commits. Each write goes into a buffer that belongs to the transaction, keyed by namespace, type and id. A read inside the transaction checks that buffer first, so the transaction always sees its own pending writes. The buffer also collapses changes as they are made: an insert followed by an update stays one insert carrying the new data, and an insert followed by a delete leaves nothing behind.
Readers outside the transaction never see the buffer, so no one reads a write that might still be rolled back. A rollback discards the buffer, and since nothing reached disk there is nothing to undo.
Conflicts are found at commit
InventDB does not lock records while a transaction is open. A transaction that waits on a person or a slow outside call therefore never blocks anyone else. Conflicts are detected when the transaction tries to commit, which is why this style is called optimistic concurrency control.
For that check the engine keeps, in memory, a version number for each record that transactions write, and raises it at the moment a committed write becomes visible. When a transaction first touches a record with a get, an update or a delete, it notes the version it saw. At commit:
- The transaction takes the next commit sequence number from an atomic counter.
- Under a short lock that does no I/O, it places a claim on every record it is about to write, then validates. For each record it both read and wrote, it asks two questions: has the version moved since the read, and does another transaction hold a live claim on the record?
- If either answer is yes, the transaction aborts. Its claims are removed, its buffer is discarded, and the commit returns 400 with the conflict named in the error. Otherwise it goes on to the log.
- Once its writes are visible, it raises the versions and drops its claims in a single step under the same lock, so another transaction validating at that instant always sees either the claim or the new version.
The effect is that the first committer wins. Two clerks each read sku-118 with one unit on hand, and each writes zero. The first commit succeeds. The second finds the version has moved, or that the first still holds its claim, and fails. Its application re-reads the record, sees zero on hand and tells the customer. Neither sale silently overwrites the other, which is the lost update this check exists to prevent.
One batch, one marker, one fsync
A validated transaction goes to the write-ahead log as a single batch: one entry per insert, update or delete, each tagged with the transaction id, followed by a commit marker. The whole batch is appended under one hold of the log buffer's lock, so the marker is always contiguous with the entries it covers. Recovery and the checkpoint apply a transaction's entries only when they find its marker, so a transaction is on disk entirely or not at all.
Durability comes from group commit. The committing thread asks the log to make everything up to its own last entry durable. One committer becomes the leader: it takes everything in the buffer, its own entries and those of every other transaction in flight, writes it and calls fsync once. The others wait until the durable position passes their entries. No commit returns before its own entries are on disk, and commits made at the same moment share one fsync instead of paying one each.
Two details keep that fsync cheap. The log file is preallocated to 64 MB, so a commit's fsync flushes data without also updating the file's size. And a transaction that only read writes nothing to the log: no marker and no fsync, because there is nothing to recover. Entry payloads are encrypted with AES-256-GCM before they are written.
Visible as soon as it is acknowledged
After the fsync, the engine applies the transaction to its in-memory structures. Inserts are grouped by type and written into each property's index in one sorted pass, and an update touches only the indexes whose value changed. Each record then goes into the MemTable, the in-memory table that every read consults before the segment files. From that point the commit is visible to every reader, and the call returns.
There are no replicas to catch up and no background apply step, so a read issued after a successful commit returns the committed data. Writes made outside a transaction follow the same rule: the call returns only after its log entry is fsynced and the MemTable holds the record.
Every five seconds, a checkpoint
The log and the MemTable are a short-term home. Every five seconds a background checkpoint moves their contents into the segment files:
- Flush anything still buffered in the log, then read every entry since the last checkpoint.
- Skip any transaction entry whose commit marker is not in that window; it stays in the log for the next checkpoint.
- For each type with changes, append the records to its record file and update its id index, with types handled in parallel.
- Write and fsync the index pages left dirty since the last checkpoint, for those types only.
- Advance the checkpoint position in the log, drain the MemTable up to it and truncate the log.
The order matters. The checkpoint position moves only after the index pages are on disk, so the entries still in the log name exactly the rows whose indexes might need repair after a crash.
If the process dies
On the next start the engine reads the log from the last checkpoint position and replays it. An entry tagged with a transaction is applied only if that transaction's commit marker is present, so a transaction that died before its fsync leaves no trace and one that reached disk is restored in full. Replay is idempotent: a row already on disk is not inserted twice, and the delete of a row that is already gone is skipped. Our crash recovery article follows that start step by step.
What it guarantees, and what it does not
| Situation | Outcome |
|---|---|
| Another transaction's uncommitted write | Never visible. |
| Two transactions read and update the same record | The first to commit wins; the second fails at commit. |
| The process dies during a commit | Applied whole if its batch reached disk, otherwise not at all. An acknowledged commit always reached disk. |
| A read after a successful commit | Returns the committed data. |
| Two transactions read the same records and write different ones | Both can commit. This is write skew. |
| A record read twice that the transaction never writes | Can change between the two reads. |
| A plain write, outside any transaction, changes a record a transaction has read | Not detected at commit. Conflict checks compare transactions with each other. |
The last three rows are the boundary of the design. InventDB keeps no old versions of records to serve each transaction a frozen snapshot: a read returns the newest committed value, and the commit check covers the records a transaction both read and wrote. The isolation is therefore weaker than serializable.
Consider a clinic rule that at least one clinician must be on call. Two transactions each read both on-call records, each sees two clinicians on call, and each takes a different one off call. Each wrote a different record, so neither conflicts, and both commit. The standard remedy works here: make the decision write a shared record. If taking someone off call also updates one roster record holding the on-call count, both transactions read and write that record, and the second commit fails.
What it costs
The main cost of a commit is one fsync, shared with every transaction committing at the same moment; a lone writer pays a full fsync for each commit. Validation takes time in proportion to the records the transaction wrote rather than the size of the database, and the lock around it covers no I/O. A failed validation costs the client a retry. Durability is local: a commit is fsynced to the instance's own volume, with no synchronous replica, and backups protect against losing the volume itself.
Using transactions well
- Read inside the transaction. The version is noted at the first touch, so a value read before the transaction began is not protected.
- Retry the whole transaction when a commit fails, reads included; the second attempt sees the other writer's change.
- Keep transactions short. They hold no locks, but a long one is more likely to meet a conflicting commit, and the limit is five minutes.
- Protect rules that span records with a shared record that every transaction relying on the rule also updates.
- Write contended records through transactions only. A plain write outside a transaction does not raise a record's version, so it is invisible to the commit check.
Transactions work the same way on InventDB Serverless and InventDB SOAR. In InventDB SOAR, an AI change set that a person approves is applied through this same commit path, so it lands whole or not at all.