InventDB
All articles Search

Embeddings computed inside the database, in pure Rust

InventDB does not call an embedding service. A MiniLM sentence model, written out by hand in Rust, runs inside the engine process; two caches keep it off the query path as much as possible, and HNSW graphs make its vectors searchable. This is how each part works and where we drew the lines.

Semantic search needs vectors, and the usual way to get them is an embedding API. The application sends text over the network, receives a list of numbers, stores it in a vector database and keeps that store in step with the records. Every step adds latency, a charge per call and another copy of the data on someone else's machine.

InventDB computes embeddings inside the engine process. The model is MiniLM-L6-v2, a small sentence-transformer that maps a passage of text to 384 numbers, and we implemented its forward pass by hand in Rust. The text never leaves the instance, and there is no second store to synchronise.

This article covers the implementation, the two caches around it, how a column of records and a folder of files each become a searchable graph, and the trade-offs we chose. If you want to see how semantic search looks from SQL first, start with Four kinds of search, from exact match to meaning.

The model, written out by hand

MiniLM-L6-v2 is a BERT-style encoder: six transformer layers, twelve attention heads of 32 dimensions each, a hidden size of 384 and a feed-forward width of 1,536. Because those shapes never change, the engine needs neither a graph interpreter nor a general tensor library. Each layer is a fixed sequence of calls: three projections for queries, keys and values, scaled dot-product attention, an output projection, a residual add and layer normalisation, then a feed-forward block with GELU, a second residual add and a second normalisation.

The weights come from the published model file. We read that file once, offline, and write the tensors out as a compact file of 32-bit floats behind a short header. At run time the engine loads the compact file on first use and never parses the original model format, so the embedding path runs without an ONNX runtime or any general inference library. The only other component on this path is the WordPiece tokenizer.

The forward pass: text is tokenised, embedded, run through six encoder layers, pooled and normalised into 384 numbers, with the matrix multiply kernel chosen from the CPU's features Text Tokenizer up to 256 tokens Embedding lookup token, position, type SIX ENCODER LAYERS Self-attention 12 heads run in parallel Feed-forward 384 to 1,536 to 384, GELU Mean pool over the real tokens Normalise to unit length 384 numbers one vector per text Matrix multiply kernel, chosen from the CPU AVX-512, AVX2+FMA, AVX2, SSE4.1 or scalar
One embedding call. The shapes are fixed, so every step is a direct function call; only the matrix multiply varies, and it is chosen from the processor's features.

A call runs in this order:

  1. Tokenise the text and drop the padding, so the sequence is as long as the text and no longer, up to 256 tokens.
  2. For each position, add the token, position and segment-type embeddings and normalise the sum.
  3. Run the six encoder layers.
  4. Average the output over the real tokens.
  5. Scale the average to unit length, giving 384 numbers ready for cosine comparison.

Matrix multiplies chosen for the processor

Almost all of the arithmetic sits in the projections, which multiply the sequence's 384-wide rows by 384×384 and 384×1,536 weight matrices. The engine checks the processor's features and uses the widest kernel it supports: AVX-512, AVX2 with fused multiply-add, AVX2, SSE4.1 or a portable scalar loop. The kernels are blocked for the cache. The AVX-512 kernel, for example, works through tiles of 6 rows by 16 columns, 256 steps deep, and handles sixteen floats per instruction. The twelve attention heads are independent of each other, so they run in parallel across cores.

Two smaller decisions matter as much as the kernels. First, the forward pass covers only the real tokens: a twelve-token product name runs as a twelve-position sequence rather than a padded block of 256. Attention costs grow with the square of the sequence length, so short inputs save the most. Second, the model keeps scratch buffers sized for 64, 128 and 256 tokens, picks one by sequence length and reuses it, so the layer activations need no memory allocation per call.

A column becomes a graph of its distinct values

Embedding every row of a large table is the expensive way to do semantic search, and much of the work is repeated, because many rows share values. For record columns, InventDB embeds distinct values instead.

An administrator asks for a column to be built through the API. The engine reads the column's distinct values from the property index, which keeps a count for every value and can list them without reading the records. It compares that list with the values already in the column's cache, embeds only the missing ones, in batches of 64, and rebuilds an HNSW graph over all the vectors. The result is written to a temporary file and renamed into place, so queries keep using the previous version until the new one is complete. A second build with no new values embeds nothing.

Values longer than 1,024 bytes are stored in the property index as hashes, so for those the engine reads the real text from the records before embedding it. The column file is encrypted, records which model produced it, survives restarts and is loaded into memory by the first query that needs it.

Building a column file from distinct values, and using it to answer a MEANING query through the property index BUILD, ON REQUEST Property index distinct values Embed new values batches of 64 HNSW graph over all vectors Column file encrypted, renamed in QUERY Search text MEANING() LIKE Query cache model on a miss Nearest values inside threshold Equality lookups ids, then rows
The build embeds each distinct value once and publishes the graph with an atomic rename. A query searches the graph for nearby values, then uses the ordinary property index to find the records that hold them.

At query time, MEANING(notes) LIKE 'leaky roof' embeds the search text, walks the graph to the nearest stored values, keeps the ones inside the similarity threshold, and turns each value into record ids with an equality lookup on the property index. The rows are then fetched through the engine like any other result, so row rules apply to them.

Two caches keep the model off the query path

Embedding a short string costs roughly 15 milliseconds of CPU time, and longer text costs more. That is long for a query and far too long to repeat. Two caches remove most of that cost.

The query cache maps search text to its vector. It is a least-recently-used cache with constant-time lookups and inserts, built from a hash map and a doubly linked list over a slab of reusable nodes. Keys are trimmed and lowercased, so "Studio Apartment" and " studio apartment " share one entry. Its capacity is set at start-up to an eighth of the memory then available, at about 1,600 bytes per entry, and never fewer than 1,000 or more than 1,000,000 entries. It holds vectors only, never rows, so it has nothing to invalidate when records change.

The distinct-value cache is the column file described above. Because it is keyed by value, a value that appears in a million rows is embedded once, and each rebuild costs in proportion to the values added since the last one rather than to the size of the table.

Files: a vector for every chunk

Files take a different route, because their text is long and a single vector cannot represent a forty-page contract. After a file is uploaded, a background worker extracts its text. Text longer than about 1,024 tokens, estimated at four characters per token, is split at sentence boundaries into overlapping chunks of about 512 tokens with 50 tokens of overlap; shorter text stays as one chunk. The worker embeds every chunk and adds the vectors to an HNSW index kept for the namespace. Uploads, new versions and deletes update that index as they happen.

A search over files returns the best-matching chunks ranked by similarity, together with each chunk's text, which is how file search can show the passage that matched. Extraction, OCR and the rest of the file pipeline are covered in Files on every record, read by OCR and searchable by meaning.

Choices we made

No embedding on every write. A write path that waited that long for each text field would make a much slower database, and a write path that queued the work would make search results depend on the length of a queue. Record columns are embedded when an administrator builds them, and each later build adds only the new values. Files are the exception, because their processing already happens in a background worker after the upload returns.

No text leaves the instance. The model runs in the instance's own process, on its own machine. For leases, clinical notes and contracts, the question of where the text is sent for embedding never arises.

One model at a time. MiniLM-L6-v2 is the default. An instance can be configured to use TinyEmbed instead, a smaller opt-in model with two layers that produces 256-number vectors. It is off unless selected. Vectors from different models cannot be compared with each other, so indexes built with one model have to be rebuilt after switching to the other.

Limits

  • The model reads at most 256 tokens of any input. Text beyond that does not affect the vector.
  • Values written after a column's last build are not found by meaning until the next build. A column that has never been built falls back to substring matching for MEANING() queries.
  • HNSW returns approximate nearest neighbours. It trades a small chance of missing a close match for search time that grows slowly with the number of vectors.
  • The vectorised kernels are for x86 processors. On ARM processors the matrix multiply uses the portable scalar kernel, which is correct but slower.

Using it

An administrator builds the columns that hold free text. One call queues several columns; each build reports progress, and a build that is already running is not started twice:

POST /api/vectors/pms/workorders/regenerate-batch
Authorization: Bearer <admin token>
Content-Type: application/json

{ "properties": ["notes", "category"] }

After that, anyone with read access can query those columns by meaning with MEANING() in SQL, under their own row rules. If you want the vectors themselves, for clustering or for a model of your own, the instance returns them from the same runtime:

POST /v1/embeddings
Authorization: Bearer <token>
Content-Type: application/json

{ "input": ["leaky roof", "water coming through the ceiling"] }

The response holds one 384-number vector per input string, in the request and response shape that most embedding clients already understand. All of this runs on both InventDB Serverless and InventDB SOAR, because the model is part of the engine rather than a service either product calls.