A business keeps as much of its information in files as in rows: a signed lease, an inspection report, a photo of a receipt. In most systems the file goes to an object store, its text is extracted by a separate service if it is extracted at all, and searching inside it needs another index that knows nothing about who may see the record the file belongs to.
In InventDB a file is attached to a record and stored by the same engine that stores the record. A background worker reads its text on the instance itself, the engine indexes that text for full-text and semantic search, and every search result is checked against the caller's access to the parent record. This article follows one file from upload to search result, and then covers versions, limits and the requests you would use.
Upload: the request returns before the reading starts
A file is uploaded to a record with a multipart request to /attach/{namespace}/{type}/{record id}, with an optional description, tags and folder path. A file uploaded without a record id goes to the type's shared vault instead. An upload that names a record id which does not exist is refused with a 404, so a file cannot be left attached to nothing.
The upload has two phases. In the first, which runs inside the request, the engine splits the bytes into chunks of 512 KB and addresses each chunk by its SHA-256 hash. A chunk that already exists in the namespace gains a reference instead of being stored a second time. The chunks and the file's metadata are written as records in the engine's own system types, so they are encrypted at rest like every other record and there is no separate file store to secure. The file is marked Pending and the request returns.
The second phase belongs to a background worker that processes one file at a time, which keeps memory use low and predictable while OCR runs. It moves each file from Pending to Processing and then to one of three end states: Ready when text was extracted and indexed, Skipped when the file has no text to index, and Failed when processing went wrong three times. The state is stored in the database, so a restart resumes where the worker stopped. A Skipped or Failed file can still be downloaded and found by its name, and a re-extract request runs the pipeline again.
Reading the text, by format
The worker picks a reader from the file's content type and extension. Every reader runs inside the engine process:
| Files | How the text is read |
|---|---|
| The text layer. When that yields almost nothing, as with a scanned PDF, the images on its pages go through OCR. | |
| DOCX | The text of the document body. |
| XLSX, XLS, XLSB, ODS | Every sheet, row by row, with cells separated by tabs. |
| EML email | The From, To, Subject and Date headers, then the body, with HTML tags stripped. |
| HTML, XML, Markdown | The text content; for HTML, scripts and styles are left out. |
| TXT, CSV, JSON, YAML, logs, source code | Read as text. |
| ZIP | The text-like files among the first 100 entries. |
| JPEG, PNG, GIF, TIFF, BMP, WebP | OCR. |
Anything else, such as video and audio, is stored, versioned and downloadable, and is marked Skipped because there is no text to index. Documents larger than 10 MB skip text extraction, so one very large scanned book cannot hold up every file queued behind it; images are not subject to that cap. Extracted text is kept up to 1 MB per file.
OCR on the instance's own processor
Images are read with the PP-OCRv4 models: a DBNet detector finds the regions that contain text, an SVTR recogniser reads each line, and two small models detect the orientation of text lines and of the whole page. The models run on the instance's CPU inside the engine process. The worker never calls a language model or a vision model, so reading a file takes a predictable amount of time and the file never leaves the instance.
Two size rules keep OCR both legible and bounded. An image whose longest side is more than 2,048 pixels is scaled down to 2,048, keeping its aspect ratio; anything smaller is read at its native resolution, which matters for small text in screenshots. Each detected line is then cropped and passed to the recogniser at up to 2,048 pixels wide with its aspect ratio kept, so a long line of a paragraph is not squeezed until its letters run together.
The text is indexed twice
The extracted text is stored as a record of its own, linked to the file, and you can fetch it with a GET request. That record feeds two indexes.
The full-text index. When the next checkpoint persists the text record, the engine adds it to the BM25 inverted index of the newest segment, splitting it into words, lowercasing and stemming them. Ranking, stemming and the per-segment layout are described in Four kinds of search, from exact match to meaning.
The vector index. Text longer than about 1,024 tokens, estimated at four characters per token, is split at sentence boundaries into chunks of about 512 tokens that overlap by 50; shorter text stays as one chunk. The worker embeds every chunk into 384 numbers and adds the vectors to an HNSW index kept for the namespace, which uploads, new versions and deletes update as they happen. How those vectors are computed is covered in Embeddings computed inside the database. Only files in the Ready state are content-searchable.
Searching inside files
File search has four modes. Keyword matches file names and descriptions. Full-text ranks files by BM25 over their extracted text; when the index has no hit for a scope, a slower scan ranks files by how many distinct query words they contain. Semantic returns the files whose chunks are nearest in meaning, together with the passage that matched. Combined, the default, runs all three and adds their scores with weights of 0.2, 0.3 and 0.5, which a request can change.
Every semantic hit must reach a cosine similarity of 0.62, in every mode. We set that floor just above what we observed for images whose only text was a short fragment, such as a number OCR'd from a stock photo: a vector built from a fragment like that had a cosine similarity between 0.47 and 0.59 to every unrelated query, so without a floor the image matched everything. The floor is a relevance bar rather than a blocklist. Searching for the fragment itself still returns the photo at a similarity close to 1.0.
POST /attach/_search
Authorization: Bearer <token>
Content-Type: application/json
{
"query": "roof leak repair estimate",
"search_type": "combined",
"types": ["properties", "workorders"],
"mime_types": ["application/pdf", "image/"],
"limit": 25
}
This global form searches every namespace and type the caller may read, with optional filters for namespaces, types, a folder prefix, MIME types or MIME families such as image/, and tags. It returns 25 results by default and at most 100 per page. Scopes the caller cannot read are absent from results, folders and totals. Files whose parent record is hidden by the caller's row rules are dropped as well, because a file is visible only to people who can see the record it belongs to. A per-type form at /attach/{namespace}/{type}/search takes the same modes.
Every version kept
Uploading a new version of a file adds to its history instead of replacing it. Any version can be listed, fetched or downloaded. Restoring version 3 writes its content as a new version with the comment "Restored from version 3", so history only ever grows. A single version can be deleted, but not the only one; deleting the current version rebuilds the file's search data from the version that becomes current.
Because chunks are addressed by hash and reference-counted, the same file attached to two records in one namespace, or uploaded again unchanged as a new version, adds no new chunk data. A chunk's bytes are removed only when its last reference goes.
Deleting many files at once is batched. A bulk delete of a folder subtree or of a whole type's files flushes the write-ahead log once per system table and saves the vector index once per batch, rather than once per file. In our measurement, deleting 60 files took 1.04 seconds batched, against 18.76 seconds one file at a time. The request can cap how many files a call deletes and returns how many remain, so a client can show progress and stop when nothing is left.
Limits
- The OCR models read English. Text in other scripts is not recognised by the default models.
- Documents over 10 MB are stored but not read, and extracted text beyond 1 MB per file is not indexed.
- A file becomes content-searchable once the worker has processed it, and full-text search sees it after the next checkpoint, which runs every five seconds. Large batches of uploads are read one file at a time.
- Inside a ZIP archive, only text-like files are read; images and office documents inside an archive are not.
- Uploads are limited by the instance's request size limit, 100 MB by default.
Using it
Attach a file to a record:
curl -X POST https://<your-instance>/attach/pms/properties/<record-id> \
-H "Authorization: Bearer $TOKEN" \
-F "file=@inspection-2026-09.pdf" \
-F "description=Annual fire safety inspection" \
-F 'tags=["inspection","2026"]'
Then check the file's state, or run a search: once the state reads Ready, its contents answer full-text and semantic queries. Attachments, OCR and file search run on both InventDB Serverless and InventDB SOAR, because the readers and the embedding model are part of the engine. In InventDB SOAR, the AI agent can also take the PDF attached to an invoice or receipt email in a connected Gmail account and file it on a record, as a step in a change set you approve.