InventDB
Research

InventDB Sparkle: our own inference engine for open models

InventDB Sparkle is the engine we wrote to run open-weight language models on our own GPUs. It came out of our research, it serves Qwen3.8-27B in the InventDB platform today, and on the same hardware it writes answers faster than vLLM and SGLang in almost every case we measured. Below we explain what it does, where it is used and how it measures, including where it is behind.

What makes InventDB Sparkle fast, and how it measured against vLLM and SGLang, in two and a quarter minutes, with music. The numbers in the film are Sparkle version 1.56, measured on 29 September 2026, losses included; the numbers further down this page are the build that runs in production, measured on 2 October.
In this article
629tokens a second for one conversation writing JSON
2,676tokens a second across eight conversations at once
262Ktokens of context in one conversation
18 sfrom start to a first answer

An inference engine is the program that runs a trained language model: it takes a prompt, carries out the model's arithmetic on a GPU, and streams the answer back a token at a time. Most engines are built to run hundreds of models on many kinds of hardware.

We asked a narrower question: how quickly can one model answer when its engine is written for that model and for one GPU, and nothing else? InventDB Sparkle is our answer. It is a single program written in Rust, with GPU code written for the NVIDIA H200, and it runs one model, Qwen3.8-27B, an open-weight model with 27 billion parameters.

Where it runs today

Sparkle is in production in the InventDB platform as an inference engine that an InventDB SOAR instance can select. An administrator chooses the instance's AI model: Claude, or an open-weight model served by Sparkle, which today is Qwen3.8-27B. Analysis, reports and workflows then run on that choice, under the same roles, row rules and monthly AI allowance as any other model.

Sparkle runs on GPUs in two regions, the US and the EU. Every request from an instance passes through the InventDB AI gateway, which checks the instance's key and allowance and sends the request to the model the administrator chose. Our article on the AI gateway describes those checks.

An InventDB SOAR instance sends its AI requests through the InventDB AI gateway either to Claude on Amazon Bedrock or to InventDB Sparkle, which runs Qwen3.8-27B on one NVIDIA H200 InventDB SOAR analysis, reports and workflows AI gateway key and allowance the model the administrator chose Claude on Amazon Bedrock INVENTDB SPARKLE Qwen3.8-27B, open weights Guess ahead, check in one pass One NVIDIA H200 GPU
An InventDB SOAR instance runs its AI on the model its administrator chose. With an open-weight model, Sparkle serves the request and streams the answer back.

What it does

  • Long context. One conversation can hold up to 262,144 tokens.
  • Images and video. A prompt can include images and video frames as well as text; the answer is text.
  • Answers in an exact shape. When a request supplies a JSON schema or requires a tool call, Sparkle lets the model write only tokens that keep the output valid, so the result parses every time.
  • Several conversations at once. One GPU serves several conversations together, and shorter jobs go first, so a short question does not wait behind a long document.
  • Fast follow-ups. After Sparkle reads a long prompt, its working state stays in GPU memory for a short while, so a follow-up question on the same document starts without reading it again.
  • Streaming. Tokens are sent as they are written.

How it answers quickly

Writing each token of an answer means reading the model and the conversation so far out of GPU memory. On a modern GPU the speed of that memory, rather than the speed of the arithmetic, usually sets the pace. Two ideas account for most of Sparkle's speed, and both are about getting more answer out of each read.

Compact numbers. Sparkle keeps the model's weights and the conversation's working state in a compact 8-bit form. Each token then costs fewer bytes to read, and a full 262,144-token conversation fits on the same GPU as the model.

Guess, then check. A small, fast helper model proposes the next several tokens, and the 27-billion-parameter model checks all of them in a single pass, keeping every token it agrees with. Each token in the answer is still the large model's choice; the helper only saves time. On structured answers such as JSON, Sparkle keeps about 8 to 11 tokens per pass, and on ordinary prose about 5.

The rest comes from writing the engine for one model and one chip. Because Sparkle runs nothing else, its GPU code can be shaped around Qwen3.8-27B's layers and the H200's memory and tensor cores, with no layer of general-purpose code in between.

Measured against vLLM and SGLang

We measure Sparkle against vLLM and SGLang, two widely used open-source engines, serving the same model. The conditions:

  • Each engine ran on an NVIDIA H200 machine of the same type, with the benchmark client on the same machine, so no network sat between them.
  • All three received the same prompts at temperature 0 with reasoning off, and each used its own guess-and-check method: vLLM 0.29.0 and SGLang 0.5.19.
  • vLLM and SGLang were measured on 16 September 2026, and Sparkle on 2 October 2026 with the build and settings that run in production, on a second machine of the same type. The same Sparkle build ran up to 3% faster on the second machine than on the first, so differences of a few percent are within what the machine alone can explain.

The table shows a request shaped like an InventDB SOAR analysis: a 33,000-token prompt and a 1,000-token answer. Each time is computed from the measured rows, as the median time to the first token plus the remaining 999 tokens at the measured writing speed. "Seen before" means the same prompt was read a moment earlier, as it is in a follow-up question.

Request, 33,000 tokens in, 1,000 outSparklevLLMSGLang
JSON, one conversation, new prompt4.06 s7.84 s6.10 s
JSON, one conversation, seen before1.65 s4.95 s4.07 s
JSON, eight at once, new prompts15.89 s28.62 s15.85 s
JSON, eight at once, seen before3.87 s8.71 s5.73 s
Prose, one conversation, new prompt5.58 s9.00 s6.33 s
Prose, one conversation, seen before3.22 s6.50 s4.35 s
Prose, eight at once, new prompts15.79 s30.07 s16.66 s
Prose, eight at once, seen before3.99 s13.89 s6.85 s

With eight new prompts at once, Sparkle and SGLang are within a few percent of each other, which we read as level. Across every case we measured:

  • Writing speed, one conversation: Sparkle was faster than both engines in all 12 cases.
  • Writing speed with eight conversations: faster in 11 of 12.
  • A batch of eight requests, from sending to the last token: faster in all 12.
  • First token for a prompt seen before: faster in 20 of 24, level in 2 and behind in 2.
  • First token for a long prompt never seen before: behind both engines in all 10 cases. Reading a large new prompt for the first time is where Sparkle trails, and it is where our work on the engine is going next.
  • Start-up: 17.7 seconds from starting the engine to a first answer, against 73.7 for vLLM and 150.4 for SGLang. In vLLM's test its model was still in memory from an earlier run, which makes its time look shorter than a true cold start.

What happens to the text you send

Sparkle runs with zero retention of request content. Prompts and completions are never written to disk or to logs, they are not used for training, and no person reads them. We keep only request metadata, such as token counts and timing, to meter usage against the instance's AI allowance and to run the service. Our privacy policy covers how InventDB SOAR handles AI requests.

Using it

In InventDB SOAR, an administrator selects an open-weight model served by Sparkle as the instance's AI model. Nothing else changes for the people using the instance: they ask questions, build reports and run workflows as before, and the answers come from Sparkle.

Sparkle is one of three research programmes at InventDB. InventDB Cluster runs one database across many machines, and InventDB Brahma works toward an intelligence that learns how its world works from its own experience. All three are listed under Research.