An AI feature inside a database raises three practical questions: who pays for each call, where the data goes, and what stops one runaway workflow from spending a month's budget in an afternoon.
InventDB SOAR, the database with the business engine, answers all three in one place. A managed instance holds no model-provider credential. It holds its account's gateway key, and sends every model call, from the Analyze agent, report authoring and workflow steps alike, to the InventDB AI gateway. The gateway makes four decisions before any tokens are spent and one after the model replies.
This article describes those decisions as a customer sees them: authentication, the monthly allowance, routing by region, metering, and how web search works on Bedrock. InventDB Serverless, the managed database, has no AI inside and never calls the gateway.
Four checks before the model, one after
- The key. Each customer account has its own gateway key, installed on the instance when it is provisioned. The instance sends it with every call, together with the user and the product feature that made the call. The gateway looks the key up on every request, so a key that has been replaced or revoked is refused on its next use.
- The subscription. The account's latest subscription must be active. A pending payment, a past-due invoice or a cancellation each get their own error code, so the console can show the right next step.
- The allowance. The gateway finds the account's usage record for the current month, creating it with the plan's allowance on the first call of the month. If the amount spent has reached the amount allocated, the call is refused with a 402 before it reaches the model.
- The model and the region. A model that is switched off is refused with a 403. Otherwise the gateway picks the Amazon Bedrock inference profile for the account's region and the requested model, and sends the call.
- After the reply. The gateway reads the token counts from the response, prices them, and adds the cost to the month's spent amount.
A monthly allowance per plan
Every InventDB SOAR plan includes a fixed amount of AI each month, stated in US dollars of model usage. The free tier includes $10; on a paid plan the allowance is part of the plan we size with you.
The month is the calendar month in UTC. The usage record for a month is created on its first AI call, so a new month starts from zero without any scheduled job. When the allowance is used up, AI calls stop until the next month or until the allowance is raised; the subscription price never changes because of AI use. Within the month, an administrator can also set daily limits per instance and per person on the instance itself, which the agent article describes.
What the console is told
Refusals carry a machine-readable code, so the console can say "complete your payment" rather than "error". The codes a customer can meet are:
| Situation | HTTP | Code |
|---|---|---|
| Key not recognised | 401 | invalid_or_revoked_key |
| Subscription awaiting payment | 402 | subscription_pending_payment |
| Subscription past due | 402 | subscription_past_due |
| Subscription cancelled | 403 | subscription_cancelled |
| No active subscription | 403 | no_active_subscription |
| Monthly allowance used up | 402 | quota_exhausted |
| Model switched off | 403 | model_disabled |
The allowance refusal includes the figures, so the console can show exactly where the account stands (abridged):
HTTP/1.1 402 Payment Required
{ "error": {
"code": "quota_exhausted",
"tier": "business_small",
"dollars_used": 50.02,
"dollars_allocated": 50.0,
"period_month": "2026-10" } }
Routing to Claude by region
An instance runs either on Claude or on an open-weight model served by InventDB Sparkle, our own inference engine, as its administrator chooses. Sparkle runs on our GPUs in the US and the EU, and the gateway applies the same key, allowance and metering checks to it. The rest of this section is about Claude.
The Claude models are Claude Sonnet 5 and Claude Opus 4.8, on Amazon Bedrock. Bedrock serves these models through inference profiles rather than through a single region, so the gateway chooses the profile from the country the account signed up in and the model requested.
- Accounts in the United States use the US profile, which Bedrock serves from US regions.
- Accounts in the EU and EEA, the UK and Switzerland use the EU profile, served from EU regions.
- Accounts in India and every other country use the global profile, which Bedrock can serve from any region. These Claude models have no profile pinned to Asia-Pacific, and we chose the current models on the global profile over older models with a regional one.
The gateway also absorbs the differences between Bedrock and Anthropic's own API, so the instance can send standard Anthropic requests. For these models Bedrock rejects the temperature, top_k and top_p parameters, so the gateway removes them. It also rewrites an extended-thinking request with a token budget into Bedrock's adaptive form with an effort level: a budget of 32,000 tokens or more becomes high, 8,000 or more becomes medium, and anything smaller becomes low.
Metering every token class
A Claude call has five kinds of tokens: input that was not cached, output, cache reads, and cache writes kept for five minutes or for an hour. Each kind is priced differently. The gateway reads every count from the response, including streamed responses, and prices it from a model catalog it refreshes every hour. Cache reads are the cheapest of the five, which is why the agent keeps the unchanging part of its prompts in the cache. The gateway writes the debit after a successful reply, so a call that fails upstream costs nothing.
For each call the gateway's usage ledger records the model, the token counts by class, the cost, the latency, the status, and which user and which feature made the call. It does not record prompt or response content; our privacy policy covers how prompts are handled. The instance keeps its own per-call usage records for the Engine page's cost breakdown, and the console shows the month's spent and allocated amounts, refreshed hourly.
Web search on Bedrock
On Anthropic's own API, Claude can call a hosted web search tool. Claude on Bedrock has no such tool, so the gateway provides it. When a person has web search switched on and a request carries the hosted search tool, the gateway replaces it with an ordinary tool of the same name and runs the loop itself: the model asks for a search, and the gateway runs it against a web index that Amazon operates and returns the results, for up to five rounds. Each query is capped at 200 characters and asks for between 1 and 25 results. Source URLs are kept through to the final answer so it can cite them, and the token usage of every round is added up and metered as one call. A failed search returns an "unavailable" result and the loop continues. Other hosted tools that Bedrock rejects are removed from the request.
Limits
- The allowance is checked before each call. The call that crosses the line completes and is charged; the next one is refused.
- The debit is written after the reply and is not part of the call. If writing it fails, the failure is logged and the call still succeeds.
- Accounts outside the US and Europe use the global profile, so their model calls are not pinned to one geography.
- The web search service runs in one US region, so search queries are processed there whatever the account's region.
- InventDB SOAR does not accept a customer's own model-provider key; all bundled AI goes through the gateway.
Using it
There is nothing to configure. Every InventDB SOAR instance is provisioned with its account's key, and the plan you choose on the pricing page sets the monthly allowance. To see where the month stands, open the Engine page in the console: it shows spent against allocated, and a cost breakdown by model built from the instance's own usage records. To keep one person or one busy workflow from using the whole allowance, set the instance's daily limits as an administrator.