AI Allowance

AI Allowance gives your application text generation under the plan's included AI allowance. Your backend calls one function with a prompt, an optional JSON Schema, and a transport function your application supplies. The platform sends the call to the provider on a key the platform keeps, and counts the tokens in units against your plan's monthly quantity. Once the quantity is spent, the platform refuses the call with allowance_exhausted. The platform never gives your application the provider key for this route, and nothing on it is billed beyond the plan you already pay for.

Use it when

Use AI Allowance wherever your backend needs an answer from a model at runtime and you have stored no provider key. Use it also where you want one part of your application to use the plan's included units while another part uses your own key.

You choose the outbound path per call. A call through this package draws on the allowance. A call through your own declared upstream, made with the Egress package's client, uses your stored key and draws nothing from the allowance. One application can use both: for example, the allowance for short questions and your own key for a content generator that uses many more tokens.

The allowance covers text generation alone at this version. A request for image generation, paid grounding, cached content, or a separately priced provider tool is refused with allowance_feature_refused. Those features need your own key through Egress.

What it provides

The current package provides:

  • One function, generateText(transport, request, options), published as the npm package @turnzero/ai_allowance. Use the library shows how to copy it into app/lib/ai_allowance/, declare it as file:lib/ai_allowance, and install it with npm install.
  • The allowance figures on every response, in its allowance member, as the second table below lists. readStanding(headers) reads the same three values from any response of the route.
  • A typed refusal, AllowanceRefusal, thrown for every response the gateway itself wrote, rather than passed on from the provider.
  • A typed upstream error, AllowanceUpstreamError, thrown for a provider response with a non-success status. Its quota member is true where the provider rejected the call because it lacked capacity, the case your application may fall back on.
  • A streamed form: pass onText in the options, and each text fragment arrives in order as the provider produces it. The returned text contains the whole answer.
Request member Required
prompt Yes
system No
responseSchema No
maxOutputTokens No
model No
thinkingBudget No

The response contains the text, the model, the usage, the finish reason, and the allowance figures.

Allowance member What it contains
allowance.used The units your application drew this month before the call.
allowance.quota The plan's quantity.
allowance.resetsAt The first instant of the next UTC month.
Refusal Status Meaning
allowance_exhausted 429 The month's units are drawn.
allowance_unavailable 503 The platform's key is not yet in place.
allowance_application_required 403 The call used an account-wide credential, which never draws on the allowance.

Any other refusal name is one of the gateway's own, passed through with its status.

A test double ships at the package's testing subpath, @turnzero/ai_allowance/testing. createAiAllowanceDouble takes the place of the allowance route behind the unchanged generateText. It returns the replies a test queues and counts units against a quota the test sets. Test your application locally shows it in use.

From version 0.4.0, the double reads an upstream name containing a NUL character (U+0000) or an unpaired UTF-16 surrogate as no name, and refuses the call with 404 unknown_upstream, as the gateway does.

transport is a function your application supplies. It takes a gateway path and a request, adds the platform's origin and your application's credential as the bearer, and returns the response. The platform credential the deploy injects is bound to its environment and takes no environment header. A call under a minted token bound to the application gives the environment in the header x-turnzero-cloud-environment. The function adds content-type and no other header.

Every environment your application has, and its local runs, draw on the same quantity. The provider the platform runs the allowance on today is Google's Gemini Flash. A new version of the package can change the provider without changing any exported name. read_usage returns the month's figure as the measure ai_allowance_units, beside the plan's quantity. When the month's units first pass eighty percent of the quantity, the platform writes the event usage warning into your application's log stream.

Where your backend closes its connection before a call has answered, the platform ends the call at the provider, streamed or not.

A call already sent to the provider can end early in any of these ways, and your backend receives what each says:

  • Your backend gives the call up: it receives no answer. On a streamed call, the text that already arrived stays delivered.
  • The call runs past the platform's time limit: a call that does not stream fails with 504 upstream_timeout. A streamed call keeps the text already sent, then ends with a final egress_error event naming upstream_timeout. Either way, generateText throws AllowanceRefusal with that name.
  • The provider's answer passes the platform's size limit: a call that does not stream fails with 502 response_too_large. A streamed call keeps the text sent within the limit, then ends with a final egress_error event naming response_too_large. Either way, generateText throws AllowanceRefusal with that name.

Such a call draws the provider's token counts where they reached the platform, in a streamed answer's text or in a whole answer over the size limit. Otherwise it draws an estimate of the request's input tokens: one for every four bytes of the request. A call the provider had already answered with an error status draws nothing, as it would if it had finished. A call stopped before the platform sent it draws nothing, such as a request over the size limit. The allowance figures and read_usage count these units like any other call's.

Availability

Version 0.5.0 is the package's current version, and list_library reports the version the platform currently publishes. The entry includes the testing subpath from version 0.2.0. The entry contains the package's product statement, its detailed contract, its integration guide, and the compiled module. The platform's egress gateway runs the route on the platform's own key, counts use per application, and returns the allowance figures on every call.

Egress sends calls on your own stored keys through the same gateway. Secrets stores those keys. Logging records the usage warning event and what your backend does with an answer.

The boundary and the fallback

Choose your fallback by the refusal's name, never by its message:

  • On allowance_exhausted or allowance_unavailable, wait for resetsAt or call your own declared upstream on your own stored key.
  • On an AllowanceUpstreamError whose quota is true, apply the same fallback.
  • Treat every other upstream error as the provider's own. It reaches you with its body unchanged.

Never retry an exhausted allowance inside the month: every further call is refused, and each refusal is counted.

The last accepted call may take the month past the quantity by its own size, and never by more. read_plan_quotas returns the included quantity of each plan as its gemini-flash-allowance row. The pricing page of the public site states the paid plans' quantities, and a larger plan raises the quantity.