Call an AI model under the allowance

Prompt:

Have the backend write a one-line summary of each new note with AI. I don't have a key for any AI service.

Also works:

  • "Use the AI that comes with my plan."
  • "Suggest a title for each upload."

What your tool does

  • Reads the platform's use-ai-allowance skill, which sets the order of the steps below.
  • Finds the ai_allowance library entry with list_library, reads it with read_library_entry, and takes it (Use the library). It copies the entry into app/lib/ai_allowance/, declares it in package.json as file:lib/ai_allowance, and runs npm install.
  • Adds the package's version to the manifest's packages member at the next submit_manifest.
  • Calls read_plan_quotas before relying on the allowance, for the plan's monthly units.
  • Writes a transport function that adds the platform's origin and the application's credential, and calls generateText with it and a prompt. Where the answer must be structured, it passes a JSON Schema too.
  • Reads the allowance figures on every answer: used, quota, and resetsAt.
  • Writes the fallback for the end of the month's units: your own declared upstream, or an answer of your application's own.
  • Tests the calls over the package's test double before the first deploy.
  • Declares no egress entry and stores no provider key for the allowance.
  • After the deploy, calls read_usage for the month's draw.
  • Asks you only where the fallback needs a provider key of your own.

What you need

  • An application on a plan whose allowance is set: read_plan_quotas shows its quantity.
  • No provider account and no provider key.

Before your AI starts

This section is for your AI tool: what it checks and gathers before it begins. You don't need to do these steps yourself.

  • Your tool connected and signed in (Connect your tool).
  • list_applications returns the application's identifier.
  • The backend's folder, the one that contains package.json.

Steps

The AI Allowance package page holds the request members, the allowance members, and the refusals. This guide links it rather than repeating its tables.

1. Take the package and declare it

Your tool takes the ai_allowance entry as Use the library describes. It copies the entry into app/lib/ai_allowance/, declares "@turnzero/ai_allowance": "file:lib/ai_allowance" in package.json, and runs npm install. At the next submit_manifest, the manifest's packages member names the package and its version.

The call goes through the platform's own route, on a key the platform keeps. So the application declares no egress entry, no upstream, and no stored key for it.

2. Call generateText

generateText takes a transport your application supplies, a request, and options. The transport adds the platform's origin and the credential as the bearer. A deployed copy reads both from the settings it receives.

This sample asks for a one-line summary and reads the allowance figures. It falls back where the month's units are drawn or the platform's key is not in place. No test checks this sample.

import { AllowanceRefusal, AllowanceUpstreamError, generateText } from '@turnzero/ai_allowance';

const origin = process.env.TURNZERO_CLOUD_GATEWAY_URL ?? process.env.TURNZERO_CLOUD_API;

// Adds the platform's origin and the deployed copy's platform credential.
const transport = (path, init = {}) =>
  fetch(origin + path, { ...init, headers: { ...init.headers, authorization: `Bearer ${process.env.TURNZERO_CLOUD_TOKEN}` } });

export async function summarize(note) {
  try {
    const answer = await generateText(transport, { prompt: `Summarize this note in one line: ${note}`, maxOutputTokens: 200 });
    console.log(`allowance ${answer.allowance.used} of ${answer.allowance.quota}`);
    return answer.text;
  } catch (error) {
    const drawn = error instanceof AllowanceRefusal && ['allowance_exhausted', 'allowance_unavailable'].includes(error.refusal);
    const providerFull = error instanceof AllowanceUpstreamError && error.quota;
    if (drawn || providerFull) return null; // Or call your own declared upstream here.
    throw error;
  }
}

For a structured answer, pass a JSON Schema as responseSchema and parse text as JSON. For a streamed answer, pass onText in the options.

3. Read what is left

Every answer carries the allowance figures in its allowance member: the units drawn this month, the plan's quantity, and the first instant of the next UTC month. read_plan_quotas returns each plan's quantity as its gemini-flash-allowance row. read_usage returns the month's draw as the measure ai_allowance_units.

Every environment of the application, and its local runs, draw on the same quantity. When the month's units first pass eighty percent of it, the platform writes the event usage warning into the application's log stream.

4. Fall back when the units run out

Choose the fallback by the refusal's name, never by its message. On allowance_exhausted or allowance_unavailable, wait for resetsAt or call your own declared upstream on your own stored key. The boundary and the fallback gives the whole rule.

Never retry an exhausted allowance inside the month. Every further call is refused, and each refusal is counted. Call an external API with an API key declares the upstream a fallback calls.

5. Test over the double

The package ships a test double at @turnzero/ai_allowance/testing. It takes the place of the allowance route behind the unchanged generateText, so a test needs no gateway, no key, and no provider. This test primes one answer and calls through the double's transport. A test runs this sample and checks that the primed answer comes back.

import { generateText } from '@turnzero/ai_allowance';
import { createAiAllowanceDouble } from '@turnzero/ai_allowance/testing';

const double = createAiAllowanceDouble({ quota: 1_000_000 });
double.prime({ text: 'A note about lunch.', usage: { input: 40, output: 12 } });
const answer = await generateText(double.transport, { prompt: 'Summarize this note in one line: lunch at noon' });

Test your application locally shows the double in a suite, its refusals among them.

6. Deploy

Your tool deploys as Deploy an application describes. The deploy injects the platform credential the transport sends, bound to its environment, so the call takes no environment header.

Expected result

A call returns the model's text, the model, the usage, the finish reason, and the allowance figures. After a call, read_usage shows ai_allowance_units above its earlier figure. Once the units are drawn, the call is refused allowance_exhausted and the application answers through its fallback.

Refusals

Refusal Status Cause Remedy
allowance_application_required 403 An allowance call used an account-wide credential, such as a session or an account-scoped token. Call under the application's platform credential, or a token for the application with the environment header.
allowance_busy 429 The application already has the most allowance calls in progress that the platform allows. Retry after one of them ends.
allowance_exhausted 429 This UTC month's allowance is used up; the refusal includes resets_at, used, and quota. Wait for resets_at, move to a larger plan, or call your own upstream. Retrying this month is refused and counted each time.
allowance_feature_refused 400 The body is not JSON, or asks for tools, cached content, or output other than text. Ask for text generation only, or use your own key through a declared upstream.
allowance_path_refused 403 The request is not a POST to the text-generation method of a model the platform allows. Call through the AI Allowance package's generateText function, with an allowed model.
allowance_unavailable 503 The platform has no key for the allowance yet. Call your own declared upstream, and report it if it persists.
plan_quantity_unset 409 An allowance call ran under a plan whose gemini-flash-allowance quantity is not set. Wait for platform staff to set the quantity, or move to another plan. read_plan_quotas shows it as a null quantity.