BlueprintService end to end
BlueprintService is a small application that runs on Turn Zero Cloud in production. This page shows how an application that must never issue the same number twice is built on one database and nothing else.
The service issues new requirement numbers and decision entry numbers for Turn Zero Blueprint projects. It also keeps a queue of a project's integrations, so that only one runs at a time. In Turn Zero Blueprint, an integration merges a work branch into the project's primary branch. The service also keeps landing holds. A landing hold is a time-limited claim on a set of files, which stops other integrations from changing those files while a release is using them.
The service also keeps a queue freeze for a project while a hotfix stands. A queue freeze is a record naming the branches the hotfix needs, which the service answers beside the holds when a caller lists them.
The service's whole source is an example project: one manifest, one server file, a copy of the Database package, a build script, and a test suite. The example project is not published, so this page quotes the parts it explains. It walks through the project from the manifest to the tests, and gives each file's path inside the project. Each sample below is a continuous part of one file, and a test checks that it matches that file byte for byte.
Turn Zero Blueprint documents what the service does for its callers. The BlueprintService page of the Turn Zero Blueprint documentation explains the numbers, the queue, and what the service keeps.
The manifest
app/app/manifest.json declares everything the application needs, and the sample is the whole file.
{
"manifest_version": 1,
"services": [
{
"kind": "database"
}
],
"health": "/health",
"region": "usa",
"egress": [],
"audience": {
"kind": "public"
},
"packages": [
{
"name": "database",
"version": "0.10.2"
}
]
}
The one service is database, so the platform gives each environment a PostgreSQL database and passes its connection string in APP_DATABASE_URL. Add a database describes what the declaration provisions. The empty egress list declares that the application calls no external host. The public audience accepts a caller with no end-user session. That suits this service, whose callers are programs that hold a project id.
The packages list declares one library package, the Database package at version 0.10.2. The server uses the package's failover pool, Node.js built-ins, and the pg driver. A copy of the package goes into the application's build from app/app/lib/. The manifest defines each member.
The server over its pool
app/app/server.mjs runs one node:http listener and exports two factory functions:
createBlueprintService({ pool, now, constants })returns the request handlers, built over a pool and a clock the caller supplies.createBlueprintServiceServer({ service, version, guarded })returns an HTTP server that is not yet listening.
Neither factory reads a setting, so a test builds both over a test double's pool and a clock the test controls.
Only the process entry at the end of app/app/server.mjs reads the settings. It runs only when the file is the process's main script.
if (isMain) {
const env = process.env;
const port = Number(env.PORT ?? 8080);
const databaseUrl = typeof env.APP_DATABASE_URL === 'string' && env.APP_DATABASE_URL.length > 0 ? env.APP_DATABASE_URL : null;
const guarded = guardedValuesOf(databaseUrl ?? '');
const print = (line) => console.error(guardLine(line, guarded));
let pool = null;
const service = createBlueprintService({ pool: () => pool, now: Date.now });
const server = createBlueprintServiceServer({ service, version: env.APP_VERSION ?? null, guarded });
server.listen(port, () => {
console.log(`blueprint-service listening on port ${port}`);
});
if (databaseUrl) {
pool = openServicePool(databaseUrl, (error) => print(`blueprint-service pool error: ${describeFault(error)}`));
service.ensureTables().then(
() => console.log('blueprint-service tables ensured'),
(error) => print(`blueprint-service boot table ensure failed, retried before the next request: ${describeFault(error)}`),
);
} else {
console.error('blueprint-service: no APP_DATABASE_URL is set; every call answers 500 until the deploy injects one');
}
}
The entry listens on PORT before it does anything else. The /health route returns the APP_VERSION the platform set and checks nothing else, so the deploy's health check succeeds even while the database is still unreachable. The pool opens afterwards, and the service reaches it through a getter. Every line that reports an error passes through guardLine, which replaces the database URL and its parts with a placeholder.
openServicePool in app/app/server.mjs builds the pool. It is the Database package's failover pool over a pg pool, with one client, a time limit on connecting, and a time limit on waiting for a lock.
export function openServicePool(connectionString, report) {
return createFailoverPool({
connect: (settings) =>
new pg.Pool({
connectionString,
...settings,
// A client-side bound beside the role's own statement_timeout, and a
// keep-alive, so a session the network drops silently is not held.
query_timeout: QUERY_TIMEOUT_MS,
// Sent as a startup parameter, so every statement's lock wait is bounded,
// the mint's row wait outside any transaction of the service's included.
lock_timeout: LOCK_TIMEOUT_MS,
keepAlive: true,
}),
max: POOL_MAX,
connectTimeoutMs: CONNECT_TIMEOUT_MS,
onConnectionLoss: report,
});
}
createFailoverPool calls the connect function once. It passes the two settings it owns, the most clients and the time limit on connecting, and the function builds the pg pool with them.
One client per process stays well within the database's limit. The database role accepts twice the plan's connection_limit, which is 20 connections on Pro, where the limit is 10. Each environment runs one copy of the process, two only while a deploy replaces it. Add a database explains the connection limit, and read_plan_quotas returns each plan's value. If the database rejects a connection, the call returns 500. A 500 in the logs is therefore the signal to check connection use.
The pool limits how long a call waits:
- A call waits at most five seconds for the one client. If calls often wait that long, the process needs more than one client.
- The
lock_timeoutstartup parameter limits every statement's wait for a lock to two seconds, including the wait of the call that issues numbers for its row.
A database failover ends every session the server had open. The failover pool listens for that error on the pg pool, for an idle client, and on each client the service has checked out. It passes each ended session to report, which prints one guarded line, and the process keeps serving. The service adds no error listener of its own.
The Database package's integration guide shows how to create the failover pool, how its boot retry works, and when to call client.release(error).
The five tables created at boot
The application ships no migration files. app/app/server.mjs keeps its schema as a frozen list of Structured Query Language (SQL) statements. Each statement creates only what is missing, so an existing copy gains a new table at its next boot and loses nothing.
export const TABLE_STATEMENTS = Object.freeze([
`CREATE TABLE IF NOT EXISTS bps_counters (
project text NOT NULL,
scope text NOT NULL,
value bigint NOT NULL,
PRIMARY KEY (project, scope)
)`,
'CREATE SEQUENCE IF NOT EXISTS bps_entry_seq',
'CREATE SEQUENCE IF NOT EXISTS bps_order_seq',
`CREATE TABLE IF NOT EXISTS bps_tickets (
project text NOT NULL,
ticket text NOT NULL,
branch text NOT NULL,
head text NOT NULL,
label text NOT NULL,
step text NOT NULL DEFAULT '',
entry bigint NOT NULL,
queue_order bigint NOT NULL,
entered_at bigint NOT NULL,
touched_at bigint NOT NULL,
turn_at bigint,
ended_at bigint,
PRIMARY KEY (project, ticket)
)`,
'CREATE INDEX IF NOT EXISTS bps_tickets_queue ON bps_tickets (project, queue_order)',
'CREATE INDEX IF NOT EXISTS bps_tickets_entered ON bps_tickets (entered_at)',
`CREATE TABLE IF NOT EXISTS bps_retired (
project text PRIMARY KEY,
replacement text NOT NULL,
retired_at bigint NOT NULL
)`,
`CREATE TABLE IF NOT EXISTS bps_holds (
project text NOT NULL,
name text NOT NULL,
paths text NOT NULL,
owner text NOT NULL,
opened_at bigint NOT NULL,
until_at bigint NOT NULL,
PRIMARY KEY (project, name)
)`,
'CREATE INDEX IF NOT EXISTS bps_holds_until ON bps_holds (until_at)',
`CREATE TABLE IF NOT EXISTS bps_freezes (
project text PRIMARY KEY,
issue text NOT NULL,
declared_at bigint NOT NULL,
opened_at bigint NOT NULL,
opener text NOT NULL,
admissions text NOT NULL
)`,
]);
The counters table has one row for each project and number line, a separate run of numbers within a project. The tickets table contains the integration queue, one ticket for each integration that entered it. The retired table records each project id that was replaced by a new id, and the holds table contains each open landing hold. Two sequences issue the entry numbers and the queue order, so both only increase.
The freezes table has at most one row for each project, its queue freeze, with the branches it admits stored as one JSON list. That table has no index beyond its key, because every call reads one project's freeze by its id.
Every time value is a bigint of milliseconds from the clock passed into the factory, never from the database's own clock. That lets a test move time forward.
createTables, inside the service factory in app/app/server.mjs, runs the list in one transaction under a transaction-scoped advisory lock. During a deploy, the old and the new container may both run it at once, and the lock makes them create the tables one after the other. The transaction first limits its lock wait, so a boot that waits for another boot for more than two seconds fails and is tried again.
const createTables = async () => {
const client = await poolOf().connect();
let failure;
try {
await client.query('BEGIN');
await client.query(SQL.lockTimeout);
await client.query(SQL.tablesLock);
for (const statement of TABLE_STATEMENTS) await client.query(statement);
await client.query('COMMIT');
} catch (error) {
failure = await rolledBack(client, error, false);
throw error;
} finally {
client.release(failure);
}
};
The boot calls this once. If the boot's attempt fails, the next request starts a fresh attempt before its first statement, and a request that arrives during an attempt waits for it. While the client is out of the pool, the failover pool's listener covers it, so the function adds none.
rolledBack rolls a failed transaction back and returns the value to release the client with. If the database session is still usable, it returns nothing, and the client goes back to the pool. A lock wait that timed out keeps its client this way. Otherwise it returns the error, and the pool discards the client.
The mint runs as one statement
A mint is the call that issues a contiguous block of numbers on one number line. During a deploy, two processes can handle two mints on the same line at the same instant. A read followed by a separate write could then give both callers the same number. So app/app/server.mjs writes the mint as one SQL statement that reads and writes together, which PostgreSQL runs atomically on the row.
/** The mint: one statement that reads and writes together. The counter
* becomes the greater of itself and the floor, plus the count; where the
* floor would lift an existing counter by more than the lift bound the
* conditional update writes nothing and no row returns. $1 project, $2 scope,
* $3 floor, $4 count, $5 the lift bound. */
export const MINT_STATEMENT =
'INSERT INTO bps_counters AS t (project, scope, value) VALUES ($1, $2, $3::bigint + $4::bigint) ' +
'ON CONFLICT (project, scope) DO UPDATE SET value = GREATEST(t.value, $3::bigint) + $4::bigint ' +
'WHERE $3::bigint - t.value <= $5::bigint RETURNING value';
The caller may send a floor, the highest number its own files already use. The WHERE clause limits how far a floor may raise an existing counter. Past the limit, the update writes nothing and returns no row, and the mint function in app/app/server.mjs then refuses the call.
const mint = (b, problem) =>
underLock(
[b.project],
async (db) => {
if ((await retiredAs(db, b.project)) !== null) throw retiredRefusal();
if (problem) throw problem;
// A member present as null was refused above; absent reads as its default.
const count = 'count' in b ? b.count : 1;
const { rows } = await db(MINT_STATEMENT, [b.project, b.scope, 'floor' in b ? b.floor : 0, count, k.floorLiftMax]);
if (rows.length === 0) {
throw new Refusal(
409,
'floor_too_far',
`this floor would lift the number line by more than ${k.floorLiftMax}; a project whose numbers moved that far takes a fresh id`,
);
}
const last = Number(rows[0].value);
return { project: b.project, scope: b.scope, first: last - count + 1, last, count };
},
{ shared: true },
);
The statement alone stops two mints from returning the same number. The mint still runs inside underLock, the transaction the next section describes, under a shared lock on its project. Several mints on one project can share that lock at once, and replacing the project id takes it exclusively. So a mint cannot find its id still current, then lose to a replacement that copies the counters, and then raise the old id's counter after the copy. The statement returns one value, the last number of the block, and the function computes the first number from it.
A queue call is a transaction under its project's lock
A queue call reads the project's live tickets, works out the result, and writes it, using several statements. underLock in app/app/server.mjs runs those statements as one unit: one client, one transaction, and a transaction-scoped advisory lock keyed by the project id. The call waits at most two seconds for the lock. A queue call, a hold call, and a freeze call take the lock exclusively, and a mint takes it shared.
const underLock = async (projects, work, { shared = false } = {}) => {
const client = await poolOf().connect();
let failure;
try {
await client.query('BEGIN');
await client.query(SQL.lockTimeout);
const lock = shared ? SQL.projectShared : SQL.projectLock;
for (const id of [...new Set(projects)].sort()) await client.query(lock, [`blueprint-service:${id}`]);
const db = (text, values) => client.query(text, values);
const answer = await work(db, now());
await client.query('COMMIT');
return answer;
} catch (error) {
failure = await rolledBack(client, error, error instanceof Refusal);
throw error;
} finally {
client.release(failure);
}
};
The lock statement calls pg_advisory_xact_lock, or pg_advisory_xact_lock_shared for a mint. PostgreSQL releases the lock at commit or rollback, so no code path can leave it held. A queue call and a mint on one project wait for each other, each wait limited to the same two seconds. Replacing a project id involves two projects and takes both locks in sorted order, so two replacements cannot deadlock.
Every failure rolls the transaction back:
- A refusal thrown inside the work rolls back, so a refused call changes nothing.
- A lock wait that times out fails the call. The database session is still usable, so the client is kept.
- Any other failure returns the client to the pool marked as broken.
Every queue call goes through one wrapper in app/app/server.mjs. Inside the lock, it checks whether the project id is retired and whether the call's members are well formed, before the call's work runs.
const queueCall = (call, b, problem) =>
underLock([b.project], async (db, at) => {
if ((await retiredAs(db, b.project)) !== null) throw retiredRefusal();
if (problem) throw problem;
return QUEUE[call](db, b, at);
});
Whether a ticket is live is never stored. Each read works it out from the supplied clock: a ticket is live while its lease was renewed recently enough and its turn has not run past its limit. The turn belongs to the live ticket whose turn has begun. If none has begun, the first read that finds no holder starts the turn of the first ticket in queue order. No background job runs, so a stopped copy of the process leaves nothing to clean up.
Two clean-up deletions run inside ordinary calls. queue/enter deletes up to 100 tickets that entered more than seven days before, the oldest first. hold/open deletes up to 100 landing holds that have ended, the earliest ended first. A queue freeze has no end time, so no clean-up deletes it. freeze/close deletes it when the call names every issue the freeze serves.
The refusals and their remedies
A refusal is a result the caller can respond to. app/app/server.mjs models one as an error with an HTTP status, a name from the contract, and a detail sentence.
class Refusal extends Error {
constructor(status, code, detail) {
super(detail);
this.status = status;
this.code = code;
this.detail = detail;
}
}
const invalid = (detail) => new Refusal(400, 'invalid_request', detail);
const retiredRefusal = () =>
new Refusal(410, 'project_retired', 'this project id is retired; fetch the trunk and merge it, which carries the id the project uses now');
Each detail says what went wrong and what the caller does next. The retired-id detail tells the caller to fetch and merge the primary branch, which has the id the project uses now. The detail does not give the replacement id. The leave handler in app/app/server.mjs shows the same pattern for a call that arrives after its ticket expired.
const leave = async (db, b, at) => {
const t = await current(db, b, at);
if (!t) throw new Refusal(409, 'ticket_gone', 'this ticket or this entry number no longer stands; nothing was changed — enter again to take a fresh place in the queue');
await db(SQL.end, [b.project, t.ticket, at]);
await queueOf(db, b.project, at);
return { left: true, outcome: b.outcome };
};
The service's handle function catches a Refusal and returns its status with { error, detail }. It rethrows any other error. The HTTP layer then returns 500 with a fixed body and prints one guarded line. The guarded line replaces the database URL and its parts, and every universally unique identifier (UUID). It also replaces each longer string of the request, whether a member's or an array member's. No printed line contains a project id, a request body, or a credential.
This page does not list the refusals. The Results and errors page of the Turn Zero Blueprint documentation lists each refusal the service returns, its cause, and the remedy a Blueprint developer follows.
One case list, two implementations
Turn Zero Blueprint ships a test double of the service that runs inside the test process. Its own test suite uses the double instead of calling the network. The double and this server are separate implementations of one contract, and one case list checks that both return the same results. The list is a JSON file in Turn Zero Blueprint's contract for the service. It contains the contract's constants and then the named cases. Each case is a sequence of steps: a call with its body and the expected status and result, an advance of the clock, or a repeat.
The constants come first, because both implementations must refuse at the same limits. app/app/server.mjs defines them once.
export const CONSTANTS = Object.freeze({
leaseMs: 300_000,
renewMs: 60_000,
turnMs: 5_400_000,
ticketAgeMs: 604_800_000,
mintCountMax: 100,
floorLiftMax: 1_000,
floorMax: 1_000_000_000,
liveTicketsMax: 200,
holdSpanMaxMs: 86_400_000,
holdPathsMax: 64,
openHoldsMax: 20,
holdPathBytesMax: 4_096,
openHoldsTotalMax: 2_000,
freezeBranchesMax: 32,
rateWindowMs: 60_000,
projectRateMax: 300,
addressRateMax: 600,
bodyBytesMax: 16_384,
});
The first case of app/tests/blueprint_service.test.ts checks three things: the whole list is valid, the server handles the same calls the runner lists, and the constants match. The body of that case follows, without its title.
expect(runner.checkCaseList(LIST)).toBe(62);
expect([...mod.CALLS]).toEqual([...runner.CALLS]);
expect({ ...mod.CONSTANTS }).toEqual(LIST.constants);
Every case of the list then runs as a test of its own, reported by the case's name. The body of each is one line of app/tests/blueprint_service.test.ts. The shared runner builds a fresh service through the factory it receives, using the database double's pool and the case's own clock.
await runner.runCase(caseDef, ({ now, constants }) => mod.createBlueprintService({ pool: double.pool, now, constants }), LIST.constants);
The handlers run over the Database package's test double, a PostgreSQL engine inside the test process, so the suite needs no running database. Test your application locally describes the double. The double's reset() empties the tables before each case. Turn Zero Blueprint's suite gives the same runner its own double. Changing the contract means changing the list, and an implementation that was not updated fails under the case's name.
The list leaves out what one database session cannot show. The double has one session, so it cannot show two processes minting at once, as happens during a deploy. The suite instead checks that the mint is one statement, by its text and by the statements the double recorded. The same file tests the HTTP layer, the health path, and the body size limit on a temporary port. app/build.mjs zips the four files under app/app/ and the package's copy under app/app/lib/ with fixed timestamps, so the same content always gives the same SHA-256 hash, which deploy compares.
Related
- The manifest — the members the application's manifest declares.
- Add a database — the declaration, the connection setting, and the connection limit.
- Test your application locally — the database double the suite runs the handlers over.
- Applications and environments — one environment by default, and development with the promote when you turn it on.
- BlueprintService — the Turn Zero Blueprint page on what the service does, receives, and keeps.
- Results and errors — the Turn Zero Blueprint page that lists the service's refusals and their remedies.
- Glossary — the terms used on this page.