Egress firewall and request limits
The platform enforces two sets of limits on a deployed backend. The egress firewall governs its outbound network access, and each plan limits how much outbound traffic it sends. The request limits govern how long a request may run and which incoming requests reach it. This page describes those limits, the two egress modes, and where to read the records of each limit. How much the records show depends on the client and route your code uses.
Outbound network routes
Application code reaches external services through the egress proxy, and reaches platform endpoints directly. The proxy's checks apply only to traffic sent through it. The platform's network blocks any other call, which then ends with the client's own connection timeout.
The proxy offers two ways to call an external service: the HTTPS tunnel and the gateway.
The HTTPS tunnel
The backend opens an HTTPS connection through the proxy on port 443.
- Allowed destinations. The manifest's
egresslist gives them as exact lowercase hostnames or patterns such as*.example.com. That pattern matches subdomains, notexample.comitself. - Modes. Observe mode allows undeclared destinations; enforce mode refuses them.
- Addresses. The proxy pins each connection to the public address it resolved when the connection opened. Private, loopback, link-local, reserved, and platform addresses are refused in both modes.
- Mail ports. Ports 25, 465, 587, and 2525 are outbound mail ports, refused like every port other than 443. Applications send no mail directly: send it through a mail provider's HTTPS application programming interface (API), declared as an upstream.
- Metering. Each connection is metered when it closes: the host and port, the bytes each way, the duration, and the outcome. The bytes sent to the destination count toward the data-transfer measure, as the gateway's do.
The gateway
The backend calls /egress/v0/<upstream>/<path> on the platform's origin with its platform credential. The gateway adds the stored credential for the declared upstream and forwards the request, so the application never receives that credential.
A declared upstream's base URL cannot name an outbound mail port (25, 465, 587, or 2525). declare_upstream refuses it. For an upstream declared before this rule, the gateway refuses each call instead. Send mail through a mail provider's HTTPS API, declared as an upstream.
The platform credential is bound to one environment. The gateway reads the upstream's key at that environment's secret scope or at account scope, never from the other environment. A session or a minted token must name the environment in the header x-turnzero-cloud-environment. Without it the call is refused 400 environment_required. Under the platform credential, a header naming another environment is refused 403 environment_mismatch. See Call an external API with an API key.
Platform endpoints
The platform origin is https://turnzero.ai, which serves the storage, egress, and logging endpoints by path. Your code reaches the platform origin and declared platform service endpoints outside the proxy, and they need no entry in egress.
Observe before enforce
A new application begins in observe mode. The tunnel records each connection attempt and allows an undeclared destination that passes its other checks. Platform staff switch an application to enforce mode with set_egress_mode; you cannot switch it yourself. read_status reports the current mode as egress_mode.
In enforce mode the proxy refuses an undeclared destination with Hypertext Transfer Protocol (HTTP) status 403 egress_undeclared. It records the host it refused and, for a hostname, the manifest edit that would allow it. A client that shows the proxy's response can also read its X-Egress-Refusal header; other clients may show only the status. Authentication, port rules, and address checks apply in both modes.
How an undeclared target is refused depends on its form:
- An address literal in a private, loopback, link-local, reserved, or platform range is refused by its address range in both modes. The refusal gives no manifest edit, because the
egresslist accepts no address literal. - A hostname is refused
egress_undeclaredin enforce mode before the proxy resolves it. In observe mode the proxy resolves it, and refuses it by address range if it resolves into one of those ranges.
Undeclared destinations in read_status
read_status lists the undeclared destinations that the serving version reached in observe mode. Read it after a deploy to find the hostnames to declare before the application moves to enforce mode. The list is the egress member inside the response's application member. Where the version reached an undeclared destination, the response's summary gives their count and names that member. Where it reached none, the summary says nothing about egress.
Each entry in undeclared gives the host and port, the number of connections, the time of the last one, and the remedy. For a hostname, the remedy is the manifest edit that declares it. For an address literal, it says to declare the hostname the address serves. The list contains at most twenty entries, and total counts every destination found.
The list covers the environment the response describes, from the later of two times: when its serving version started, or seven days before the read. since gives that start. Each read_status call reads a limited number of records. Where the read stops at that limit before reaching the start, since is the oldest record it read, and the list covers that shorter window. For anything before since, use read_logs with filter: "undeclared". The member's earlier says so in every response.
read says how the list was made. It is logs where the records were read, unavailable where the log store did not respond, in time or at all, and skipped while a deploy of that environment is in flight. A skipped list is empty and has no since.
The list leaves out refused connections: refusals in enforce mode, refusals of an undeclared host by port or address range, and connections a client outside the proxy failed to open. It also leaves out a host that was declared when it was reached and has since been removed from the manifest. read_logs with filter: "undeclared" returns both kinds of record. A host the manifest now declares is dropped from the list.
The refusal in your code
Through the runtime harness, a refused call reaches your code as an EgressRefusedError. The error has code: 'EGRESS_REFUSED', the refusal's name, the status, the host, the port, and the remedy. Global fetch rejects with TypeError: fetch failed, whose cause is that error. An undici client routed by the harness receives the error itself. Your code can read the refusal's name from the error without reading the logs.
When the control plane cannot be reached
Each proxy replica keeps its own cached copy of an application's authorization, declarations, and mode. The cache normally reflects a change within its usual interval, but only while the control plane responds.
If the control plane cannot be reached, a replica keeps using a cached answer for new connections until the answer's age reaches the proxy's maximum authorization age, 3,600 seconds by default. The age counts from the replica's last successful read. Within that time, a change such as an account suspension can take effect late.
Once the age is reached, the replica refuses new connections with HTTP 503 egress_resolve_unavailable until the control plane responds again. A replica with no cached answer refuses the same way, which can also follow a cache eviction or a replica restart. Connections already open are not cut; they run to their own limits.
Limits per plan
Each plan sets two limits on an application's outbound traffic. All of the application's outbound traffic counts toward the same two limits, whichever environment it comes from.
- Connections per minute. The new connections the application opens through the HTTPS tunnel in one UTC minute. Gateway calls do not count toward it.
- Bytes per day. The bytes the application sends and receives in one UTC day, through the tunnel and through gateway calls to its own declared upstreams. Calls to the platform's own upstreams,
ai-allowanceandissue-tracking, do not count toward it.
The proxy runs one to three replicas, and each replica counts the minute's connections on its own. So an application whose connections spread over the replicas can open up to the limit through each one. The gateway keeps no count per minute. Its calls are limited instead by the platform's request limit per credential per minute, which refuses with HTTP 429 rate_capped.
The limits apply in observe and enforce mode alike. They limit how much traffic leaves, never where it goes. In enforce mode, an undeclared destination is still refused egress_undeclared first.
At the limit
At either limit, the proxy refuses a new connection with HTTP 429: egress_connections_capped at the minute limit, and egress_daily_bytes_capped at the day limit. At the minute limit, connections already open keep running.
At the day limit, the proxy also cuts each open tunnel connection of the application at its next byte. It records the cut as refused_bounds in the connection's record under the egress source of read_logs, naming egress-bytes-per-day as the limit passed. Your client receives no refusal name for the cut. The gateway refuses calls to the application's own upstreams with HTTP 429 egress_daily_bytes_capped. A refused call reads no key and sends nothing.
Each refusal's body contains the count, the limit, and resets_at, the time the limit resets. Your code reads them in a gateway refusal, and in a proxy refusal where its client shows the proxy's response. On the runtime harness, a tunnel refusal reaches your code with the refusal's name and the status 429, without the three figures. The tunnel's record of the refusal, under the egress source of read_logs, contains all three.
The minute limit resets at the start of the next UTC minute, and the day limit at the start of the next UTC day.
Where the figures are read
read_usage returns an egress_limits member for each application. It contains the two limits, bytes_today, refused_today, a state, and resets_at, the start of the next UTC day. state is ok under the day limit, capped at or past it, and unset where the day limit is not set. read_plan_quotas lists each plan's two limits as egress-connections-per-minute and egress-bytes-per-day. Plan and usage describes both actions.
Each plan's two figures are set from the start. When platform staff clear a figure, that limit refuses nothing: read_plan_quotas then shows a null quantity, and egress_limits shows that limit as null. The platform counts the day's bytes either way. A changed figure applies on the gateway within a minute, and on the tunnel within a minute plus the proxy's usual cache interval.
How exact the limits are
The day limit is close, not exact. The figure each check reads trails the traffic, so an application can pass the limit by a margin before the refusals start:
- Each proxy replica reads the application's day total from the platform with its cached copy, at the cache's usual interval. Between reads it adds its own connections' bytes, but not the other replicas'.
- A connection's bytes reach the day total within two reporting intervals of its close, so a replica's read can miss them.
- On the gateway, the figure trails its own calls by one reporting interval. Calls already running when the limit is reached still complete, each up to the gateway's response size limit.
- A tunnel connection open across midnight UTC counts toward the day it opened. Once it closes, the bytes it transferred after midnight leave the new day's total. Each such connection can pass the new day's limit by up to the tunnel's byte limit for one connection.
- The platform's servers can differ slightly in their clocks, which moves the start of the day by that difference.
bytes_today in read_usage trails the traffic in the same way. If the gateway cannot read the day's figures, it allows calls without the day check for the next minute.
Covered clients
The runtime harness routes the runtime's global fetch, undici, and SDKs built on them through the proxy. Axios uses the proxy environment variables the harness sets. Other clients, including node:https, node-fetch, got, and ws, need a proxy agent configured from HTTPS_PROXY.
A call that is not routed through the proxy usually fails with the client's connection timeout, not a named proxy refusal. Native addons and child processes can open connections the harness does not observe, so its logs may not show all their traffic.
The execution window
Ordinary HTTP request handlers have a set time limit. At the deadline the serving router responds with HTTP 504 window_ended if the response headers have not been sent; otherwise it cuts the streaming response. The handler's request.signal is aborted. Upgraded connections are outside this deadline, but the runtime harness still checks the router mark on their initial upgrade requests.
If the handler does not stop within the grace interval, the harness ends the process, which also interrupts the process's other requests. The harness records process_ended. Each interrupted caller that had not yet received response headers gets the router's 502 upstream_ended. An intermediary can retry an idempotent request on the restarted process, and that caller then receives the new response.
Work does not reliably continue after a handler responds. A timer, polling loop, detached promise, or child process can be ended when its container process ends.
The per-address and per-user limits
The serving router limits how fast one client address, and one signed-in user, can reach your application. It keeps three counts per application, each per minute and per router replica:
- A request with no valid session counts against its client address, which allows 300 a minute.
- A request with a valid session, as the cookie or as a bearer token, counts against the user, who is allowed 300 a minute. It also counts against the address's signed-in count, which allows 30,000 a minute. It does not count against the address's anonymous count.
- An IPv6 client address counts as its /64 prefix, and the production and development hostnames share one count.
A request past any bound it counted against is refused with HTTP 429 rate_capped and a Retry-After header giving the seconds to the next minute. It never reaches your application and uses none of its monthly backend actions or data transfer. A bearer the router cannot verify counts once against the source address before it is refused (Verifying an end user).
The bound covers every request that would reach your application. It does not cover the page files the router serves from the assets store, the end-user sign-in routes under /__account/, the three well-known documents, or scheduled runs.
The counts are kept per router replica, and the router runs at least three replicas. One client address whose requests spread over the replicas can therefore get up to 300 a minute through each, and more as the platform adds replicas. Many users behind one address, such as an office or a mobile carrier's gateway, share the address's counts. Each signed-in user among them has a count of their own, and the address allows a thousand signed-in users at thirty requests a minute each.
The bound slows one source; it does not stop traffic from many sources using up the month's allowance. For each minute in which the bound refused requests, the router writes one source_refused record into your application's logs, under the router source below. The record gives the number of refusals under each of the three counts, the total, and the number of addresses refused. It never gives an address or a user.
The platform request limit
The platform also limits how fast one client address can call the platform itself: the MCP server, the management actions, the sign-in routes, and the pages of the site. This limit is separate from the limits above, which apply to your application.
Every request to the platform counts once, against the address it comes from:
- A page view of the site or its documentation counts toward the address's page limit, 600 a minute. The stylesheet and the files under
/assets/do not count. - A read of the library over HTTP,
list_libraryorread_library_entry, counts toward the address's library limit, 3,000 a minute. Taking a library entry reads its list and then each of its files, so one take makes many reads. They do not count toward the request limit below. - Every other request counts toward the address's request limit, 600 a minute. That covers
/mcp, every other management action, and the sign-in routes. The actions that need no sign-in, such aslist_contextandread_context, count too. - A dashboard page or the approval page counts toward the request limit only while you are signed out. Once you are signed in, opening those pages is not counted.
- A call on the storage, egress, logging, or push wire that presents a valid credential counts against that credential instead, 600 a minute. Applications behind one address therefore do not share a count there, and two credentials of one application are counted apart.
- A call on one of those wires with a credential the platform does not know, such as a revoked or mistyped one, returns 401. It counts toward a separate limit for the address, 600 a minute, and not toward the address's other limits.
The page limit, the library limit, the request limit, and the limit for calls with an unknown credential are counted apart. Reading many pages or taking library entries does not use up the address's calls to /mcp, and many calls do not stop the address reading pages.
A request past a limit is refused with HTTP 429 rate_capped, before the platform reads the request's body. The response has a Retry-After header giving the seconds to the next minute, when the count starts again. Wait that long, then repeat the request. A page view past the limit gets the same JSON response, not a page.
People behind one address, such as an office, share that address's limits.
When the gateway is busy checking credentials
On the storage, egress, logging, and push wires, the gateway checks each credential it is shown and keeps the answer for 30 seconds. A call under a credential it checked within that time needs no new check.
Each gateway replica checks a limited number of new credentials a second. A call it cannot check within two seconds is refused with HTTP 503 credential_lookup_busy and a Retry-After header of 1. A call can also be refused this way at once, while many new credentials wait to be checked, especially from one address. Nothing was done, so repeat the same call after one second. The refusal does not count toward any limit on this page.
A call through https://turnzero.ai from an address past its limit for unknown credentials is refused 429 rate_capped before the check. A credential the gateway checked within the last 30 seconds is not refused this way.
An upload or download grant is kept for those 30 seconds too. A grant the platform ends, such as one whose application was deleted, can serve calls until the 30 seconds pass. Its expiry and an upload grant's one write are checked on every call, as before. In those 30 seconds, a deploy upload's grant that a later deploy call replaced writes nothing, but other bytes sent under it can be refused 409 upload_hash_mismatch rather than naming the replacement.
The router mark
The application is reachable only through its managed hostnames: <label>.ai.host for production and <label>-dev.ai.host for development. Every response on the development hostname has the header X-Robots-Tag: noindex, nofollow. The serving router marks each request it forwards and sets its forwarded headers. A request that reaches the application's container by any other address, without that mark, is refused 403 router_mark_required before your own listener runs.
What the router sets on each request
The serving router sets three headers on every request it forwards to your application:
| Header | What it contains |
|---|---|
x-forwarded-host |
The hostname the request arrived on. |
x-forwarded-proto |
Always https, because the encrypted connection ends before the router. |
x-forwarded-for |
The client's address as the router knows it. Read the first entry, since a hop after the router may append its own. |
The router sets the three after the client's headers, so a client's copy never arrives. The Host header names the platform's internal address for the environment, never the public hostname. Build an absolute address from x-forwarded-proto and x-forwarded-host, or from APP_PUBLIC_HOST. The router sets no other forwarding header, such as Forwarded or x-forwarded-port, so trust none. The router also sets the headers of a scheduled run and of a signed-in end user, which those pages describe.
A local run has no router and no APP_PUBLIC_HOST. Fall back to the request's Host header and http. A test that requests a local server expects http addresses, or sends x-forwarded-proto itself.
Where the records are
read_logs takes a source that names the records to read:
| Source | What it returns |
|---|---|
egress |
The tunnel's records for this application: one egress_establishments record per host, port, outcome, and declared flag per minute, and each refusal as its own record. |
harness |
The harness's records of connections from inside the application's process: one egress_establishments record per host, port, and outcome per minute, and each failed connection as its own record. |
router |
The serving router's records for this application, listed below. |
container |
The application's raw console, including the harness's window_ended, process_ended, and harness_refusal events. It is the one source read from the console rather than the logging service. |
all |
Every logging-service source (app, platform, router, egress, and harness) merged by time. It leaves out the console; read console lines with container. |
The router source contains:
- one
legs_endedrecord per kind of call (a request or a scheduled run) and outcome, per minute, a call your application responded to also split by itsstatus_class,1xxto5xx; - one
answered_5xxrecord per call your application responded to with a 5xx status, at most twenty per environment a minute on each router replica, the rest counted on the5xxrow; - one
window_endedrecord per request or scheduled run whose window ended; - one
upstream_endedrecord per request cut by a process end; and - one
source_refusedrecord per minute in which the per-address limit refused requests.
read_logs reads the production environment's records when no environment is given. The other two values are development, for the development environment, and local, the stream a local run writes. The local stream has no console, so a container read naming local is refused invalid_request. A development pod scaled to zero has no console lines, because the console exists only while the pod runs. The logging-service sources return records either way.
Finding hostnames to declare
filter: "declared" and filter: "undeclared" select recorded egress destinations against the current manifest. Use the undeclared list to find the hostnames to declare before platform staff switch the application to enforce mode.
- The filter reads the per-minute
egressandharnessrows; with no source, it reads both. - It leaves out platform endpoints, loopback targets, and the harness's records of connections that went through the proxy.
- It filters the records it retrieved and does not fetch older ones to fill the requested limit.
- An address literal in the
undeclaredresults is not a host to declare, because theegresslist accepts no address literal. Declare the hostname the address serves instead. - An address in a private, loopback, link-local, reserved, or platform range is refused by its range in enforce mode too.
A container query reads since and ignores until. None of these sources contains the gateway's per-call usage records. Node.js Runtime describes the harness and the request time limit, and the platform's deploy skill describes what a deployment must do.
Related
- Call an external API with an API key declares an upstream and calls it through the gateway or the tunnel.
- Read logs and counters shows how to read the records this page describes.
- Plan and usage reads each plan's outbound limits and the day's figures.
- Egress is the package page for the proxy and the gateway.
- Node.js Runtime describes the harness and the request time limit.
- The manifest describes the
egressmember and the other declarations. - Applications and environments explains the two managed hostnames.
- read_logs is the action's generated reference entry.