The deploy health check
The deploy health check is the platform's test of a new running copy before an environment switches to it: it requests the manifest's health path until the path returns status 200. A deploy, a promote, and a platform redeploy each run it, and a copy that fails it never serves.
How the check probes
The check sends GET to the manifest's health path on the new copy every three seconds, for up to 180 seconds. It passes only when the path returns status 200 exactly. The check follows no redirect, so the health path must return 200 itself. A 3xx response counts as a failed probe and never ends the check early, so a path that only redirects fails at the 180-second bound.
Each probe waits at most ten seconds for the response to begin. Near the end of the check it waits only for the time left, but never less than one second. So a health handler must begin its response within ten seconds, and a probe waiting longer is recorded as timed_out.
When the check ends early
A stable client error ends the check early. When twenty probes in a row return the same 4xx status, other than 408, 425, or 429, the platform checks whether your own process is responding. It sends one control probe, a request without the router mark, right after the twentieth. Where your process responds to it, the check ends early, about a minute in, and the record's outcome.gate contains ended_early: true.
Any other response during the twenty probes starts the count again, and a 200 passes the check. A 5xx response or a transport error, such as a failed or timed-out connection, never counts toward the twenty. A process that is still starting returns them. Nor does a 4xx from something in front of your process end the check early.
What the check expects of your server
The check expects a server to answer on its health path once it listens. It treats a 503 as starting, so a server may answer 503 until the application is ready, and it treats exactly 200 as ready. It reads twenty 4xx responses in a row as a stable client error, so a working health path gives no 4xx. A server with a database may wait to listen until its migrations finish, as the Database package describes, and until it listens, the check treats it as starting.
A route mounted after the server starts listening answers 404 until it is mounted, and twenty of those 404 responses end the check early.
Where the manifest declares the database kind, the deploy skill asks for a health route that performs one read over the connection and returns 503 while that read fails. The Database package's integration guide describes what the read proves: a pass shows the connection setting reaches the database, not only that the process responds. Build a database-backed service shows such a route.
A fixed body suits that 503, such as {"ok":false}, with no error text, credential, or host. Where the check fails, the deploy's record keeps the first 512 bytes of the last non-200 body your own process returned, and list_versions returns it too.
The evidence a failed check keeps
A failed deploy ends its record in the state failed, with the failure's name and detail in the record's outcome. A failed deploy or promote also gives the step it stopped at in step, such as image_build, compute_apply, or health_gate. Where the health check failed (health_gate_failed), the outcome also includes gate, the check's evidence, collected before the platform removes the new copy:
| Member | What it contains |
|---|---|
probe_sequence |
The probe responses as runs of an HTTP status or a connection word, the last sixteen runs. |
last |
The last non-200 response's status and content type, with the first 512 bytes of its body where your own process responded. Null where no probe got an HTTP response. |
answered_by |
application where your process responded to the control probe with the harness's router_mark_required refusal. intermediary where something in front of your process responded, nothing where nothing responded, and unknown otherwise. |
compute |
The new copy's state and restart count, or unavailable with a cause word where the platform could not read it. |
console |
The last forty lines of the new copy's console, at most 4,096 bytes, with truncated where the limit cut them. |
ended_early |
True where the check ended early on a stable client error from your own process, and false where it ran to its bound. |
Each control probe may leave one harness_refusal line in your console. For a development pod, answered_by is usually unknown, because the pod receives no request until it is ready. The record's detail sums up the evidence and gives no address. Where the evidence shows a likely cause, such as no route at the health path or a process that stops, the detail opens with it and its remedy. Where the last response was a redirect, the detail says so and that the check does not follow it. list_versions returns the same outcome.
For an environment running as a container, console lines reach the log store after a delay. The platform waits up to 60 seconds for a first line before recording the tail. Where none arrived, console is empty and says so. Where a version is running, the lines are then readable through read_logs. After a failed first deploy they are not: read_logs returns only the lines the check kept.
Related
- Deploy an application runs the check at each deploy, and its step 5 reads the record the check leaves.
- Deploys in Troubleshooting starts from what a failed check shows: a run of 404 responses, connection errors while the server starts, or a failed check with its evidence.
- Author the manifest says how to choose the health path.
- Build a database-backed service shows a health route that reads over the database connection.
- Read logs and counters explains
read_logs, which returns a running copy's console. - Add a database explains why a server waiting for its migrations does not listen yet, which the check reads as starting.