Git Push → Live Without Downtime: Auto-Deploy a Next.js App with Docker, a Signed Webhook and Blue/Green Slots
The PHP recipe breaks the moment a build step appears. This is the version for Next.js: build next to the running app, health-check it, switch with one rename() — and a broken commit never reaches your visitors. No CI platform, no docker.sock.
In my last post a push to main updated a PHP website in 1.4 seconds: fetch, export, switch a symlink. Then I tried the same idea on the Next.js front end that sits next to one of my Oracle APEX applications — and it fell apart, because a Next.js app has to be built before it can run. The build takes time, it needs RAM, and it can fail.
My first attempt did what most forum answers do: one node:22-alpine container that clones the repository, runs npm install and npm run build, then starts server.js — and a second container that restarts it through docker.sock when a webhook arrives. I rebuilt it in my lab and measured it: every deploy takes the site offline for the whole build, and one commit with a typo kept it offline until the next fix was pushed.
This post is the setup I use now. Who it is for: anyone who runs a Next.js (or any Node.js) app in Docker on their own server and wants Git-based deployments without a CI platform. What you have at the end: push → live in about 10 seconds, 0 failed requests during deploys, switches and rollbacks (measured with four clients hammering the site), a broken build that is reported in the log and never goes live, rollback in 0.3 seconds — and no container that can control your Docker daemon.
deploy.mjs, build.sh, slot.sh, hook.py, four small Dockerfiles, Caddyfile, docker-compose.yml and a measuring tool. Originally developed by S&H Software Solutions – first published on APEX Note (October 2026).TL;DR — what the setup does:
- Builds next to the running version — a separate
buildercontainer runsnpm ciandnext build; the live app is not touched until the new build has passed - Two app slots, blue and green — the new release starts in the idle slot and must answer
/api/healthwith the right commit andGET /with200before it gets any traffic - Switches with one
rename()— Caddy reads the name of the live slot from a file on every request; no reload, no admin API, no dropped connection - Broken build = log entry, not an outage — the failing step and the compiler error are in the log, the old version keeps serving
- Instant rollback — the previous release keeps running in the other slot; switching back is one file write
- Signed webhooks only — the listener from the PHP post, unchanged: HMAC-SHA256, constant-time compare, doorbell pattern
- No
docker.sock, read-only containers, dropped capabilities, the app on an internal network, the build without any secret
My First Attempt — and What It Did to the Site
Here is the draft, with the token and the IP address replaced. It is short, it is readable, and it "works":
# My first attempt (sanitised: token and IP replaced) - do NOT use this
services:
my_store_locator_app:
image: node:22-alpine
container_name: my_store_locator_app
working_dir: /app
command: >
sh -c "apk add --no-cache git &&
if [ -d /app/.git ]; then
git -C /app fetch origin &&
git -C /app reset --hard origin/main &&
git -C /app clean -fd -e node_modules -e .next;
else
git clone https://<GITHUB_TOKEN>@github.com/<owner>/my-store-locator.git /app;
fi &&
npm install && npm run build &&
cp -r /app/public /app/.next/standalone/public 2>/dev/null;
cp -r /app/.next/static /app/.next/standalone/.next/static &&
HOSTNAME=0.0.0.0 PORT=3000 node /app/.next/standalone/server.js"
ports:
- "4000:3000"
volumes:
- my_store_locator_data:/app
restart: always
my_store_locator_hook:
image: docker:cli
container_name: my_store_locator_hook
command: >
sh -c "while true; do
printf 'HTTP/1.1 200 OK\r\nContent-Length: 2\r\n\r\nok' | nc -l -p 9001 > /dev/null || { sleep 5; continue; };
echo 'Webhook received - restarting the app';
docker restart my_store_locator_app;
done"
ports:
- "9001:9001"
volumes:
- /var/run/docker.sock:/var/run/docker.sock
restart: always
volumes:
my_store_locator_data:The problems are not visible in the YAML, so I rebuilt it in the lab and put load on it. Four parallel clients request the page in a loop while I run exactly what the hook container runs — docker restart:
docker restart — what its hook does on every push — costs 276 failed requests and five seconds without any answer, after ten seconds of waiting for a process that ignores SIGTERM.276 failed requests and five seconds without an answer for an app that builds in three seconds. Your real app builds longer, and every second of next build is a second of downtime. Two more things hide in the numbers: the failures only start ten seconds after the restart, because sh -c as PID 1 ignores SIGTERM and Docker waits ten seconds before it kills the container — so in-flight requests are never finished gracefully. And the restart re-runs npm install against whatever the volume holds.
Then I pushed a commit with a JSX typo and restarted again. The build failed, && never reached node server.js, the container exited and restart: always started the next round:
$ git log --oneline -1; git show HEAD -U0 --format= app/page.js | tail -2
7b99a41 (HEAD -> main, origin/main) Version 8: broken JSX again
+ ["Version", VERSION],,
+ ["Broken", <b>oops</i>],
$ docker restart blog3-draft-app >/dev/null; node ~/next-app/tools/watch-downtime.mjs http://127.0.0.1:14000/ 45 4 | tail -5
requests: 3282 ok: 0 failed: 3282
ECONNRESET: 1580
ECONNREFUSED: 1702
site down from +0.0s to +45.0s
longest time without a good answer: 45046 ms
$ docker inspect blog3-draft-app --format "status={{.State.Status}} exit={{.State.ExitCode}} restarts={{.RestartCount}}"; docker logs blog3-draft-app 2>&1 | grep -m1 -A3 "Build error"
status=restarting exit=1 restarts=8
> Build error occurred
Error: Turbopack build failed with 1 error:
./app/page.js:14:25
Error: Expected corresponding JSX closing tag for <b>
$ node ~/next-app/tools/watch-downtime.mjs https://www.example.com/ 10 4 | tail -4 # the blue/green stack, same broken commit
---
requests: 1909 ok: 1909 failed: 0
Version 7 (slot blue): 1909
longest time without a good answer: 70 msERR_CONNECTION_RESET — for as long as the bug is in main.The site stayed down until I pushed a fix — and even then it took 41 seconds and 10 restarts until Docker's back-off let the loop pick it up. The same broken commit hit the setup from this post at the same moment: 0 failed requests, the old version kept serving (last line above). That difference is what this post is about.
| Problem in the draft | Why it hurts | Fix in this post |
|---|---|---|
| Build and run in the same container | The old version must stop before the new one is built → downtime = build time | Separate builder, two app slots, switch after the build |
| No check before going live | A failing build or a crashing start takes the site down | Build result + health check gate every switch |
docker.sock in the hook container | Root on the host for anyone who breaks the internet-facing listener | Doorbell file; nobody talks to the Docker daemon |
| Unsigned trigger on port 9001 | Anyone can restart (= take down) the site | HMAC-SHA256 signature, only through Caddy |
| Token in the clone URL | Stored in /app/.git/config on the volume | Read-only deploy key as a Compose secret |
npm install in production | May change the lock file, runs package scripts next to the token | npm ci in a container without secrets |
The Design: Build Beside, Check, Switch
Six containers, each with exactly the access it needs. The numbers in the diagram are the order of one deployment:
| Container | Holds | Can reach | Cannot |
|---|---|---|---|
| proxy (Caddy) | TLS certificates | hook, both slots | change anything — it only reads control/live |
| hook (Python) | webhook secret | nothing but its doorbell file | run git, see the app, start anything |
| deployer (Node.js) | read-only deploy key | GitHub, the slots (health check) | run npm; it has no published port |
| builder (Node.js) | nothing secret | the npm registry | reach the app, the proxy or the key |
| app-blue / app-green | runtime variables (app.env) | internal network only | reach the internet, change a release |
node:22-alpine (Node.js 22.23.3, npm 10.9.9), Next.js 16.4.0 (Turbopack) with React 19.3.0, Caddy 2.11.7 and python:3.12-alpine. A local Gitea 1.24 plays GitHub: it hosts apexnote/demo-app, holds the read-only deploy key and sends the webhook with the same X-GitHub-Event and X-Hub-Signature-256 headers. Two lab-only changes, both in test/: the Alpine package mirror was blocked, so git and ssh in the deployer image were copied from alpine/git instead of apk add; and the lab's egress proxy re-signs TLS, so the builder got its CA via NODE_EXTRA_CA_CERTS. Not tested: GitHub's own UI and Let's Encrypt (described from their documentation).Step-by-Step: From Repository to Zero-Downtime Deploys
Project Layout on the Server
The deployment stack is its own small project on the Docker host (I use ~/next-app); your Next.js repository stays a normal Next.js project.
~/next-app
├── .env # Step 1
├── docker-compose.yml # Step 7
├── Caddyfile # Step 7
├── app.env / build.env # Step 7 (optional, from the .example files)
├── known_hosts # Step 3
├── secrets/ # never commit this folder
│ ├── deploy_key # Step 3
│ └── webhook_secret # Step 3
├── builder/ Dockerfile build.sh # Step 4
├── app-slot/ Dockerfile slot.sh # Step 5
├── deployer/ Dockerfile deploy.mjs # Step 6
├── hook/ Dockerfile hook.py # Step 7 (from the PHP post)
└── tools/ watch-downtime.mjs send-test-webhook.sh# copy to .env and adjust
SITE_DOMAIN=www.example.com
REPO_URL=git@github.com:your-name/your-next-app.git
REPO_FULL_NAME=your-name/your-next-app
DEPLOY_BRANCH=mainPrepare the Next.js App: Standalone Output and a Health Route
Two changes in your app repository. output: "standalone" makes next build write .next/standalone/server.js plus only the node_modules files the server really needs — 67 MB for the demo instead of 352 MB for a full install. And a health route that tells the deployer which build is answering, so it can never switch to a slot that still runs the old version. Here is the complete APEX Note Demo App I used for every test:
/** @type {import('next').NextConfig} */ const nextConfig = { output: "standalone", // .next/standalone/server.js + only the node_modules it needs poweredByHeader: false, }; export default nextConfig;
// Health check for the deployer: is THIS process up, and which build is it? export const dynamic = "force-dynamic"; export function GET() { return Response.json({ status: "ok", commit: process.env.NEXT_PUBLIC_COMMIT || "dev", slot: process.env.APP_SLOT || "-", uptime: Math.round(process.uptime()), }); }
import { VERSION, INTRO } from "./version"; // Rendered per request: APP_SLOT is read at RUNTIME, NEXT_PUBLIC_* was fixed at BUILD time export const dynamic = "force-dynamic"; export default function Home() { const commit = process.env.NEXT_PUBLIC_COMMIT || "dev"; const subject = process.env.NEXT_PUBLIC_COMMIT_SUBJECT || "local build"; const builtAt = process.env.NEXT_PUBLIC_BUILT_AT || "-"; const slot = process.env.APP_SLOT || "-"; const rows = [ ["Status", <span key="s" className="ok">✓ live</span>], ["Version", VERSION], ["Commit", `${commit} · ${subject}`], ["Built at", builtAt], ["Served by", `slot ${slot}`], ["Rendered", new Date().toISOString().replace("T", " ").slice(0, 19) + " UTC"], ["Runtime", `Node.js ${process.versions.node} · Next.js standalone`], ]; return ( <main> <div className="card"> <div className="stripe"><i /><i /><i /></div> <header> <span className="brand"><b>APEX</b> Note · Lab</span> <span className="badge">Version {VERSION}</span> </header> <section> <h1>APEX Note Demo App</h1> <p className="intro">{INTRO}</p> <table> <tbody> {rows.map(([k, v]) => ( <tr key={k}><th>{k}</th><td>{v}</td></tr> ))} </tbody> </table> </section> </div> </main> ); }
export const VERSION = "1"; export const INTRO = "Version 1 — built from main and switched in without downtime.";
{ "name": "apexnote-demo-app", "version": "1.0.0", "private": true, "scripts": { "dev": "next dev", "build": "next build", "start": "node .next/standalone/server.js" }, "dependencies": { "next": "16.4.0", "react": "19.3.0", "react-dom": "19.3.0" }, "engines": { "node": ">=20.9" } }
Note the two kinds of variables in page.js. NEXT_PUBLIC_COMMIT is set by the builder and inlined at build time — it is part of the JavaScript. APP_SLOT comes from the container and is read at runtime, because the page is rendered per request (dynamic = "force-dynamic"). Mixing these up is the most common Next.js deployment bug; How it works inside shows the proof.
Deploy Key, Host Key, Webhook Secret
This part is identical to the PHP post, where every line is explained. The deploy key is read by the deployer (UID 1000, the node user), the secret by the hook (UID 1001):
mkdir -p secrets
ssh-keygen -t ed25519 -N "" -C "deploy@www.example.com" -f secrets/deploy_key
chown 1000:1000 secrets/deploy_key && chmod 400 secrets/deploy_key # read by the deployer (UID 1000)
cat secrets/deploy_key.pub # -> GitHub, read-only deploy key
ssh-keyscan -t ed25519 github.com > known_hosts
ssh-keygen -lf known_hosts # compare with GitHub's published Ed25519 fingerprint
openssl rand -hex 32 > secrets/webhook_secret
chown 1001:1001 secrets/webhook_secret && chmod 400 secrets/webhook_secret # read by the hook (UID 1001)
cat secrets/webhook_secret # paste into the GitHub webhook, then clear the terminalIn GitHub: Settings → Deploy keys → paste deploy_key.pub, leave "Allow write access" unchecked. Settings → Webhooks → Payload URL https://www.example.com/_hooks/github, content type application/json, the secret, Just the push event.
The Builder: npm ci + next build, Isolated
The builder watches a request file. For each request it copies the exported source into its workspace, runs npm ci (skipped when package.json, package-lock.json and the Node.js version are unchanged), runs next build and assembles an immutable release in /srv/releases/<commit>: the standalone server, .next/static and public. The release is built in a temporary folder and renamed when complete — a slot can never start half a release. Every step is logged; on failure the last 20 lines go to the container log:
#!/bin/sh # build.sh - the ONLY container that runs npm. It gets source code, never secrets. # # build.sh watch /srv/source/request (written by the deployer) and build # build.sh once build the current request once and exit # # For every request it runs "npm ci" and "next build" in /build/app and turns # the standalone output into an immutable release /srv/releases/<commit>/. # The result goes to /srv/releases/.results/<commit>.json, the full log to # /srv/releases/.logs/<commit>.log. A broken build never touches a release # that is running. # # Originally developed by S&H Software Solutions - first published on # APEX Note (www.apexnote.de), October 2026. MIT licence. set -u umask 022 SOURCE=/srv/source RELEASES=/srv/releases CONTROL=/srv/control WORK=/build/app CACHE=/cache KEEP=${KEEP_RELEASES:-5} SELF=$(readlink -f "$0") log() { printf '%s build %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$*"; } field() { sed -n "s/^$1=//p" "$SOURCE/request" | head -n 1; } build_once() { id=$(field id); sha=$(field sha); short=$(field short); subject=$(field subject) case "$sha" in *[!0-9a-f]* | "") log "ignoring invalid request"; return 0 ;; esac [ ${#sha} -eq 40 ] || { log "ignoring invalid request"; return 0; } mkdir -p "$RELEASES/.results" "$RELEASES/.logs" || return 1 if grep -qs "\"id\":\"$id\"" "$RELEASES/.results/$sha.json"; then return 0 # answered before this container restarted fi logf=$RELEASES/.logs/$sha.log started=$(date +%s) took="" result() { # status step - written atomically printf '{"id":"%s","sha":"%s","status":"%s","step":"%s","seconds":%s}\n' \ "$id" "$sha" "$1" "$2" "$(( $(date +%s) - started ))" > "$RELEASES/.results/.$sha.tmp" mv -f "$RELEASES/.results/.$sha.tmp" "$RELEASES/.results/$sha.json" } step() { # step NAME command... - output into the log name=$1; shift printf '\n### %s\n' "$name" >> "$logf" s0=$(date +%s) if "$@" >> "$logf" 2>&1; then case "$name" in npm*|next*) took="$took, $name $(( $(date +%s) - s0 ))s" ;; esac else log "FAILED in step \"$name\" after $(( $(date +%s) - started ))s - last lines of $logf:" tail -n 20 "$logf" | sed 's/^/ | /' result failed "$name" return 1 fi } if [ -f "$RELEASES/$sha/server.js" ]; then log "release $short already built - reusing it" result ok cached return 0 fi : > "$logf" log "building $short \"$subject\"" # node_modules is reused only if package.json, package-lock.json and Node.js are unchanged deps=$( (cat "$SOURCE/$sha/package.json" "$SOURCE/$sha/package-lock.json"; node -v) 2>/dev/null | sha256sum | cut -c1-16) keep="" if [ "$(cat /build/deps 2>/dev/null)" = "$deps" ] && [ -d "$WORK/node_modules" ]; then rm -rf /build/node_modules.keep && mv "$WORK/node_modules" /build/node_modules.keep && keep=1 fi step "copy source" sh -c 'rm -rf "$1" && cp -a "$2" "$1" && mkdir -p "$1/.next" "$3/next" && ln -s "$3/next" "$1/.next/cache"' - "$WORK" "$SOURCE/$sha" "$CACHE" || return 1 cd "$WORK" || return 1 if [ -n "$keep" ]; then mv /build/node_modules.keep node_modules printf '\n### npm ci skipped - dependencies unchanged (%s)\n' "$deps" >> "$logf" took="$took, npm ci skipped" else rm -f /build/deps step "npm ci" npm ci --no-audit --no-fund --cache "$CACHE/npm" --prefer-offline || return 1 echo "$deps" > /build/deps fi # NEXT_PUBLIC_* values are inlined into the bundle NOW - they cannot change at runtime step "next build" env NEXT_PUBLIC_COMMIT="$short" NEXT_PUBLIC_COMMIT_SUBJECT="$subject" \ NEXT_PUBLIC_BUILT_AT="$(date -u '+%Y-%m-%d %H:%M:%S UTC')" npm run build || return 1 step "check output" test -f .next/standalone/server.js || return 1 tmp=$RELEASES/.tmp-$sha step "assemble release" sh -c 'rm -rf "$1" && cp -a .next/standalone "$1" && rm -rf "$1/.next/cache" && cp -a .next/static "$1/.next/static" && if [ -d public ]; then cp -a public "$1/public"; fi' - "$tmp" || return 1 mv "$tmp" "$RELEASES/$sha" # appears complete or not at all result ok built log "built $short in $(( $(date +%s) - started ))s (${took#, }, release $(du -sh "$RELEASES/$sha" | cut -f1))" # keep the newest $KEEP releases - never one that a slot is using in_use=$(cat "$CONTROL"/slots/* 2>/dev/null || true) ls -1t "$RELEASES" | tail -n +"$((KEEP + 1))" | while read -r old; do case " $in_use " in *"$old"*) continue ;; esac rm -rf "${RELEASES:?}/$old" "$RELEASES/.logs/$old.log" "$RELEASES/.results/$old.json" done } case "${1:-watch}" in once) build_once ;; watch) rm -rf "$RELEASES"/.tmp-* 2>/dev/null || true log "waiting for build requests in $SOURCE/request" last="" while :; do req=$(cat "$SOURCE/request" 2>/dev/null || true) if [ -n "$req" ] && [ "$req" != "$last" ]; then last=$req "$SELF" once || true # a child process: one failed build never stops the loop fi sleep 1 done ;; *) echo "usage: build.sh [watch|once]" >&2; exit 2 ;; esac
# Runs npm ci + next build. Has internet access for npm - and nothing worth stealing: # no deploy key, no webhook secret, no admin socket, no route to the running app. FROM node:22-alpine RUN mkdir -p /srv/source /srv/releases /srv/control /build /cache \ && chown node:node /srv/source /srv/releases /srv/control /build /cache COPY --chmod=755 build.sh /usr/local/bin/build.sh USER node ENV HOME=/tmp NEXT_TELEMETRY_DISABLED=1 NPM_CONFIG_UPDATE_NOTIFIER=false \ NODE_OPTIONS=--max-old-space-size=1536 ENTRYPOINT ["build.sh"]
The App Slots: Blue and Green
Without docker.sock nobody can restart a container — so the containers restart themselves. Each slot runs a tiny supervisor that reads its slot file (/srv/control/slots/blue) and keeps node server.js of that release running. When the file changes it stops the old process (SIGTERM, after 10 seconds SIGKILL) and starts the new release; when node crashes it starts it again:
#!/bin/sh # slot.sh - keeps "node server.js" of ONE release running in this slot (blue or green). # # The deployer chooses the release by writing its commit into # /srv/control/slots/<slot>. When the file changes, the old process gets # SIGTERM and the new release starts. If node crashes, it is started again. # No docker.sock needed: the deployer never restarts containers. # # Originally developed by S&H Software Solutions - first published on # APEX Note (www.apexnote.de), October 2026. MIT licence. set -u SLOT=${APP_SLOT:?set APP_SLOT to blue or green} WANT_FILE=/srv/control/slots/$SLOT RELEASES=/srv/releases pid="" running="" log() { printf '%s slot-%s %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$SLOT" "$*"; } stop_app() { # SIGTERM, then SIGKILL after 10 s if [ -n "$pid" ]; then kill -TERM "$pid" 2>/dev/null i=0 while kill -0 "$pid" 2>/dev/null && [ $i -lt 10 ]; do sleep 1; i=$((i + 1)); done kill -KILL "$pid" 2>/dev/null wait "$pid" 2>/dev/null pid="" fi } trap 'stop_app; exit 0' TERM INT log "waiting for $WANT_FILE" while :; do want=$(cat "$WANT_FILE" 2>/dev/null || true) crashed="" if [ -n "$pid" ] && ! kill -0 "$pid" 2>/dev/null; then wait "$pid" 2>/dev/null; crashed="exit $?"; pid="" fi if [ "$want" != "$running" ] || [ -n "$crashed" ]; then [ -n "$crashed" ] && log "node stopped ($crashed) - starting it again" stop_app if [ -n "$want" ] && [ -f "$RELEASES/$want/server.js" ]; then log "starting release ${want%"${want#???????}"}" (cd "$RELEASES/$want" && PORT=3000 HOSTNAME=0.0.0.0 exec node server.js) & pid=$! elif [ -n "$want" ]; then log "ERROR: release $want not found" fi running=$want [ -n "$crashed" ] && sleep 2 # do not spin on a crash loop fi sleep 1 done
# Runs ONE Next.js release (standalone server.js) - no npm, no git, no build tools needed FROM node:22-alpine # same owner as in the builder image - whichever container creates a volume first RUN mkdir -p /srv/releases /srv/control && chown node:node /srv/releases /srv/control COPY --chmod=755 slot.sh /usr/local/bin/slot.sh USER node ENV NODE_ENV=production NEXT_TELEMETRY_DISABLED=1 EXPOSE 3000 ENTRYPOINT ["slot.sh"]
The Deployer: Fetch, Build, Check, Switch
The deployer is the only part that makes decisions. It fetches main into a bare clone, exports the commit with git archive (the builder never sees .git or the key), rings the builder's doorbell and waits for its answer. If the build is good, it writes the commit into the idle slot's file, waits until that slot reports the new commit on /api/health and answers GET / with 200 — and only then writes the slot's name into control/live. Read deploy() from top to bottom; every return before pointProxyTo() leaves the live version alone:
#!/usr/bin/env node // deploy.mjs - zero-downtime deploys for a Next.js app (Node.js standard library only). // // deploy.mjs watch loop: deploy at start, on every doorbell from hook.py // and every FALLBACK_INTERVAL seconds // deploy.mjs once deploy once and exit // deploy.mjs status show live slot, standby slot and the kept releases // deploy.mjs rollback switch back to the standby slot (previous release) and pin it // deploy.mjs unpin remove the pin and deploy the branch again // // One deploy = fetch -> export source -> builder runs npm ci + next build -> // start the release in the IDLE slot -> health check -> switch the proxy file. // If any step fails, the live slot is never touched. // // Originally developed by S&H Software Solutions - first published on // APEX Note (www.apexnote.de), October 2026. MIT licence. import { execFileSync } from "node:child_process"; import fs from "node:fs"; const env = process.env; const REPO_URL = env.REPO_URL || fail("set REPO_URL, e.g. git@github.com:owner/app.git"); const BRANCH = env.DEPLOY_BRANCH || "main"; const FALLBACK = Number(env.FALLBACK_INTERVAL || 300); // seconds const BUILD_TIMEOUT = Number(env.BUILD_TIMEOUT || 900); // seconds const HEALTH_TIMEOUT = Number(env.HEALTH_TIMEOUT || 60); // seconds const SIGNAL = env.REQUEST_FILE || "/signal/request"; const KEY = env.DEPLOY_KEY_FILE || "/run/secrets/deploy_key"; const KNOWN_HOSTS = env.KNOWN_HOSTS_FILE || "/etc/deploy/known_hosts"; const MIRROR = "/srv/git/repo.git"; // bare clone, never visible to anybody else const SOURCE = "/srv/source"; // exported source for the builder (read-only there) const RELEASES = "/srv/releases"; // written by the builder, read-only here const CONTROL = "/srv/control"; // live + slots/<slot> + state.json, read-only for proxy and slots const OTHER = { blue: "green", green: "blue" }; process.env.GIT_TERMINAL_PROMPT = "0"; process.env.GIT_SSH_COMMAND = `ssh -i ${KEY} -o IdentitiesOnly=yes -o BatchMode=yes ` + `-o StrictHostKeyChecking=yes -o UserKnownHostsFile=${KNOWN_HOSTS}`; function fail(msg) { console.error(`deploy: ${msg}`); process.exit(2); } const log = (...m) => console.log(new Date().toISOString().slice(0, 19) + "Z deploy", ...m); const sleep = (ms) => new Promise((r) => setTimeout(r, ms)); const git = (...args) => execFileSync("git", args, { encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] }).trim(); function writeAtomic(path, text) { // readers never see half a file fs.writeFileSync(`${path}.tmp`, text); fs.renameSync(`${path}.tmp`, path); } const readState = () => { try { return JSON.parse(fs.readFileSync(`${CONTROL}/state.json`, "utf8")); } catch { return {}; } }; const saveState = (s) => writeAtomic(`${CONTROL}/state.json`, JSON.stringify(s, null, 1) + "\n"); // Short exclusive lock for "start slot + switch" - so a rollback never races a deploy async function withLock(fn) { for (let i = 0; ; i++) { try { fs.mkdirSync(`${CONTROL}/.lock`); break; } catch (e) { if (e.code !== "EEXIST") throw e; if (i === 0) log("waiting for a running switch to finish"); await sleep(250); } } try { return await fn(); } finally { fs.rmSync(`${CONTROL}/.lock`, { recursive: true, force: true }); } } // ---------------------------------------------------------------- the switch // Caddy reads /srv/control/live on EVERY request ({file./srv/control/live}:3000). // Switching = one rename(): no reload, no admin API, no connection is dropped. function pointProxyTo(slot) { writeAtomic(`${CONTROL}/live`, `app-${slot}\n`); } // ---------------------------------------------------------------- health check of one slot async function healthy(slot, short, timeoutS = HEALTH_TIMEOUT) { const until = Date.now() + timeoutS * 1000; let last = "no answer"; while (Date.now() < until) { try { const r = await fetch(`http://app-${slot}:3000/api/health`, { signal: AbortSignal.timeout(2000) }); const h = await r.json(); if (r.ok && h.commit === short) { const page = await fetch(`http://app-${slot}:3000/`, { signal: AbortSignal.timeout(5000) }); if (page.ok) return true; last = `GET / answered ${page.status}`; } else last = `health: ${r.status} commit ${h.commit}`; } catch (e) { last = e.cause?.code || e.name; } await sleep(500); } log(`slot ${slot} not healthy after ${timeoutS}s (${last})`); return false; } // ---------------------------------------------------------------- steps function exportSource(sha) { // git archive: the builder never sees .git or the key if (fs.existsSync(`${SOURCE}/${sha}`)) return; const tmp = `${SOURCE}/.tmp-${sha}`; fs.rmSync(tmp, { recursive: true, force: true }); fs.mkdirSync(tmp); git("-C", MIRROR, "archive", "--format=tar", `--output=${tmp}.tar`, sha); execFileSync("tar", ["-x", "-f", `${tmp}.tar`, "-C", tmp]); fs.rmSync(`${tmp}.tar`); fs.renameSync(tmp, `${SOURCE}/${sha}`); for (const old of fs.readdirSync(SOURCE).filter((d) => /^[0-9a-f]{40}$/.test(d) && d !== sha)) { fs.rmSync(`${SOURCE}/${old}`, { recursive: true, force: true }); // keep only the newest export } } async function build(sha, short, subject) { // ring the builder's doorbell and wait for its answer const id = `${Date.now()}-${process.pid}`; writeAtomic(`${SOURCE}/request`, `id=${id}\nsha=${sha}\nshort=${short}\nsubject=${subject.replace(/[\r\n]/g, " ")}\n`); const until = Date.now() + BUILD_TIMEOUT * 1000; while (Date.now() < until) { try { const r = JSON.parse(fs.readFileSync(`${RELEASES}/.results/${sha}.json`, "utf8")); if (r.id === id) return r; } catch { /* not there yet */ } await sleep(1000); } return { status: "failed", step: `timeout after ${BUILD_TIMEOUT}s`, seconds: BUILD_TIMEOUT }; } async function deploy() { let state = readState(); if (state.pinned) { log(`pinned to ${state.live.short} - not deploying (run: deploy.mjs unpin)`); return true; } const t0 = Date.now(); if (!fs.existsSync(MIRROR)) { log(`first run: cloning branch ${BRANCH}`); fs.rmSync(`${MIRROR}.tmp`, { recursive: true, force: true }); git("clone", "--quiet", "--bare", "--single-branch", "--branch", BRANCH, REPO_URL, `${MIRROR}.tmp`); fs.renameSync(`${MIRROR}.tmp`, MIRROR); } git("-C", MIRROR, "fetch", "--quiet", "--prune", "origin", `+refs/heads/${BRANCH}:refs/heads/${BRANCH}`); const sha = git("-C", MIRROR, "rev-parse", "--verify", `refs/heads/${BRANCH}^{commit}`); const short = sha.slice(0, 7); const subject = git("-C", MIRROR, "log", "-1", "--format=%s", sha); fs.writeFileSync(`${CONTROL}/last_ok`, String(Date.now())); // read by the health check if (state.live?.sha === sha) { log(`already live: ${short} on ${state.live.slot}`); return true; } if (state.failed === sha) { log(`${short} failed before - waiting for a new commit`); return true; } log(`new commit ${short} "${subject}" - building`); exportSource(sha); const res = await build(sha, short, subject); const live = state.live ? `${state.live.short} on ${state.live.slot}` : "nothing"; if (res.status !== "ok") { state = readState(); state.failed = sha; saveState(state); log(`ERROR: build of ${short} failed in step "${res.step}" after ${res.seconds}s - ${live} stays live`); log(` full log: docker compose exec builder cat ${RELEASES}/.logs/${sha}.log`); return false; } log(`build ok in ${res.seconds}s (${res.step})`); return withLock(async () => { state = readState(); if (state.pinned) { log("pinned meanwhile - not switching"); return true; } const slot = state.live ? OTHER[state.live.slot] : "blue"; writeAtomic(`${CONTROL}/slots/${slot}`, sha); // slot.sh starts the release log(`starting ${short} in slot ${slot}, waiting for /api/health`); if (!(await healthy(slot, short))) { state.failed = sha; saveState(state); log(`ERROR: ${short} did not become healthy in slot ${slot} - ${live} stays live`); if (state.standby?.slot === slot) { // keep the rollback target alive writeAtomic(`${CONTROL}/slots/${slot}`, state.standby.sha); log(`slot ${slot} goes back to ${state.standby.short} (standby for rollback)`); } return false; } pointProxyTo(slot); state.standby = state.live || null; state.live = { slot, sha, short, subject, at: new Date().toISOString() }; delete state.failed; saveState(state); log(`live: ${short} "${subject}" on ${slot} in ${Math.round((Date.now() - t0) / 1000)}s ` + `(build ${res.seconds}s)`); return true; }); } async function rollback() { return withLock(async () => { const state = readState(); const prev = state.standby; if (!prev) { log("no standby release to roll back to"); return false; } if (!(await healthy(prev.slot, prev.short, 10))) { log(`standby slot ${prev.slot} is not healthy`); return false; } pointProxyTo(prev.slot); [state.live, state.standby] = [prev, state.live]; state.pinned = true; saveState(state); log(`rolled back to ${prev.short} on ${prev.slot} and pinned - run 'deploy.mjs unpin' to deploy ${BRANCH} again`); return true; }); } function status() { const s = readState(); const line = (k, v) => console.log(`${k.padEnd(9)}${v ? `${v.short} slot ${v.slot} ${v.at} ${v.subject}` : "-"}`); line("live:", s.live); line("standby:", s.standby); if (s.pinned) console.log("pinned: yes (deploy.mjs unpin)"); if (s.failed) console.log(`failed: ${s.failed.slice(0, 7)} (build or health check failed)`); const rel = fs.readdirSync(RELEASES).filter((d) => /^[0-9a-f]{40}$/.test(d)); console.log(`releases: ${rel.map((d) => d.slice(0, 7)).join(" ") || "-"}`); } async function watch() { fs.mkdirSync(`${CONTROL}/slots`, { recursive: true }); fs.rmSync(`${CONTROL}/.lock`, { recursive: true, force: true }); // left over from a crash log(`watching ${SIGNAL} - branch ${BRANCH}, fallback sync every ${FALLBACK}s`); let last = null, next = 0; for (;;) { let req = ""; try { req = fs.readFileSync(SIGNAL, "utf8"); } catch { /* no doorbell yet */ } if (req !== last || Date.now() >= next) { if (last !== null && req !== last) log(`doorbell: ${req.trim()}`); last = req; next = Date.now() + FALLBACK * 1000; try { await deploy(); } catch (e) { log(`ERROR: ${String(e.stderr || e.message).trim()} - live slot unchanged, retry in 30s`); next = Date.now() + 30000; } } await sleep(1000); } } const cmd = process.argv[2] || "watch"; const run = { watch, once: deploy, rollback, status, unpin: async () => { const s = readState(); delete s.pinned; delete s.failed; saveState(s); return deploy(); } }[cmd]; if (!run) fail("usage: deploy.mjs [watch|once|status|rollback|unpin]"); Promise.resolve(run()).then((ok) => process.exit(ok === false ? 1 : 0), (e) => { log(`ERROR: ${String(e.stderr || e.message).trim()}`); process.exit(1); });
# Has the deploy key and decides which slot is live - but no published port and never runs npm FROM node:22-alpine RUN apk add --no-cache git openssh-client \ && mkdir -p /srv/git /srv/source /srv/control/slots /srv/releases /etc/deploy \ && chown -R node:node /srv/git /srv/source /srv/control /srv/releases \ && mkdir /signal && chown 1001:1001 /signal # owner = hook (UID 1001), whoever mounts it first COPY --chmod=755 deploy.mjs /usr/local/bin/deploy.mjs USER node ENV HOME=/tmp ENTRYPOINT ["deploy.mjs"]
Caddy, docker-compose.yml — and Start It
The interesting line in the Caddyfile is reverse_proxy {file./srv/control/live}:3000: Caddy's {file.*} placeholder reads the file on every request, so the upstream changes the moment the deployer renames a new live file into place. The lb_try_duration lines make Caddy retry the dial for up to five seconds instead of answering 502 if node is restarting. The hook is the one from the PHP post, unchanged:
# Zero-downtime Next.js deploys: proxy / hook / deployer / builder / app-blue + app-green name: nextapp x-hardening: &hardening init: true restart: unless-stopped read_only: true tmpfs: [/tmp] cap_drop: [ALL] security_opt: ["no-new-privileges:true"] x-slot: &slot <<: *hardening build: ./app-slot volumes: - releases:/srv/releases:ro # runs a release, can never change one - control:/srv/control:ro env_file: # RUNTIME variables (DATABASE_URL, API keys ...) - read per request - path: ./app.env required: false mem_limit: 512m healthcheck: test: ["CMD-SHELL", "wget -q -O /dev/null http://127.0.0.1:3000/api/health || exit 1"] interval: 30s timeout: 3s networks: [app] # internal network: no internet, no hook, no builder services: proxy: image: caddy:2.11-alpine restart: unless-stopped ports: - "80:80" - "443:443" - "443:443/udp" environment: SITE_DOMAIN: ${SITE_DOMAIN:?set SITE_DOMAIN in .env} volumes: - ./Caddyfile:/etc/caddy/Caddyfile:ro - caddy_data:/data - caddy_config:/config - control:/srv/control:ro # "live" = which slot gets the traffic depends_on: [deployer] # the deployer image creates /srv/control with the right owner networks: [front, app] hook: # unchanged from the PHP post <<: *hardening build: ./hook environment: HOOK_PATH: /_hooks/github DEPLOY_BRANCH: ${DEPLOY_BRANCH:-main} REPO_FULL_NAME: ${REPO_FULL_NAME:-} secrets: [webhook_secret] volumes: - signal:/signal # the only thing it can write to healthcheck: test: ["CMD", "python3", "-c", "import urllib.request as u; u.urlopen('http://127.0.0.1:9000/healthz', timeout=2)"] interval: 30s timeout: 3s networks: [front] deployer: <<: *hardening build: ./deployer environment: REPO_URL: ${REPO_URL:?set REPO_URL in .env} DEPLOY_BRANCH: ${DEPLOY_BRANCH:-main} FALLBACK_INTERVAL: "300" BUILD_TIMEOUT: "900" HEALTH_TIMEOUT: "60" secrets: [deploy_key] volumes: - git:/srv/git - source:/srv/source - releases:/srv/releases:ro - control:/srv/control - signal:/signal:ro - ./known_hosts:/etc/deploy/known_hosts:ro healthcheck: test: ["CMD-SHELL", "find /srv/control/last_ok -mmin -15 | grep -q ."] interval: 60s timeout: 5s networks: [app, git] builder: <<: *hardening build: ./builder environment: KEEP_RELEASES: "5" env_file: # BUILD-time variables: NEXT_PUBLIC_* end up in the browser bundle - path: ./build.env required: false volumes: - source:/srv/source:ro - releases:/srv/releases - control:/srv/control:ro # to never prune a release a slot is using - build:/build - build_cache:/cache # npm cache + .next/cache survive between builds mem_limit: 2g # next build needs RAM - see "memory" in the post networks: [build] app-blue: <<: *slot environment: APP_SLOT: blue app-green: <<: *slot environment: APP_SLOT: green secrets: webhook_secret: file: ./secrets/webhook_secret deploy_key: file: ./secrets/deploy_key volumes: git: source: releases: control: signal: build: build_cache: caddy_data: caddy_config: networks: front: # proxy + hook app: # proxy + deployer + slots, no internet internal: true git: # deployer -> GitHub build: # builder -> npm registry
{$SITE_DOMAIN} { encode zstd gzip # Only POST to this one path reaches the webhook listener @hook { method POST path /_hooks/github } handle @hook { request_body { max_size 5MB } reverse_proxy hook:9000 } # Blue/green switch: the deployer writes "app-blue" or "app-green" into this # file with one atomic rename. Caddy reads it on every request - no reload. handle { reverse_proxy {file./srv/control/live}:3000 { # node restarting in the live slot? retry the dial instead of answering 502 lb_try_duration 5s lb_try_interval 250ms } } log }
# Runtime variables for app-blue and app-green (server-side only, never in the browser bundle). # Changing them needs a restart of the slots: docker compose up -d app-blue app-green # STORE_API_URL=https://api.example.com/stores # DATABASE_URL=postgres://app:<PASSWORD>@db:5432/app
# Build-time variables for the builder. NEXT_PUBLIC_* values are inlined into the JavaScript # that every visitor downloads - never put a secret here. # NEXT_PUBLIC_MAPS_STYLE=light
#!/usr/bin/env python3 """hook.py - a tiny GitHub webhook listener (Python standard library only). It checks the X-Hub-Signature-256 HMAC, accepts push events for one branch and then only "rings the doorbell": it writes a small request file that sync.sh watches. It never runs git and never trusts the payload for WHAT to deploy - sync.sh always fetches the branch from the repository itself. Originally developed by S&H Software Solutions - first published on APEX Note (www.apexnote.de), October 2026. MIT licence. """ import hashlib import hmac import json import os import signal import sys import threading import time from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer from urllib.parse import parse_qs PORT = int(os.environ.get("HOOK_PORT", "9000")) HOOK_PATH = os.environ.get("HOOK_PATH", "/hooks/github") BRANCH = os.environ.get("DEPLOY_BRANCH", "main") REPO = os.environ.get("REPO_FULL_NAME", "") # optional: "owner/repo" REQUEST_FILE = os.environ.get("REQUEST_FILE", "/signal/request") MAX_BODY = int(os.environ.get("MAX_BODY_BYTES", str(5 * 1024 * 1024))) SECRET_FILE = os.environ.get("WEBHOOK_SECRET_FILE", "/run/secrets/webhook_secret") def log(msg): print(time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), "hook", msg, flush=True) try: with open(SECRET_FILE, "rb") as f: SECRET = f.read().strip() # "echo secret > file" adds a newline except OSError as e: sys.exit(f"hook: cannot read webhook secret: {e}") if len(SECRET) < 16: sys.exit("hook: webhook secret is empty or shorter than 16 characters") def ring_doorbell(after, delivery): """Atomically replace the request file - sync.sh never reads half a file.""" tmp = f"{REQUEST_FILE}.{os.getpid()}.{threading.get_ident()}.tmp" with open(tmp, "w", encoding="utf-8") as f: json.dump({"after": after, "delivery": delivery, "at": int(time.time())}, f) os.replace(tmp, REQUEST_FILE) class Hook(BaseHTTPRequestHandler): server_version = "hook" sys_version = "" timeout = 10 # slow or stuck clients are dropped def log_message(self, *args): # one line per delivery is enough pass def reply(self, code, text): body = (text + "\n").encode() self.send_response(code) self.send_header("Content-Type", "text/plain; charset=utf-8") self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) def do_GET(self): if self.path == "/healthz": self.reply(200, "ok") else: self.reply(404, "not found") def do_POST(self): delivery = self.headers.get("X-GitHub-Delivery", "-")[:64] event = self.headers.get("X-GitHub-Event", "-")[:64] try: code, text = self.handle_delivery(event, delivery) except Exception as e: # never leak a traceback to the caller log(f"internal error: {e!r}") code, text = 500, "internal error" log(f"{code} event={event} delivery={delivery} {text}") self.reply(code, text) def handle_delivery(self, event, delivery): if self.path.split("?", 1)[0] != HOOK_PATH: return 404, "not found" try: length = int(self.headers.get("Content-Length", "")) except ValueError: return 411, "length required" if length < 0 or length > MAX_BODY: return 413, "payload too large" body = self.rfile.read(length) # 1. Authenticity: HMAC-SHA256 over the RAW body, compared in constant time expected = "sha256=" + hmac.new(SECRET, body, hashlib.sha256).hexdigest() received = self.headers.get("X-Hub-Signature-256", "") if not hmac.compare_digest(expected.encode(), received.encode("utf-8", "replace")): return 401, "bad signature" # 2. Only push events for the deploy branch if event == "ping": return 200, "pong" if event != "push": return 202, f"ignored event {event}" try: if self.headers.get("Content-Type", "").startswith("application/x-www-form-urlencoded"): body = parse_qs(body.decode("utf-8"))["payload"][0] data = json.loads(body) except (ValueError, KeyError, IndexError): return 400, "invalid payload" if not isinstance(data, dict): return 400, "invalid payload" repo = (data.get("repository") or {}).get("full_name", "") if REPO and repo != REPO: return 403, f"wrong repository {repo}" ref = str(data.get("ref", "")) if ref != f"refs/heads/{BRANCH}": return 202, f"ignored {ref[:80]}" if data.get("deleted"): return 202, "ignored branch deletion" # 3. Ring the doorbell - the actual deploy runs in sync.sh after = str(data.get("after", ""))[:40] ring_doorbell(after, delivery) return 202, f"deploy queued for {after[:7]}" def main(): signal.signal(signal.SIGTERM, lambda *_: sys.exit(0)) os.umask(0o022) server = ThreadingHTTPServer(("0.0.0.0", PORT), Hook) server.daemon_threads = True log(f"listening on :{PORT}{HOOK_PATH} - branch {BRANCH}" + (f", repository {REPO}" if REPO else "")) server.serve_forever() if __name__ == "__main__": main()
# Internet-facing listener: Python standard library only - no git, no ssh, no shell tools needed FROM python:3.12-alpine RUN adduser -D -H -u 1001 hook && mkdir /signal && chown hook:hook /signal COPY --chmod=755 hook.py /usr/local/bin/hook.py USER hook EXPOSE 9000 ENTRYPOINT ["hook.py"]
#!/usr/bin/env bash # Send a push event signed exactly like GitHub signs it (X-Hub-Signature-256). # usage: send-test-webhook.sh URL SECRET_FILE [branch] [event] # extra curl options via CURL_OPTS, e.g. CURL_OPTS="--cacert root.crt" set -euo pipefail url=$1 secret_file=$2 branch=${3:-main} event=${4:-push} payload=$(printf '{"ref":"refs/heads/%s","after":"%s","deleted":false,"repository":{"full_name":"%s"}}' \ "$branch" "$(openssl rand -hex 20)" "${REPO_FULL_NAME:-your-name/your-site}") sig=$(printf '%s' "$payload" | openssl dgst -sha256 -hmac "$(cat "$secret_file")" | sed 's/^.*= //') curl -sS ${CURL_OPTS:-} -X POST "$url" -w ' [HTTP %{http_code}]\n' \ -H 'Content-Type: application/json' \ -H 'User-Agent: GitHub-Hookshot/send-test-webhook' \ -H "X-GitHub-Event: $event" \ -H "X-GitHub-Delivery: $(cat /proc/sys/kernel/random/uuid)" \ -H "X-Hub-Signature-256: sha256=$sig" \ --data-binary "$payload"
cp .env.example .env && nano .env # SITE_DOMAIN, REPO_URL, REPO_FULL_NAME
cp app.env.example app.env # optional: runtime variables of your app
docker compose up -d --build
docker compose logs -f deployer builder # first clone + build, then: deploy live: <commit> ...$ docker compose up -d --build 2>/dev/null
$ docker compose logs -f --no-log-prefix deployer builder | sed "/deploy live:/q"
2026-10-07T20:28:50Z build waiting for build requests in /srv/source/request
2026-10-07T20:28:50Z deploy watching /signal/request - branch main, fallback sync every 120s
2026-10-07T20:28:50Z deploy first run: cloning branch main
2026-10-07T20:28:50Z deploy new commit f384989 "Version 1: first release" - building
2026-10-07T20:28:51Z build building f384989 "Version 1: first release"
2026-10-07T20:29:08Z build built f384989 in 17s (npm ci 10s, next build 6s, release 67.1M)
2026-10-07T20:29:08Z deploy build ok in 17s (built)
2026-10-07T20:29:08Z deploy starting f384989 in slot blue, waiting for /api/health
2026-10-07T20:29:10Z deploy live: f384989 "Version 1: first release" on blue in 20s (build 17s)
$ docker compose ps --format "table {{.Service}}\t{{.Status}}\t{{.Ports}}"
SERVICE STATUS PORTS
app-blue Up 21 seconds (health: starting) 3000/tcp
app-green Up 21 seconds (health: starting) 3000/tcp
builder Up 21 seconds
deployer Up 21 seconds (health: starting)
gitea Up 24 seconds 22/tcp, 127.0.0.1:3300->3000/tcp
hook Up 21 seconds (health: starting) 9000/tcp
proxy Up 20 seconds 127.0.0.1:80->80/tcp, 127.0.0.1:443->443/tcp, 443/udp, 2019/tcpThe first deploy has no cache: 17 seconds from clone to live, 10 of them npm ci. Only Caddy publishes ports; the slots show 3000/tcp on the internal network.
The First Push
Version 1 is live in slot blue. Commit, build time and slot come from the deployment itself:
Now change the intro text, commit, push — and measure until the page shows the new version:
9.7 seconds from git push to the new page. About 2–3 seconds of that is the webhook delivery itself; the build took 3 seconds because the dependencies had not changed. The new version runs in slot green — and slot blue still runs Version 1, ready for a rollback:
Prove It: Downtime, Broken Build, Rollback
"Zero downtime" is a claim; I wanted a number. watch-downtime.mjs runs N parallel clients with keep-alive connections (the harder case — a reused connection that gets closed fails, a new one would just retry) and counts every request that does not return 200 with a complete page:
#!/usr/bin/env node // watch-downtime.mjs - hammer a URL with N parallel keep-alive clients and count // every request that does not return a complete page. Ctrl+C or the time limit ends it. // usage: node watch-downtime.mjs URL SECONDS [CLIENTS] // A request counts as "ok" only with HTTP 200 and a complete HTML page (</html>). import http from "node:http"; import https from "node:https"; const [url, seconds = "30", clients = "4"] = process.argv.slice(2); if (!url) { console.error("usage: node watch-downtime.mjs URL SECONDS [CLIENTS]"); process.exit(2); } const lib = url.startsWith("https") ? https : http; const agent = new lib.Agent({ keepAlive: true, maxSockets: Number(clients) }); const start = Date.now(), end = start + Number(seconds) * 1000; const t = () => ((Date.now() - start) / 1000).toFixed(1).padStart(5) + "s"; let ok = 0, failed = 0, lastOk = start, maxGap = 0, firstFail = null, lastFail = null, last = ""; const versions = new Map(), errors = new Map(); function once() { return new Promise((done) => { let settled = false; const resolve = (r) => { if (!settled) { settled = true; done(r); } }; const req = lib.get(url, { agent, timeout: 5000 }, (res) => { let body = ""; res.setEncoding("utf8"); res.on("data", (c) => (body += c)); res.on("error", (e) => resolve({ error: e.code || e.message })); res.on("close", () => resolve({ error: "response cut off" })); // only if "end" never came res.on("end", () => resolve(res.statusCode === 200 && body.includes("</html>") ? { version: (body.match(/Version (\d+)/) || [, "?"])[1], slot: (body.match(/slot (blue|green)/) || [, "?"])[1] } : { error: `HTTP ${res.statusCode}` })); }); req.on("timeout", () => req.destroy(new Error("timeout"))); req.on("error", (e) => resolve({ error: e.code || e.message })); }); } async function client() { while (Date.now() < end) { const r = await once(); const now = Date.now(); if (r.error) { failed++; firstFail ??= now; lastFail = now; errors.set(r.error, (errors.get(r.error) || 0) + 1); if (last !== r.error) { console.log(`${t()} FAIL ${r.error}`); last = r.error; } await new Promise((s) => setTimeout(s, 50)); } else { ok++; maxGap = Math.max(maxGap, now - lastOk); lastOk = now; const key = `Version ${r.version} (slot ${r.slot})`; versions.set(key, (versions.get(key) || 0) + 1); if (last !== key) { console.log(`${t()} ok ${key}`); last = key; } } } } await Promise.all(Array.from({ length: Number(clients) }, client)); maxGap = Math.max(maxGap, Date.now() - lastOk); // a site that never answers counts too console.log(`---\nrequests: ${ok + failed} ok: ${ok} failed: ${failed}`); for (const [k, v] of versions) console.log(` ${k}: ${v}`); for (const [k, v] of errors) console.log(` ${k}: ${v}`); if (failed) console.log(`site down from +${((firstFail - start) / 1000).toFixed(1)}s to +${((lastFail - start) / 1000).toFixed(1)}s`); console.log(`longest time without a good answer: ${maxGap} ms`); agent.destroy();
Deploy and switch under load
Four clients, one real deploy (Version 3) and then three rollbacks and three un-pins in a row — seven switches within seven seconds:
7,543 requests, 0 failed. The longest pause between two good answers was 113 ms — the CPU was busy with next build at that moment. A final regression run with the published code from an empty lab (cold build, one deploy, six switches, one broken build) gave 15,463 requests, 0 failed.
A broken commit
The same JSX typo that killed my draft, pushed to the new setup while the site is under load:
8f6cb02 live — 5,843 requests, 0 failed. (npm's update notice cropped.)The builder names the step and shows the compiler error, the deployer says in one line what happened and what stays live, and the clients did not notice anything. The full log stays in /srv/releases/.logs/<commit>.log. The deployer remembers the failed commit, so the 5-minute fallback sync does not rebuild it again and again; the next push (or deploy.mjs unpin) tries again.
Rollback
A release that builds and passes the health check can still be wrong. Because the previous release keeps running in the other slot, rollback is a file write — 0.3 seconds including docker compose exec — and it pins the version so the next fallback sync does not undo it:
$ docker compose exec deployer deploy.mjs status
live: f3487a4 slot green 2026-10-07T20:33:19.734Z Version 5: fix JSX
standby: 8f6cb02 slot blue 2026-10-07T20:30:10.188Z Version 3: deploy under load
releases: 03ebd6e 8f6cb02 f3487a4 f384989
$ time docker compose exec deployer deploy.mjs rollback
2026-10-07T20:33:25Z deploy rolled back to 8f6cb02 on blue and pinned - run 'deploy.mjs unpin' to deploy main again
real 0m0.292s
user 0m0.044s
sys 0m0.037s
$ curl -s https://www.example.com/api/health; echo
{"status":"ok","commit":"8f6cb02","slot":"blue","uptime":202}
$ docker compose exec deployer deploy.mjs status
live: 8f6cb02 slot blue 2026-10-07T20:30:10.188Z Version 3: deploy under load
standby: f3487a4 slot green 2026-10-07T20:33:19.734Z Version 5: fix JSX
pinned: yes (deploy.mjs unpin)
releases: 03ebd6e 8f6cb02 f3487a4 f384989
$ docker compose exec deployer deploy.mjs unpin
2026-10-07T20:33:25Z deploy new commit f3487a4 "Version 5: fix JSX" - building
2026-10-07T20:33:26Z deploy build ok in 0s (cached)
2026-10-07T20:33:26Z deploy starting f3487a4 in slot green, waiting for /api/health
2026-10-07T20:33:26Z deploy live: f3487a4 "Version 5: fix JSX" on green in 1s (build 0s)How It Works Inside
1. The switch is a file, not a reload — because my first switch dropped requests
My first version switched Caddy through its admin API: post the Caddyfile with the new upstream to /load over a unix socket. Caddy calls that a graceful reload, and for new connections it is. But the reload closes idle keep-alive connections, and a client that sends its next request on such a connection at that exact moment gets a reset. My load test caught it on the first run:
# Console output of my first zero-downtime test (20:21 UTC, 7 Oct 2026).
# At that time the deployer switched Caddy via its admin API: POST /load with the
# Caddyfile (upstream replaced) over a unix socket. Copied from the terminal - this
# run was not recorded with `script`. Command:
# node tools/watch-downtime.mjs https://www.example.com/ 12 4 &
# docker compose exec -T deployer deploy.mjs rollback; ...; deploy.mjs unpin
0.1s ok Version 2 (slot green)
3.2s FAIL ECONNRESET
2026-10-07T20:21:57Z deploy rolled back to 79c2258 on blue in 55 ms and pinned - run 'deploy.mjs unpin' to deploy main again
3.2s ok Version 2 (slot green)
3.3s ok Version 1 (slot blue)
live: 79c2258 slot blue 2026-10-07T20:20:22.974Z Version 1: first release
standby: 4aa4dd4 slot green 2026-10-07T20:21:27.554Z Version 2: new intro text
pinned: yes (deploy.mjs unpin)
releases: 4aa4dd4 79c2258
2026-10-07T20:22:00Z deploy new commit 4aa4dd4 "Version 2: new intro text" - building
2026-10-07T20:22:01Z deploy build ok in 0s (cached)
2026-10-07T20:22:01Z deploy starting 4aa4dd4 in slot green, waiting for /api/health
2026-10-07T20:22:02Z deploy live: 4aa4dd4 "Version 2: new intro text" on green in 1s (build 0s, switch 53 ms)
8.2s FAIL ECONNRESET
8.2s ok Version 1 (slot blue)
8.3s ok Version 2 (slot green)
---
requests: 1880 ok: 1878 failed: 2
Version 2 (slot green): 1096
Version 1 (slot blue): 782
ECONNRESET: 2
site down from +3.2s to +8.2s
longest time without a good answer: 109 ms
# Note: "site down from ... to ..." is first/last failure - there were exactly two
# failed requests, one per config reload. Same test with the file switch: 7,313 ok, 0 failed.Two failures in 1,880 requests — one per reload. Browsers usually retry a GET on a reused connection silently, but a POST (a form, a server action) is not retried. The fix was simpler than the original: Caddy's {file.*} placeholder reads /srv/control/live on every request, and deploy.mjs replaces that file with writeFileSync + renameSync. Nothing is reloaded, no connection is closed, the admin API is not needed at all, and a restarted Caddy reads the same file — no state to restore. The cost is one cached file read per request.
2. npm runs in a container that has nothing to steal
npm ci installs hundreds of packages, and any of them can run an install script. In my draft that script would have run next to the GitHub token. Here the builder has no deploy key, no webhook secret, no access to the running app or to Caddy — its only network goes to the npm registry, and its only output is a folder that has to pass the health check before it serves anything. The deployer, which holds the key, never runs npm.
3. The health check knows which build it talks to
"Port 3000 answers" is not enough: during the switch the slot may still run the old process for a second. The deployer waits until /api/health reports its commit, then requests GET / — because an app can be healthy and still crash on its main page. I tested exactly that with a commit that throws when STORE_API_URL is missing. It built fine, /api/health said ok, the page answered 500:
$ git diff -U0 app/page.js | tail -1; git commit -qam "Version 6: needs STORE_API_URL" && git push -q origin main 2>/dev/null
+ if (!process.env.STORE_API_URL) throw new Error("STORE_API_URL is not set");
$ docker compose logs -f --no-log-prefix deployer | sed -n "/Version 6/,/standby for rollback/p"
2026-10-07T20:33:49Z deploy new commit 6687188 "Version 6: needs STORE_API_URL" - building
2026-10-07T20:33:53Z deploy build ok in 3s (built)
2026-10-07T20:33:53Z deploy starting 6687188 in slot blue, waiting for /api/health
2026-10-07T20:34:54Z deploy slot blue not healthy after 60s (GET / answered 500)
2026-10-07T20:34:54Z deploy ERROR: 6687188 did not become healthy in slot blue - f3487a4 on green stays live
2026-10-07T20:34:54Z deploy slot blue goes back to 8f6cb02 (standby for rollback)
$ curl -s https://www.example.com/api/health; echo
{"status":"ok","commit":"f3487a4","slot":"green","uptime":159}
$ docker compose logs --no-log-prefix --since 5m app-blue | grep -E "slot-blue|STORE_API_URL" | head -4
2026-10-07T20:33:54Z slot-blue starting release 6687188
⨯ Error: STORE_API_URL is not set
⨯ Error: STORE_API_URL is not set
⨯ Error: STORE_API_URL is not setAfter 60 seconds the deployer gives up, the live slot is untouched, and the idle slot goes back to the previous release so a rollback target still exists.
4. node_modules, npm cache and .next/cache
The builder keeps three things between builds in volumes: the npm cache, Next.js' .next/cache (linked into each build), and the last node_modules together with a hash of package.json, package-lock.json and node -v. Measured on the demo app:
| Situation | npm ci | next build | Build total |
|---|---|---|---|
| First build, empty caches | 10–11 s | 6 s | 17 s |
| Dependencies changed, npm cache warm | 8 s | 2 s | 12 s |
| Dependencies unchanged (most pushes) | skipped | 2–3 s | 3 s |
$ docker compose logs --no-log-prefix --since 2m builder | grep "Rebuild\|built"
2026-10-07T20:39:06Z build built b787fc8 in 3s (npm ci skipped, next build 2s, release 67.1M)
2026-10-07T20:39:31Z build building 85b5c70 "Rebuild: npm ci with warm cache"
$ docker compose exec builder rm -f /build/deps # forget node_modules again, keep the npm cache
$ (while :; do docker stats --no-stream --format "{{.MemUsage}}" blog3-builder; done) > /tmp/mem.log &
$ git commit -q --allow-empty -m "Rebuild 2: measure memory" && git push -q origin main 2>/dev/null
$ timeout 120 docker compose logs -f --no-log-prefix --since 1s builder | sed "/built/q"; kill %1; sort -h /tmp/mem.log | tail -1
2026-10-07T20:39:43Z build built 85b5c70 in 12s (npm ci 8s, next build 2s, release 67.1M)
608.2MiB / 2GiBMost of npm ci is unpacking Next.js' native SWC binaries, not downloading — that is why even a warm cache costs 8 seconds. Skipping it when nothing changed is safe because the hash includes the Node.js version (native modules) and the lock file.
5. NEXT_PUBLIC_* is frozen at build time
I started the live release a second time inside its slot, on port 3001, with both variables "changed at runtime":
$ docker compose exec app-green sh -c 'cd /srv/releases/$(cat /srv/control/slots/green) && NEXT_PUBLIC_COMMIT=changed-at-runtime APP_SLOT=changed-at-runtime HOSTNAME=127.0.0.1 PORT=3001 timeout 4 node server.js >/dev/null & sleep 2; wget -qO- 127.0.0.1:3001 | grep -oE "<th>(Commit|Served by)</th><td>[^<]*" | sed "s/<[^>]*>/ /g"; wait'
Commit f3487a4 · Version 5: fix JSX
Served by slot changed-at-runtime
$ docker compose exec app-green sh -c 'cd /srv/releases/$(cat /srv/control/slots/green) && grep -rl "f3487a4" .next | head -3'
.next/server/chunks/ssr/[root-of-the-server]__02u33qrikb078._.js
.next/server/chunks/[root-of-the-server]__0sp20yik1olc-._.jsThe commit did not change — the value is a string in the compiled server chunks (second command). APP_SLOT did. Rule of thumb: everything with NEXT_PUBLIC_ goes into build.env (builder) and ends up in every visitor's browser, so never put a secret there; everything else goes into app.env (slots) and is read per request. After changing app.env, recreate the slots — thanks to the dial retry in Caddy the live slot was recreated with 0 failed requests in my second test (longest wait 1 second). In the first attempt one request that was in flight when the old container stopped was cut off — so recreate the live slot in a quiet minute:
$ cat app.env; curl -s https://www.example.com/api/health; echo
STORE_API_URL=https://api.example.com/stores
{"status":"ok","commit":"da42f32","slot":"green","uptime":19}
$ node tools/watch-downtime.mjs https://www.example.com/ 15 4 > /tmp/env.log & sleep 3; docker compose up -d --force-recreate app-green 2>&1 | tail -1; wait; tail -4 /tmp/env.log
Container blog3-app-green Started
---
requests: 2514 ok: 2514 failed: 0
Version 9 (slot green): 2514
longest time without a good answer: 999 ms
$ docker compose exec app-green printenv STORE_API_URL
https://api.example.com/stores
^@6. Memory: the build needs more than the app
The running app needs about 40–70 MB. The build peaked at about 600 MiB in docker stats for this tiny app; real apps need 1–3 GB. On a small VPS the kernel kills the build — I simulated that with a 256 MB limit:
$ docker update --memory 256m --memory-swap 256m blog3-builder >/dev/null # simulate a small VPS
$ git commit -qam "Version 7: STORE_API_URL optional" && git push -q origin main 2>/dev/null
$ timeout 120 docker compose logs -f --no-log-prefix --since 5s deployer builder | sed "/stays live/q"
2026-10-07T20:36:52Z deploy doorbell: {"after": "b787fc8acbc5ff11bcd85fa3b195feb05451a9ef", "delivery": "2f4947b5-fdd8-47fc-9efb-293446ad3e2d", "at": 1791405412}
2026-10-07T20:36:52Z deploy new commit b787fc8 "Version 7: STORE_API_URL optional" - building
2026-10-07T20:36:53Z build building b787fc8 "Version 7: STORE_API_URL optional"
2026-10-07T20:36:55Z build FAILED in step "next build" after 2s - last lines of /srv/releases/.logs/b787fc8acbc5ff11bcd85fa3b195feb05451a9ef.log:
|
| ### copy source
|
| ### npm ci skipped - dependencies unchanged (ef0e40a1b74993f6)
|
| ### next build
|
| > apexnote-demo-app@1.0.0 build
| > next build
|
| ▲ Next.js 16.4.0 (Turbopack)
| ✓ Running next.config.mjs took 22ms
|
| Creating an optimized production build ...
| Killed
2026-10-07T20:36:55Z deploy ERROR: build of b787fc8 failed in step "next build" after 2s - f3487a4 on green stays live
$ docker inspect blog3-builder --format "OOMKilled={{.State.OOMKilled}} RestartCount={{.RestartCount}}"; dmesg 2>/dev/null | grep -i "killed process" | tail -1
OOMKilled=true RestartCount=0
[ 4512.248865] Memory cgroup out of memory: Killed process 13441 (next-build (v16) total-vm:22931848kB, anon-rss:239392kB, file-rss:125656kB, shmem-rss:0kB, UID:1000 pgtables:2228kB oom_score_adj:0
$ docker update --memory 2g --memory-swap 2g blog3-builder >/dev/null
$ docker compose exec deployer deploy.mjs unpin
2026-10-07T20:39:03Z deploy new commit b787fc8 "Version 7: STORE_API_URL optional" - building
2026-10-07T20:39:06Z deploy build ok in 3s (built)
2026-10-07T20:39:06Z deploy starting b787fc8 in slot blue, waiting for /api/health
2026-10-07T20:39:07Z deploy live: b787fc8 "Version 7: STORE_API_URL optional" on blue in 5s (build 3s)
$ docker compose logs --no-log-prefix --since 30s builder | tail -1; docker stats --no-stream --format "table {{.Name}}\t{{.MemUsage}}" blog3-builder blog3-app-blue blog3-app-green
2026-10-07T20:39:06Z build built b787fc8 in 3s (npm ci skipped, next build 2s, release 67.1M)
NAME MEM USAGE / LIMIT
blog3-builder 16.61MiB / 2GiB
blog3-app-blue 37.51MiB / 512MiB
blog3-app-green 69.81MiB / 512MiBBecause build and app live in different containers, the out-of-memory kill hit only the builder; the site kept serving. In my draft the same kill would have taken the site down with it. Give the builder its own mem_limit (2 GB here) and keep --max-old-space-size below it, add swap on small servers, or build in CI (see below).
7. Why releases are folders, not Docker images
The textbook answer is "build an image per commit, tag it, roll back to the previous tag". On a single server without a CI platform that needs the Docker API — docker build and docker run — which means docker.sock, which means root on the host for whoever controls the deployer. A standalone release is already self-contained (server, traced node_modules, static files), so an immutable folder per commit gives the same properties: a release never changes, rollback is "run the previous one", the last five are kept. If you build images in GitHub Actions anyway, push them to a registry and let a pull-based agent deploy them — but that is a different post, and I have not tested it here.
Security & Edge Cases
| Risk | What this setup does |
|---|---|
| Anyone triggers deploys | HMAC-SHA256 over the raw body, constant-time compare, 401 without a valid signature (tested below) |
| Forged payload deploys foreign code | Doorbell pattern: the deployer always fetches main of REPO_URL itself; a payload with a made-up commit only causes already live |
docker.sock | Not mounted anywhere — the slots restart their own process |
| Malicious npm package | Runs in the builder: no secrets, no route to app or proxy; its output must pass the health check |
| Compromised app | Read-only filesystem, read-only releases, internal network without internet, no capabilities, non-root |
| Secrets in the browser bundle | Only build.env reaches the build; runtime secrets live in app.env |
| Build eats the server's RAM | mem_limit: 2g on the builder, 512 MB per slot |
$ tools/send-test-webhook.sh https://www.example.com/_hooks/github secrets/webhook_secret main
deploy queued for fdbb9e7
[HTTP 202]
$ openssl rand -hex 32 > /tmp/wrong; tools/send-test-webhook.sh https://www.example.com/_hooks/github /tmp/wrong main
bad signature
[HTTP 401]
$ curl -s -X POST https://www.example.com/_hooks/github -H "X-GitHub-Event: push" -d "{}" -w " [HTTP %{http_code}]\n"
bad signature
[HTTP 401]
$ tools/send-test-webhook.sh https://www.example.com/_hooks/github secrets/webhook_secret feature/x
ignored refs/heads/feature/x
[HTTP 202]
$ curl -s -o /dev/null -w "GET /_hooks/github -> %{http_code}\n" https://www.example.com/_hooks/github
GET /_hooks/github -> 404
$ sleep 2; docker compose logs --no-log-prefix --since 10s hook deployer
2026-10-07T20:45:16Z hook 202 event=push delivery=d9f1c9a1-c64f-4525-8227-69c2d6729c18 deploy queued for fdbb9e7
2026-10-07T20:45:16Z hook 401 event=push delivery=8e600439-9328-4268-8e09-f97d0e61b929 bad signature
2026-10-07T20:45:16Z hook 401 event=push delivery=- bad signature
2026-10-07T20:45:17Z hook 202 event=push delivery=62af2342-e3e2-4cbf-ab18-0150259b3932 ignored refs/heads/feature/x
2026-10-07T20:45:17Z deploy doorbell: {"after": "fdbb9e7be5a3a1e3d698c7a918a147559ee76271", "delivery": "d9f1c9a1-c64f-4525-8227-69c2d6729c18", "at": 1791405916}
2026-10-07T20:45:17Z deploy already live: da42f32 on greenEdge cases you should know before you rely on it:
- A crash of the live process —
slot.shrestarts node within a second. Without the dial retry my test showed 48 ×502; withlb_try_duration 5sit was 0 failed requests, the slowest answer waited 1 second (output below) - Restarting Caddy itself is the one thing that drops connections (42 failed requests, 0.6 s in my test). The deploy never does it; do it in a quiet minute
- Database migrations — during a switch, old and new code run at the same time. Make migrations backwards-compatible (expand first, contract in a later release)
- WebSockets and long requests stay on the old slot until it is replaced by the next deploy — that is a feature, but plan for it
- ISR and the image optimizer want to write into
.next/cache; releases are read-only here. For a fully dynamic app like the demo that is fine — if you use ISR, mount a writable cache volume per slot or configure a cache handler - More than one server — this is a single-host design; with several hosts, use a registry and an orchestrator
$ node tools/watch-downtime.mjs https://www.example.com/ 12 4 > /tmp/crash.log & sleep 3; docker compose exec -T app-green pkill -9 -f next-server; wait; cat /tmp/crash.log
0.1s ok Version 9 (slot green)
3.1s FAIL HTTP 502
4.1s ok Version 9 (slot green)
---
requests: 1977 ok: 1929 failed: 48
Version 9 (slot green): 1929
HTTP 502: 48
site down from +3.1s to +3.6s
longest time without a good answer: 1024 ms
$ docker compose logs --no-log-prefix --since 15s app-green | grep slot-
2026-10-07T20:47:15Z slot-green node stopped (exit 137) - starting it again
2026-10-07T20:47:15Z slot-green starting release da42f32
$ node tools/watch-downtime.mjs https://www.example.com/ 15 4 > /tmp/proxy.log & sleep 3; docker compose restart proxy 2>&1 | tail -1; wait; tail -6 /tmp/proxy.log
Container blog3-proxy Started
requests: 2822 ok: 2780 failed: 42
Version 9 (slot green): 2780
ECONNRESET: 7
ECONNREFUSED: 35
site down from +3.2s to +3.7s
longest time without a good answer: 623 ms
$ grep -A4 "reverse_proxy {file" Caddyfile
reverse_proxy {file./srv/control/live}:3000 {
# node restarting in the live slot? retry the dial instead of answering 502
lb_try_duration 5s
lb_try_interval 250ms
}
$ node tools/watch-downtime.mjs https://www.example.com/ 12 4 > /tmp/crash.log & sleep 3; docker compose exec -T app-green pkill -9 -f next-server; wait; cat /tmp/crash.log
0.1s ok Version 9 (slot green)
---
requests: 1765 ok: 1765 failed: 0
Version 9 (slot green): 1765
longest time without a good answer: 1052 msTroubleshooting: Real Errors and Their Fixes
Every message in this section came out of my lab while building this post.
Killed / OOMKilled=true during next build
The kernel killed the build for lack of memory (inside, point 6). Raise mem_limit of the builder, keep NODE_OPTIONS=--max-old-space-size about 25 % below it, add swap. The site is not affected — fix the limit and run docker compose exec deployer deploy.mjs unpin to retry the same commit.
Error: Turbopack build failed with 1 error
A real compile error in your code. The old version stays live; the builder log names file, line and column. Fix it, push again. If a commit was force-pushed away, the deployer simply builds the new head.
did not become healthy in slot … (GET / answered 500)
The build is fine, but the app fails at runtime — almost always a missing runtime variable. docker compose logs app-blue shows the exception (⨯ Error: STORE_API_URL is not set in my test). Add it to app.env, recreate the slots, deploy.mjs unpin.
npm error SELF_SIGNED_CERT_IN_CHAIN
My first npm view next version inside a container failed with this, because the lab's egress proxy re-signs TLS. If your server sits behind such a proxy, give the builder the proxy's CA (NODE_EXTRA_CA_CERTS, mounted read-only) — never strict-ssl=false.
mkdir: can't create directory '/srv/releases/.results': Permission denied
My very first start. A named volume gets the owner of the directory in the image of the first container that mounts it — and that was the deployer, whose image had /srv/releases owned by root. All images in this post now create every shared directory with the same owner (UID 1000), and the proxy, which cannot be changed, starts after the deployer (depends_on). If you hit it: docker compose down, docker volume rm the affected volume, start again.
dial tcp: lookup 127.0.0.1:9001:80: no such host
From my first {file.*} test. If the upstream is a placeholder, Caddy cannot see a port and adds :80. Write the port outside the placeholder ({file./srv/control/live}:3000) and only the host name into the file.
deploy ERROR: Internal Server Connection Error … Could not read from remote repository
The first clone ran while Gitea was still starting. The deployer retries after 30 seconds instead of waiting for the 5-minute fallback; with GitHub you see this only during outages or when the deploy key is missing (Permission denied (publickey) — see the PHP post).
Complete code in this post — copy it from Steps 2–7; every code box has a Copy button. MIT licensed and free for personal and commercial projects. Originally developed by S&H Software Solutions – first published on APEX Note (October 2026).
Questions? Leave a comment below.
Final Thoughts
The PHP version of this idea was about security: who may trigger a deploy, and what a reachable .git folder gives away. The Next.js version is about time: as soon as a build sits between git push and a running server, "restart the container" means "take the site offline for as long as the build runs — or forever, if it fails". The answer is not a bigger CI platform; it is three small decisions: build next to the live version, prove the new one works, and make the switch so small that nobody can catch it halfway.
The most useful finding was the one I did not plan for: a "graceful" proxy reload that still dropped one request per switch. Without a load test I would have shipped that and called it zero-downtime. Measure it — watch-downtime.mjs is in the post for exactly that reason.
If you run this with a monorepo, pnpm, or an app that needs ISR, and something does not fit — leave a comment below. I read all of them.
{fullWidth}