Auto-Deploy a Next.js App from GitHub with Docker – Zero Downtime

Docker Next.js GitHub Zero Downtime Open Source

Git Push → Live Without Downtime: Auto-Deploy a Next.js App with Docker, a Signed Webhook and Blue/Green Slots

The PHP recipe breaks the moment a build step appears. This is the version for Next.js: build next to the running app, health-check it, switch with one rename() — and a broken commit never reaches your visitors. No CI platform, no docker.sock.

📅 October 2026 👤 Sajjad Hanifa ⏱ ~20 min read 🏢 S&H Software Solutions
Cover: Auto-deploy a Next.js app. Zero downtime. No docker.sock. git push to live in 9.7 seconds with 0 failed requests - with the real terminal run of seven switches under load and the browser screenshot of the APEX Note Demo App, Version 2

In my last post a push to main updated a PHP website in 1.4 seconds: fetch, export, switch a symlink. Then I tried the same idea on the Next.js front end that sits next to one of my Oracle APEX applications — and it fell apart, because a Next.js app has to be built before it can run. The build takes time, it needs RAM, and it can fail.

My first attempt did what most forum answers do: one node:22-alpine container that clones the repository, runs npm install and npm run build, then starts server.js — and a second container that restarts it through docker.sock when a webhook arrives. I rebuilt it in my lab and measured it: every deploy takes the site offline for the whole build, and one commit with a typo kept it offline until the next fix was pushed.

This post is the setup I use now. Who it is for: anyone who runs a Next.js (or any Node.js) app in Docker on their own server and wants Git-based deployments without a CI platform. What you have at the end: push → live in about 10 seconds, 0 failed requests during deploys, switches and rollbacks (measured with four clients hammering the site), a broken build that is reported in the log and never goes live, rollback in 0.3 seconds — and no container that can control your Docker daemon.

Free and open source. Complete code in this post — copy it from Steps 4–7. MIT licensed. Files: deploy.mjs, build.sh, slot.sh, hook.py, four small Dockerfiles, Caddyfile, docker-compose.yml and a measuring tool. Originally developed by S&H Software Solutions – first published on APEX Note (October 2026).

TL;DR — what the setup does:

  • Builds next to the running version — a separate builder container runs npm ci and next build; the live app is not touched until the new build has passed
  • Two app slots, blue and green — the new release starts in the idle slot and must answer /api/health with the right commit and GET / with 200 before it gets any traffic
  • Switches with one rename() — Caddy reads the name of the live slot from a file on every request; no reload, no admin API, no dropped connection
  • Broken build = log entry, not an outage — the failing step and the compiler error are in the log, the old version keeps serving
  • Instant rollback — the previous release keeps running in the other slot; switching back is one file write
  • Signed webhooks only — the listener from the PHP post, unchanged: HMAC-SHA256, constant-time compare, doorbell pattern
  • No docker.sock, read-only containers, dropped capabilities, the app on an internal network, the build without any secret

My First Attempt — and What It Did to the Site

Here is the draft, with the token and the IP address replaced. It is short, it is readable, and it "works":

YAMLdocker-compose.yml — my draftdo not use · sanitised
# My first attempt (sanitised: token and IP replaced) - do NOT use this
services:
  my_store_locator_app:
    image: node:22-alpine
    container_name: my_store_locator_app
    working_dir: /app
    command: >
      sh -c "apk add --no-cache git &&
             if [ -d /app/.git ]; then
               git -C /app fetch origin &&
               git -C /app reset --hard origin/main &&
               git -C /app clean -fd -e node_modules -e .next;
             else
               git clone https://<GITHUB_TOKEN>@github.com/<owner>/my-store-locator.git /app;
             fi &&
             npm install && npm run build &&
             cp -r /app/public /app/.next/standalone/public 2>/dev/null;
             cp -r /app/.next/static /app/.next/standalone/.next/static &&
             HOSTNAME=0.0.0.0 PORT=3000 node /app/.next/standalone/server.js"
    ports:
      - "4000:3000"
    volumes:
      - my_store_locator_data:/app
    restart: always

  my_store_locator_hook:
    image: docker:cli
    container_name: my_store_locator_hook
    command: >
      sh -c "while true; do
               printf 'HTTP/1.1 200 OK\r\nContent-Length: 2\r\n\r\nok' | nc -l -p 9001 > /dev/null || { sleep 5; continue; };
               echo 'Webhook received - restarting the app';
               docker restart my_store_locator_app;
             done"
    ports:
      - "9001:9001"
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock
    restart: always

volumes:
  my_store_locator_data:

The problems are not visible in the YAML, so I rebuilt it in the lab and put load on it. Four parallel clients request the page in a loop while I run exactly what the hook container runs — docker restart:

Terminal: watch-downtime.mjs with four clients against the draft on port 14000, then docker restart blog3-draft-app; the log shows failures (ECONNRESET, ECONNREFUSED) from 13.3 s to 17.9 s: 18,874 requests, 276 failed, longest time without a good answer 4,980 ms.
Real runFigure 1 — My draft under load: one docker restart — what its hook does on every push — costs 276 failed requests and five seconds without any answer, after ten seconds of waiting for a process that ignores SIGTERM.

276 failed requests and five seconds without an answer for an app that builds in three seconds. Your real app builds longer, and every second of next build is a second of downtime. Two more things hide in the numbers: the failures only start ten seconds after the restart, because sh -c as PID 1 ignores SIGTERM and Docker waits ten seconds before it kills the container — so in-flight requests are never finished gracefully. And the restart re-runs npm install against whatever the volume holds.

Then I pushed a commit with a JSX typo and restarted again. The build failed, && never reached node server.js, the container exited and restart: always started the next round:

Outputdraft + broken commitreal run · port 14000 in the lab
$ git log --oneline -1; git show HEAD -U0 --format= app/page.js | tail -2
7b99a41 (HEAD -> main, origin/main) Version 8: broken JSX again

+    ["Version", VERSION],,
+    ["Broken", <b>oops</i>],
$ docker restart blog3-draft-app >/dev/null; node ~/next-app/tools/watch-downtime.mjs http://127.0.0.1:14000/ 45 4 | tail -5
requests: 3282   ok: 0   failed: 3282
  ECONNRESET: 1580
  ECONNREFUSED: 1702
site down from +0.0s to +45.0s
longest time without a good answer: 45046 ms
$ docker inspect blog3-draft-app --format "status={{.State.Status}} exit={{.State.ExitCode}} restarts={{.RestartCount}}"; docker logs blog3-draft-app 2>&1 | grep -m1 -A3 "Build error"
status=restarting exit=1 restarts=8
> Build error occurred
Error: Turbopack build failed with 1 error:
./app/page.js:14:25
Error: Expected corresponding JSX closing tag for <b>
$ node ~/next-app/tools/watch-downtime.mjs https://www.example.com/ 10 4 | tail -4   # the blue/green stack, same broken commit
---
requests: 1909   ok: 1909   failed: 0
  Version 7 (slot blue): 1909
longest time without a good answer: 70 ms
Chromium showing http://127.0.0.1:14000/ with the error page This site can't be reached, The connection was reset, ERR_CONNECTION_RESET, while the draft container is in a restart loop after a broken commit.
Real runFigure 2 — The draft after the broken push: the container loops between failed build and restart, the browser gets ERR_CONNECTION_RESET — for as long as the bug is in main.

The site stayed down until I pushed a fix — and even then it took 41 seconds and 10 restarts until Docker's back-off let the loop pick it up. The same broken commit hit the setup from this post at the same moment: 0 failed requests, the old version kept serving (last line above). That difference is what this post is about.

Problem in the draftWhy it hurtsFix in this post
Build and run in the same containerThe old version must stop before the new one is built → downtime = build timeSeparate builder, two app slots, switch after the build
No check before going liveA failing build or a crashing start takes the site downBuild result + health check gate every switch
docker.sock in the hook containerRoot on the host for anyone who breaks the internet-facing listenerDoorbell file; nobody talks to the Docker daemon
Unsigned trigger on port 9001Anyone can restart (= take down) the siteHMAC-SHA256 signature, only through Caddy
Token in the clone URLStored in /app/.git/config on the volumeRead-only deploy key as a Compose secret
npm install in productionMay change the lock file, runs package scripts next to the tokennpm ci in a container without secrets

The Design: Build Beside, Check, Switch

Six containers, each with exactly the access it needs. The numbers in the diagram are the order of one deployment:

Zero-downtime Next.js deployment: GitHub, Caddy, hook, deployer, builder, two app slots and two volumes GitHub sends a signed push event through Caddy to the hook container, which only rings a doorbell. The deployer fetches the branch over SSH with a read-only deploy key and hands the exported source to the builder. The builder runs npm ci and next build and writes an immutable release into the releases volume. The deployer starts the release in the idle app slot, checks its health and then writes the name of that slot into the file control/live. Caddy reads that file on every request, so the switch needs no reload. GitHub repository · signed webhook Visitors https://www.example.com ① signed push event HTTPS proxy · Caddy reads control/live on every request → app-blue:3000 or app-green:3000 POST /_hooks/github all traffic → live slot hook · hook.py checks the HMAC signature no git, no keys, no app access app-blue LIVE · v2 node server.js app-green idle → v3 node server.js ② doorbell ⑤ start v3 · health check deployer · deploy.mjs read-only deploy key · decides no published port · never runs npm networks: app (internal) + git volume: control live · slots/blue · slots/green ⑥ switch = rename(live) ④ source + build request builder · build.sh npm ci · next build · mem_limit 2g internet for npm, nothing to steal network: build (no app, no proxy) volume: releases <commit>/server.js · immutable read-only ③ git fetch · read-only deploy key
Figure 3 — Diagram: one deployment from push to switch. Only the deployer decides; the builder runs npm without any secret; Caddy only reads which slot is live.
ContainerHoldsCan reachCannot
proxy (Caddy)TLS certificateshook, both slotschange anything — it only reads control/live
hook (Python)webhook secretnothing but its doorbell filerun git, see the app, start anything
deployer (Node.js)read-only deploy keyGitHub, the slots (health check)run npm; it has no published port
builder (Node.js)nothing secretthe npm registryreach the app, the proxy or the key
app-blue / app-greenruntime variables (app.env)internal network onlyreach the internet, change a release
Tested on 7 October 2026 in my APEX Note Lab with Docker 29.6.2 and Compose 5.3.1, node:22-alpine (Node.js 22.23.3, npm 10.9.9), Next.js 16.4.0 (Turbopack) with React 19.3.0, Caddy 2.11.7 and python:3.12-alpine. A local Gitea 1.24 plays GitHub: it hosts apexnote/demo-app, holds the read-only deploy key and sends the webhook with the same X-GitHub-Event and X-Hub-Signature-256 headers. Two lab-only changes, both in test/: the Alpine package mirror was blocked, so git and ssh in the deployer image were copied from alpine/git instead of apk add; and the lab's egress proxy re-signs TLS, so the builder got its CA via NODE_EXTRA_CA_CERTS. Not tested: GitHub's own UI and Let's Encrypt (described from their documentation).

Step-by-Step: From Repository to Zero-Downtime Deploys

01
Step one

Project Layout on the Server

one folder · nothing of it lives in your app repository

The deployment stack is its own small project on the Docker host (I use ~/next-app); your Next.js repository stays a normal Next.js project.

Text~/next-app
~/next-app
├── .env                      # Step 1
├── docker-compose.yml        # Step 7
├── Caddyfile                 # Step 7
├── app.env / build.env       # Step 7 (optional, from the .example files)
├── known_hosts               # Step 3
├── secrets/                  # never commit this folder
│   ├── deploy_key            # Step 3
│   └── webhook_secret        # Step 3
├── builder/   Dockerfile  build.sh       # Step 4
├── app-slot/  Dockerfile  slot.sh        # Step 5
├── deployer/  Dockerfile  deploy.mjs     # Step 6
├── hook/      Dockerfile  hook.py        # Step 7 (from the PHP post)
└── tools/     watch-downtime.mjs  send-test-webhook.sh
.env.env
# copy to .env and adjust
SITE_DOMAIN=www.example.com
REPO_URL=git@github.com:your-name/your-next-app.git
REPO_FULL_NAME=your-name/your-next-app
DEPLOY_BRANCH=main
02
Step two

Prepare the Next.js App: Standalone Output and a Health Route

next.config.mjs · app/api/health/route.js · 2 small changes

Two changes in your app repository. output: "standalone" makes next build write .next/standalone/server.js plus only the node_modules files the server really needs — 67 MB for the demo instead of 352 MB for a full install. And a health route that tells the deployer which build is answering, so it can never switch to a slot that still runs the old version. Here is the complete APEX Note Demo App I used for every test:

JavaScriptapp-example/next.config.mjscomplete · 7 lines
/** @type {import('next').NextConfig} */
const nextConfig = {
  output: "standalone",      // .next/standalone/server.js + only the node_modules it needs
  poweredByHeader: false,
};

export default nextConfig;
JavaScriptapp-example/app/api/health/route.jscomplete · 11 lines
// Health check for the deployer: is THIS process up, and which build is it?
export const dynamic = "force-dynamic";

export function GET() {
  return Response.json({
    status: "ok",
    commit: process.env.NEXT_PUBLIC_COMMIT || "dev",
    slot: process.env.APP_SLOT || "-",
    uptime: Math.round(process.uptime()),
  });
}
JSXapp-example/app/page.jscomplete · 42 lines
import { VERSION, INTRO } from "./version";

// Rendered per request: APP_SLOT is read at RUNTIME, NEXT_PUBLIC_* was fixed at BUILD time
export const dynamic = "force-dynamic";

export default function Home() {
  const commit = process.env.NEXT_PUBLIC_COMMIT || "dev";
  const subject = process.env.NEXT_PUBLIC_COMMIT_SUBJECT || "local build";
  const builtAt = process.env.NEXT_PUBLIC_BUILT_AT || "-";
  const slot = process.env.APP_SLOT || "-";
  const rows = [
    ["Status", <span key="s" className="ok">✓ live</span>],
    ["Version", VERSION],
    ["Commit", `${commit} · ${subject}`],
    ["Built at", builtAt],
    ["Served by", `slot ${slot}`],
    ["Rendered", new Date().toISOString().replace("T", " ").slice(0, 19) + " UTC"],
    ["Runtime", `Node.js ${process.versions.node} · Next.js standalone`],
  ];
  return (
    <main>
      <div className="card">
        <div className="stripe"><i /><i /><i /></div>
        <header>
          <span className="brand"><b>APEX</b> Note · Lab</span>
          <span className="badge">Version {VERSION}</span>
        </header>
        <section>
          <h1>APEX Note Demo App</h1>
          <p className="intro">{INTRO}</p>
          <table>
            <tbody>
              {rows.map(([k, v]) => (
                <tr key={k}><th>{k}</th><td>{v}</td></tr>
              ))}
            </tbody>
          </table>
        </section>
      </div>
    </main>
  );
}
JavaScriptapp-example/app/version.jscomplete · 2 lines
export const VERSION = "1";
export const INTRO = "Version 1 — built from main and switched in without downtime.";
JSONapp-example/package.jsoncomplete · 18 lines
{
  "name": "apexnote-demo-app",
  "version": "1.0.0",
  "private": true,
  "scripts": {
    "dev": "next dev",
    "build": "next build",
    "start": "node .next/standalone/server.js"
  },
  "dependencies": {
    "next": "16.4.0",
    "react": "19.3.0",
    "react-dom": "19.3.0"
  },
  "engines": {
    "node": ">=20.9"
  }
}

Note the two kinds of variables in page.js. NEXT_PUBLIC_COMMIT is set by the builder and inlined at build time — it is part of the JavaScript. APP_SLOT comes from the container and is read at runtime, because the page is rendered per request (dynamic = "force-dynamic"). Mixing these up is the most common Next.js deployment bug; How it works inside shows the proof.

03
Step three

Deploy Key, Host Key, Webhook Secret

same as in the PHP post · 3 minutes

This part is identical to the PHP post, where every line is explained. The deploy key is read by the deployer (UID 1000, the node user), the secret by the hook (UID 1001):

Bashon the Docker host, in ~/next-app
mkdir -p secrets
ssh-keygen -t ed25519 -N "" -C "deploy@www.example.com" -f secrets/deploy_key
chown 1000:1000 secrets/deploy_key && chmod 400 secrets/deploy_key       # read by the deployer (UID 1000)
cat secrets/deploy_key.pub                                                # -> GitHub, read-only deploy key

ssh-keyscan -t ed25519 github.com > known_hosts
ssh-keygen -lf known_hosts        # compare with GitHub's published Ed25519 fingerprint

openssl rand -hex 32 > secrets/webhook_secret
chown 1001:1001 secrets/webhook_secret && chmod 400 secrets/webhook_secret   # read by the hook (UID 1001)
cat secrets/webhook_secret        # paste into the GitHub webhook, then clear the terminal

In GitHub: Settings → Deploy keys → paste deploy_key.pub, leave "Allow write access" unchecked. Settings → Webhooks → Payload URL https://www.example.com/_hooks/github, content type application/json, the secret, Just the push event.

04
Step four

The Builder: npm ci + next build, Isolated

build.sh · POSIX sh · no secrets · mem_limit 2g

The builder watches a request file. For each request it copies the exported source into its workspace, runs npm ci (skipped when package.json, package-lock.json and the Node.js version are unchanged), runs next build and assembles an immutable release in /srv/releases/<commit>: the standalone server, .next/static and public. The release is built in a temporary folder and renamed when complete — a slot can never start half a release. Every step is logged; on failure the last 20 lines go to the container log:

Shellbuilder/build.shcomplete · 119 lines
#!/bin/sh
# build.sh - the ONLY container that runs npm. It gets source code, never secrets.
#
#   build.sh          watch /srv/source/request (written by the deployer) and build
#   build.sh once     build the current request once and exit
#
# For every request it runs "npm ci" and "next build" in /build/app and turns
# the standalone output into an immutable release /srv/releases/<commit>/.
# The result goes to /srv/releases/.results/<commit>.json, the full log to
# /srv/releases/.logs/<commit>.log. A broken build never touches a release
# that is running.
#
# Originally developed by S&H Software Solutions - first published on
# APEX Note (www.apexnote.de), October 2026. MIT licence.
set -u
umask 022
SOURCE=/srv/source
RELEASES=/srv/releases
CONTROL=/srv/control
WORK=/build/app
CACHE=/cache
KEEP=${KEEP_RELEASES:-5}
SELF=$(readlink -f "$0")

log() { printf '%s build %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$*"; }
field() { sed -n "s/^$1=//p" "$SOURCE/request" | head -n 1; }

build_once() {
  id=$(field id); sha=$(field sha); short=$(field short); subject=$(field subject)
  case "$sha" in *[!0-9a-f]* | "") log "ignoring invalid request"; return 0 ;; esac
  [ ${#sha} -eq 40 ] || { log "ignoring invalid request"; return 0; }
  mkdir -p "$RELEASES/.results" "$RELEASES/.logs" || return 1
  if grep -qs "\"id\":\"$id\"" "$RELEASES/.results/$sha.json"; then
    return 0                                   # answered before this container restarted
  fi
  logf=$RELEASES/.logs/$sha.log
  started=$(date +%s)
  took=""

  result() {                                   # status step - written atomically
    printf '{"id":"%s","sha":"%s","status":"%s","step":"%s","seconds":%s}\n' \
      "$id" "$sha" "$1" "$2" "$(( $(date +%s) - started ))" > "$RELEASES/.results/.$sha.tmp"
    mv -f "$RELEASES/.results/.$sha.tmp" "$RELEASES/.results/$sha.json"
  }
  step() {                                     # step NAME command... - output into the log
    name=$1; shift
    printf '\n### %s\n' "$name" >> "$logf"
    s0=$(date +%s)
    if "$@" >> "$logf" 2>&1; then
      case "$name" in npm*|next*) took="$took, $name $(( $(date +%s) - s0 ))s" ;; esac
    else
      log "FAILED in step \"$name\" after $(( $(date +%s) - started ))s - last lines of $logf:"
      tail -n 20 "$logf" | sed 's/^/    | /'
      result failed "$name"
      return 1
    fi
  }

  if [ -f "$RELEASES/$sha/server.js" ]; then
    log "release $short already built - reusing it"
    result ok cached
    return 0
  fi
  : > "$logf"
  log "building $short \"$subject\""
  # node_modules is reused only if package.json, package-lock.json and Node.js are unchanged
  deps=$( (cat "$SOURCE/$sha/package.json" "$SOURCE/$sha/package-lock.json"; node -v) 2>/dev/null | sha256sum | cut -c1-16)
  keep=""
  if [ "$(cat /build/deps 2>/dev/null)" = "$deps" ] && [ -d "$WORK/node_modules" ]; then
    rm -rf /build/node_modules.keep && mv "$WORK/node_modules" /build/node_modules.keep && keep=1
  fi
  step "copy source" sh -c 'rm -rf "$1" && cp -a "$2" "$1" && mkdir -p "$1/.next" "$3/next" &&
        ln -s "$3/next" "$1/.next/cache"' - "$WORK" "$SOURCE/$sha" "$CACHE" || return 1
  cd "$WORK" || return 1
  if [ -n "$keep" ]; then
    mv /build/node_modules.keep node_modules
    printf '\n### npm ci skipped - dependencies unchanged (%s)\n' "$deps" >> "$logf"
    took="$took, npm ci skipped"
  else
    rm -f /build/deps
    step "npm ci" npm ci --no-audit --no-fund --cache "$CACHE/npm" --prefer-offline || return 1
    echo "$deps" > /build/deps
  fi
  # NEXT_PUBLIC_* values are inlined into the bundle NOW - they cannot change at runtime
  step "next build" env NEXT_PUBLIC_COMMIT="$short" NEXT_PUBLIC_COMMIT_SUBJECT="$subject" \
       NEXT_PUBLIC_BUILT_AT="$(date -u '+%Y-%m-%d %H:%M:%S UTC')" npm run build || return 1
  step "check output" test -f .next/standalone/server.js || return 1
  tmp=$RELEASES/.tmp-$sha
  step "assemble release" sh -c 'rm -rf "$1" && cp -a .next/standalone "$1" &&
        rm -rf "$1/.next/cache" && cp -a .next/static "$1/.next/static" &&
        if [ -d public ]; then cp -a public "$1/public"; fi' - "$tmp" || return 1
  mv "$tmp" "$RELEASES/$sha"                   # appears complete or not at all
  result ok built
  log "built $short in $(( $(date +%s) - started ))s (${took#, }, release $(du -sh "$RELEASES/$sha" | cut -f1))"

  # keep the newest $KEEP releases - never one that a slot is using
  in_use=$(cat "$CONTROL"/slots/* 2>/dev/null || true)
  ls -1t "$RELEASES" | tail -n +"$((KEEP + 1))" | while read -r old; do
    case " $in_use " in *"$old"*) continue ;; esac
    rm -rf "${RELEASES:?}/$old" "$RELEASES/.logs/$old.log" "$RELEASES/.results/$old.json"
  done
}

case "${1:-watch}" in
  once) build_once ;;
  watch)
    rm -rf "$RELEASES"/.tmp-* 2>/dev/null || true
    log "waiting for build requests in $SOURCE/request"
    last=""
    while :; do
      req=$(cat "$SOURCE/request" 2>/dev/null || true)
      if [ -n "$req" ] && [ "$req" != "$last" ]; then
        last=$req
        "$SELF" once || true                   # a child process: one failed build never stops the loop
      fi
      sleep 1
    done ;;
  *) echo "usage: build.sh [watch|once]" >&2; exit 2 ;;
esac
Dockerfilebuilder/Dockerfilecomplete · 10 lines
# Runs npm ci + next build. Has internet access for npm - and nothing worth stealing:
# no deploy key, no webhook secret, no admin socket, no route to the running app.
FROM node:22-alpine
RUN mkdir -p /srv/source /srv/releases /srv/control /build /cache \
 && chown node:node /srv/source /srv/releases /srv/control /build /cache
COPY --chmod=755 build.sh /usr/local/bin/build.sh
USER node
ENV HOME=/tmp NEXT_TELEMETRY_DISABLED=1 NPM_CONFIG_UPDATE_NOTIFIER=false \
    NODE_OPTIONS=--max-old-space-size=1536
ENTRYPOINT ["build.sh"]
05
Step five

The App Slots: Blue and Green

slot.sh · the same image twice · internal network

Without docker.sock nobody can restart a container — so the containers restart themselves. Each slot runs a tiny supervisor that reads its slot file (/srv/control/slots/blue) and keeps node server.js of that release running. When the file changes it stops the old process (SIGTERM, after 10 seconds SIGKILL) and starts the new release; when node crashes it starts it again:

Shellapp-slot/slot.shcomplete · 52 lines
#!/bin/sh
# slot.sh - keeps "node server.js" of ONE release running in this slot (blue or green).
#
# The deployer chooses the release by writing its commit into
# /srv/control/slots/<slot>. When the file changes, the old process gets
# SIGTERM and the new release starts. If node crashes, it is started again.
# No docker.sock needed: the deployer never restarts containers.
#
# Originally developed by S&H Software Solutions - first published on
# APEX Note (www.apexnote.de), October 2026. MIT licence.
set -u
SLOT=${APP_SLOT:?set APP_SLOT to blue or green}
WANT_FILE=/srv/control/slots/$SLOT
RELEASES=/srv/releases
pid="" running=""

log() { printf '%s slot-%s %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$SLOT" "$*"; }

stop_app() {                                 # SIGTERM, then SIGKILL after 10 s
  if [ -n "$pid" ]; then
    kill -TERM "$pid" 2>/dev/null
    i=0
    while kill -0 "$pid" 2>/dev/null && [ $i -lt 10 ]; do sleep 1; i=$((i + 1)); done
    kill -KILL "$pid" 2>/dev/null
    wait "$pid" 2>/dev/null
    pid=""
  fi
}
trap 'stop_app; exit 0' TERM INT

log "waiting for $WANT_FILE"
while :; do
  want=$(cat "$WANT_FILE" 2>/dev/null || true)
  crashed=""
  if [ -n "$pid" ] && ! kill -0 "$pid" 2>/dev/null; then
    wait "$pid" 2>/dev/null; crashed="exit $?"; pid=""
  fi
  if [ "$want" != "$running" ] || [ -n "$crashed" ]; then
    [ -n "$crashed" ] && log "node stopped ($crashed) - starting it again"
    stop_app
    if [ -n "$want" ] && [ -f "$RELEASES/$want/server.js" ]; then
      log "starting release ${want%"${want#???????}"}"
      (cd "$RELEASES/$want" && PORT=3000 HOSTNAME=0.0.0.0 exec node server.js) &
      pid=$!
    elif [ -n "$want" ]; then
      log "ERROR: release $want not found"
    fi
    running=$want
    [ -n "$crashed" ] && sleep 2              # do not spin on a crash loop
  fi
  sleep 1
done
Dockerfileapp-slot/Dockerfilecomplete · 9 lines
# Runs ONE Next.js release (standalone server.js) - no npm, no git, no build tools needed
FROM node:22-alpine
# same owner as in the builder image - whichever container creates a volume first
RUN mkdir -p /srv/releases /srv/control && chown node:node /srv/releases /srv/control
COPY --chmod=755 slot.sh /usr/local/bin/slot.sh
USER node
ENV NODE_ENV=production NEXT_TELEMETRY_DISABLED=1
EXPOSE 3000
ENTRYPOINT ["slot.sh"]
06
Step six

The Deployer: Fetch, Build, Check, Switch

deploy.mjs · Node.js standard library only · 231 lines

The deployer is the only part that makes decisions. It fetches main into a bare clone, exports the commit with git archive (the builder never sees .git or the key), rings the builder's doorbell and waits for its answer. If the build is good, it writes the commit into the idle slot's file, waits until that slot reports the new commit on /api/health and answers GET / with 200 — and only then writes the slot's name into control/live. Read deploy() from top to bottom; every return before pointProxyTo() leaves the live version alone:

JavaScriptdeployer/deploy.mjscomplete · 231 lines
#!/usr/bin/env node
// deploy.mjs - zero-downtime deploys for a Next.js app (Node.js standard library only).
//
//   deploy.mjs           watch loop: deploy at start, on every doorbell from hook.py
//                        and every FALLBACK_INTERVAL seconds
//   deploy.mjs once      deploy once and exit
//   deploy.mjs status    show live slot, standby slot and the kept releases
//   deploy.mjs rollback  switch back to the standby slot (previous release) and pin it
//   deploy.mjs unpin     remove the pin and deploy the branch again
//
// One deploy = fetch -> export source -> builder runs npm ci + next build ->
// start the release in the IDLE slot -> health check -> switch the proxy file.
// If any step fails, the live slot is never touched.
//
// Originally developed by S&H Software Solutions - first published on
// APEX Note (www.apexnote.de), October 2026. MIT licence.
import { execFileSync } from "node:child_process";
import fs from "node:fs";

const env = process.env;
const REPO_URL = env.REPO_URL || fail("set REPO_URL, e.g. git@github.com:owner/app.git");
const BRANCH = env.DEPLOY_BRANCH || "main";
const FALLBACK = Number(env.FALLBACK_INTERVAL || 300);      // seconds
const BUILD_TIMEOUT = Number(env.BUILD_TIMEOUT || 900);     // seconds
const HEALTH_TIMEOUT = Number(env.HEALTH_TIMEOUT || 60);    // seconds
const SIGNAL = env.REQUEST_FILE || "/signal/request";
const KEY = env.DEPLOY_KEY_FILE || "/run/secrets/deploy_key";
const KNOWN_HOSTS = env.KNOWN_HOSTS_FILE || "/etc/deploy/known_hosts";
const MIRROR = "/srv/git/repo.git";        // bare clone, never visible to anybody else
const SOURCE = "/srv/source";              // exported source for the builder (read-only there)
const RELEASES = "/srv/releases";          // written by the builder, read-only here
const CONTROL = "/srv/control";            // live + slots/<slot> + state.json, read-only for proxy and slots
const OTHER = { blue: "green", green: "blue" };

process.env.GIT_TERMINAL_PROMPT = "0";
process.env.GIT_SSH_COMMAND = `ssh -i ${KEY} -o IdentitiesOnly=yes -o BatchMode=yes ` +
  `-o StrictHostKeyChecking=yes -o UserKnownHostsFile=${KNOWN_HOSTS}`;

function fail(msg) { console.error(`deploy: ${msg}`); process.exit(2); }
const log = (...m) => console.log(new Date().toISOString().slice(0, 19) + "Z deploy", ...m);
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const git = (...args) => execFileSync("git", args, { encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] }).trim();

function writeAtomic(path, text) {         // readers never see half a file
  fs.writeFileSync(`${path}.tmp`, text);
  fs.renameSync(`${path}.tmp`, path);
}
const readState = () => {
  try { return JSON.parse(fs.readFileSync(`${CONTROL}/state.json`, "utf8")); } catch { return {}; }
};
const saveState = (s) => writeAtomic(`${CONTROL}/state.json`, JSON.stringify(s, null, 1) + "\n");

// Short exclusive lock for "start slot + switch" - so a rollback never races a deploy
async function withLock(fn) {
  for (let i = 0; ; i++) {
    try { fs.mkdirSync(`${CONTROL}/.lock`); break; } catch (e) {
      if (e.code !== "EEXIST") throw e;
      if (i === 0) log("waiting for a running switch to finish");
      await sleep(250);
    }
  }
  try { return await fn(); } finally { fs.rmSync(`${CONTROL}/.lock`, { recursive: true, force: true }); }
}

// ---------------------------------------------------------------- the switch
// Caddy reads /srv/control/live on EVERY request ({file./srv/control/live}:3000).
// Switching = one rename(): no reload, no admin API, no connection is dropped.
function pointProxyTo(slot) {
  writeAtomic(`${CONTROL}/live`, `app-${slot}\n`);
}

// ---------------------------------------------------------------- health check of one slot
async function healthy(slot, short, timeoutS = HEALTH_TIMEOUT) {
  const until = Date.now() + timeoutS * 1000;
  let last = "no answer";
  while (Date.now() < until) {
    try {
      const r = await fetch(`http://app-${slot}:3000/api/health`, { signal: AbortSignal.timeout(2000) });
      const h = await r.json();
      if (r.ok && h.commit === short) {
        const page = await fetch(`http://app-${slot}:3000/`, { signal: AbortSignal.timeout(5000) });
        if (page.ok) return true;
        last = `GET / answered ${page.status}`;
      } else last = `health: ${r.status} commit ${h.commit}`;
    } catch (e) { last = e.cause?.code || e.name; }
    await sleep(500);
  }
  log(`slot ${slot} not healthy after ${timeoutS}s (${last})`);
  return false;
}

// ---------------------------------------------------------------- steps
function exportSource(sha) {               // git archive: the builder never sees .git or the key
  if (fs.existsSync(`${SOURCE}/${sha}`)) return;
  const tmp = `${SOURCE}/.tmp-${sha}`;
  fs.rmSync(tmp, { recursive: true, force: true });
  fs.mkdirSync(tmp);
  git("-C", MIRROR, "archive", "--format=tar", `--output=${tmp}.tar`, sha);
  execFileSync("tar", ["-x", "-f", `${tmp}.tar`, "-C", tmp]);
  fs.rmSync(`${tmp}.tar`);
  fs.renameSync(tmp, `${SOURCE}/${sha}`);
  for (const old of fs.readdirSync(SOURCE).filter((d) => /^[0-9a-f]{40}$/.test(d) && d !== sha)) {
    fs.rmSync(`${SOURCE}/${old}`, { recursive: true, force: true });   // keep only the newest export
  }
}

async function build(sha, short, subject) {  // ring the builder's doorbell and wait for its answer
  const id = `${Date.now()}-${process.pid}`;
  writeAtomic(`${SOURCE}/request`, `id=${id}\nsha=${sha}\nshort=${short}\nsubject=${subject.replace(/[\r\n]/g, " ")}\n`);
  const until = Date.now() + BUILD_TIMEOUT * 1000;
  while (Date.now() < until) {
    try {
      const r = JSON.parse(fs.readFileSync(`${RELEASES}/.results/${sha}.json`, "utf8"));
      if (r.id === id) return r;
    } catch { /* not there yet */ }
    await sleep(1000);
  }
  return { status: "failed", step: `timeout after ${BUILD_TIMEOUT}s`, seconds: BUILD_TIMEOUT };
}

async function deploy() {
  let state = readState();
  if (state.pinned) { log(`pinned to ${state.live.short} - not deploying (run: deploy.mjs unpin)`); return true; }
  const t0 = Date.now();
  if (!fs.existsSync(MIRROR)) {
    log(`first run: cloning branch ${BRANCH}`);
    fs.rmSync(`${MIRROR}.tmp`, { recursive: true, force: true });
    git("clone", "--quiet", "--bare", "--single-branch", "--branch", BRANCH, REPO_URL, `${MIRROR}.tmp`);
    fs.renameSync(`${MIRROR}.tmp`, MIRROR);
  }
  git("-C", MIRROR, "fetch", "--quiet", "--prune", "origin", `+refs/heads/${BRANCH}:refs/heads/${BRANCH}`);
  const sha = git("-C", MIRROR, "rev-parse", "--verify", `refs/heads/${BRANCH}^{commit}`);
  const short = sha.slice(0, 7);
  const subject = git("-C", MIRROR, "log", "-1", "--format=%s", sha);
  fs.writeFileSync(`${CONTROL}/last_ok`, String(Date.now()));       // read by the health check

  if (state.live?.sha === sha) { log(`already live: ${short} on ${state.live.slot}`); return true; }
  if (state.failed === sha) { log(`${short} failed before - waiting for a new commit`); return true; }

  log(`new commit ${short} "${subject}" - building`);
  exportSource(sha);
  const res = await build(sha, short, subject);
  const live = state.live ? `${state.live.short} on ${state.live.slot}` : "nothing";
  if (res.status !== "ok") {
    state = readState(); state.failed = sha; saveState(state);
    log(`ERROR: build of ${short} failed in step "${res.step}" after ${res.seconds}s - ${live} stays live`);
    log(`       full log: docker compose exec builder cat ${RELEASES}/.logs/${sha}.log`);
    return false;
  }
  log(`build ok in ${res.seconds}s (${res.step})`);

  return withLock(async () => {
    state = readState();
    if (state.pinned) { log("pinned meanwhile - not switching"); return true; }
    const slot = state.live ? OTHER[state.live.slot] : "blue";
    writeAtomic(`${CONTROL}/slots/${slot}`, sha);                   // slot.sh starts the release
    log(`starting ${short} in slot ${slot}, waiting for /api/health`);
    if (!(await healthy(slot, short))) {
      state.failed = sha; saveState(state);
      log(`ERROR: ${short} did not become healthy in slot ${slot} - ${live} stays live`);
      if (state.standby?.slot === slot) {                          // keep the rollback target alive
        writeAtomic(`${CONTROL}/slots/${slot}`, state.standby.sha);
        log(`slot ${slot} goes back to ${state.standby.short} (standby for rollback)`);
      }
      return false;
    }
    pointProxyTo(slot);
    state.standby = state.live || null;
    state.live = { slot, sha, short, subject, at: new Date().toISOString() };
    delete state.failed;
    saveState(state);
    log(`live: ${short} "${subject}" on ${slot} in ${Math.round((Date.now() - t0) / 1000)}s ` +
      `(build ${res.seconds}s)`);
    return true;
  });
}

async function rollback() {
  return withLock(async () => {
    const state = readState();
    const prev = state.standby;
    if (!prev) { log("no standby release to roll back to"); return false; }
    if (!(await healthy(prev.slot, prev.short, 10))) { log(`standby slot ${prev.slot} is not healthy`); return false; }
    pointProxyTo(prev.slot);
    [state.live, state.standby] = [prev, state.live];
    state.pinned = true;
    saveState(state);
    log(`rolled back to ${prev.short} on ${prev.slot} and pinned - run 'deploy.mjs unpin' to deploy ${BRANCH} again`);
    return true;
  });
}

function status() {
  const s = readState();
  const line = (k, v) => console.log(`${k.padEnd(9)}${v ? `${v.short}  slot ${v.slot}  ${v.at}  ${v.subject}` : "-"}`);
  line("live:", s.live);
  line("standby:", s.standby);
  if (s.pinned) console.log("pinned:  yes (deploy.mjs unpin)");
  if (s.failed) console.log(`failed:  ${s.failed.slice(0, 7)} (build or health check failed)`);
  const rel = fs.readdirSync(RELEASES).filter((d) => /^[0-9a-f]{40}$/.test(d));
  console.log(`releases: ${rel.map((d) => d.slice(0, 7)).join(" ") || "-"}`);
}

async function watch() {
  fs.mkdirSync(`${CONTROL}/slots`, { recursive: true });
  fs.rmSync(`${CONTROL}/.lock`, { recursive: true, force: true });   // left over from a crash
  log(`watching ${SIGNAL} - branch ${BRANCH}, fallback sync every ${FALLBACK}s`);
  let last = null, next = 0;
  for (;;) {
    let req = "";
    try { req = fs.readFileSync(SIGNAL, "utf8"); } catch { /* no doorbell yet */ }
    if (req !== last || Date.now() >= next) {
      if (last !== null && req !== last) log(`doorbell: ${req.trim()}`);
      last = req;
      next = Date.now() + FALLBACK * 1000;
      try { await deploy(); } catch (e) {
        log(`ERROR: ${String(e.stderr || e.message).trim()} - live slot unchanged, retry in 30s`);
        next = Date.now() + 30000;
      }
    }
    await sleep(1000);
  }
}

const cmd = process.argv[2] || "watch";
const run = { watch, once: deploy, rollback, status, unpin: async () => {
  const s = readState(); delete s.pinned; delete s.failed; saveState(s); return deploy();
} }[cmd];
if (!run) fail("usage: deploy.mjs [watch|once|status|rollback|unpin]");
Promise.resolve(run()).then((ok) => process.exit(ok === false ? 1 : 0),
  (e) => { log(`ERROR: ${String(e.stderr || e.message).trim()}`); process.exit(1); });
Dockerfiledeployer/Dockerfilecomplete · 10 lines
# Has the deploy key and decides which slot is live - but no published port and never runs npm
FROM node:22-alpine
RUN apk add --no-cache git openssh-client \
 && mkdir -p /srv/git /srv/source /srv/control/slots /srv/releases /etc/deploy \
 && chown -R node:node /srv/git /srv/source /srv/control /srv/releases \
 && mkdir /signal && chown 1001:1001 /signal       # owner = hook (UID 1001), whoever mounts it first
COPY --chmod=755 deploy.mjs /usr/local/bin/deploy.mjs
USER node
ENV HOME=/tmp
ENTRYPOINT ["deploy.mjs"]
07
Step seven

Caddy, docker-compose.yml — and Start It

7 services · 4 networks · only 80/443 public

The interesting line in the Caddyfile is reverse_proxy {file./srv/control/live}:3000: Caddy's {file.*} placeholder reads the file on every request, so the upstream changes the moment the deployer renames a new live file into place. The lb_try_duration lines make Caddy retry the dial for up to five seconds instead of answering 502 if node is restarting. The hook is the one from the PHP post, unchanged:

YAMLdocker-compose.ymlcomplete · 134 lines
# Zero-downtime Next.js deploys: proxy / hook / deployer / builder / app-blue + app-green
name: nextapp

x-hardening: &hardening
  init: true
  restart: unless-stopped
  read_only: true
  tmpfs: [/tmp]
  cap_drop: [ALL]
  security_opt: ["no-new-privileges:true"]

x-slot: &slot
  <<: *hardening
  build: ./app-slot
  volumes:
    - releases:/srv/releases:ro              # runs a release, can never change one
    - control:/srv/control:ro
  env_file:                                  # RUNTIME variables (DATABASE_URL, API keys ...) - read per request
    - path: ./app.env
      required: false
  mem_limit: 512m
  healthcheck:
    test: ["CMD-SHELL", "wget -q -O /dev/null http://127.0.0.1:3000/api/health || exit 1"]
    interval: 30s
    timeout: 3s
  networks: [app]                            # internal network: no internet, no hook, no builder

services:
  proxy:
    image: caddy:2.11-alpine
    restart: unless-stopped
    ports:
      - "80:80"
      - "443:443"
      - "443:443/udp"
    environment:
      SITE_DOMAIN: ${SITE_DOMAIN:?set SITE_DOMAIN in .env}
    volumes:
      - ./Caddyfile:/etc/caddy/Caddyfile:ro
      - caddy_data:/data
      - caddy_config:/config
      - control:/srv/control:ro              # "live" = which slot gets the traffic
    depends_on: [deployer]                   # the deployer image creates /srv/control with the right owner
    networks: [front, app]

  hook:                                      # unchanged from the PHP post
    <<: *hardening
    build: ./hook
    environment:
      HOOK_PATH: /_hooks/github
      DEPLOY_BRANCH: ${DEPLOY_BRANCH:-main}
      REPO_FULL_NAME: ${REPO_FULL_NAME:-}
    secrets: [webhook_secret]
    volumes:
      - signal:/signal                       # the only thing it can write to
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request as u; u.urlopen('http://127.0.0.1:9000/healthz', timeout=2)"]
      interval: 30s
      timeout: 3s
    networks: [front]

  deployer:
    <<: *hardening
    build: ./deployer
    environment:
      REPO_URL: ${REPO_URL:?set REPO_URL in .env}
      DEPLOY_BRANCH: ${DEPLOY_BRANCH:-main}
      FALLBACK_INTERVAL: "300"
      BUILD_TIMEOUT: "900"
      HEALTH_TIMEOUT: "60"
    secrets: [deploy_key]
    volumes:
      - git:/srv/git
      - source:/srv/source
      - releases:/srv/releases:ro
      - control:/srv/control
      - signal:/signal:ro
      - ./known_hosts:/etc/deploy/known_hosts:ro
    healthcheck:
      test: ["CMD-SHELL", "find /srv/control/last_ok -mmin -15 | grep -q ."]
      interval: 60s
      timeout: 5s
    networks: [app, git]

  builder:
    <<: *hardening
    build: ./builder
    environment:
      KEEP_RELEASES: "5"
    env_file:                                 # BUILD-time variables: NEXT_PUBLIC_* end up in the browser bundle
      - path: ./build.env
        required: false
    volumes:
      - source:/srv/source:ro
      - releases:/srv/releases
      - control:/srv/control:ro               # to never prune a release a slot is using
      - build:/build
      - build_cache:/cache                    # npm cache + .next/cache survive between builds
    mem_limit: 2g                             # next build needs RAM - see "memory" in the post
    networks: [build]

  app-blue:
    <<: *slot
    environment:
      APP_SLOT: blue

  app-green:
    <<: *slot
    environment:
      APP_SLOT: green

secrets:
  webhook_secret:
    file: ./secrets/webhook_secret
  deploy_key:
    file: ./secrets/deploy_key

volumes:
  git:
  source:
  releases:
  control:
  signal:
  build:
  build_cache:
  caddy_data:
  caddy_config:

networks:
  front:                                     # proxy + hook
  app:                                       # proxy + deployer + slots, no internet
    internal: true
  git:                                       # deployer -> GitHub
  build:                                     # builder -> npm registry
CaddyfileCaddyfilecomplete · 27 lines
{$SITE_DOMAIN} {
	encode zstd gzip

	# Only POST to this one path reaches the webhook listener
	@hook {
		method POST
		path /_hooks/github
	}
	handle @hook {
		request_body {
			max_size 5MB
		}
		reverse_proxy hook:9000
	}

	# Blue/green switch: the deployer writes "app-blue" or "app-green" into this
	# file with one atomic rename. Caddy reads it on every request - no reload.
	handle {
		reverse_proxy {file./srv/control/live}:3000 {
			# node restarting in the live slot? retry the dial instead of answering 502
			lb_try_duration 5s
			lb_try_interval 250ms
		}
	}

	log
}
.envapp.env.examplecomplete · 4 lines
# Runtime variables for app-blue and app-green (server-side only, never in the browser bundle).
# Changing them needs a restart of the slots: docker compose up -d app-blue app-green
# STORE_API_URL=https://api.example.com/stores
# DATABASE_URL=postgres://app:<PASSWORD>@db:5432/app
.envbuild.env.examplecomplete · 3 lines
# Build-time variables for the builder. NEXT_PUBLIC_* values are inlined into the JavaScript
# that every visitor downloads - never put a secret here.
# NEXT_PUBLIC_MAPS_STYLE=light
Pythonhook/hook.pycomplete · 142 lines
#!/usr/bin/env python3
"""hook.py - a tiny GitHub webhook listener (Python standard library only).

It checks the X-Hub-Signature-256 HMAC, accepts push events for one branch
and then only "rings the doorbell": it writes a small request file that
sync.sh watches. It never runs git and never trusts the payload for WHAT to
deploy - sync.sh always fetches the branch from the repository itself.

Originally developed by S&H Software Solutions - first published on
APEX Note (www.apexnote.de), October 2026. MIT licence.
"""
import hashlib
import hmac
import json
import os
import signal
import sys
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from urllib.parse import parse_qs

PORT = int(os.environ.get("HOOK_PORT", "9000"))
HOOK_PATH = os.environ.get("HOOK_PATH", "/hooks/github")
BRANCH = os.environ.get("DEPLOY_BRANCH", "main")
REPO = os.environ.get("REPO_FULL_NAME", "")          # optional: "owner/repo"
REQUEST_FILE = os.environ.get("REQUEST_FILE", "/signal/request")
MAX_BODY = int(os.environ.get("MAX_BODY_BYTES", str(5 * 1024 * 1024)))
SECRET_FILE = os.environ.get("WEBHOOK_SECRET_FILE", "/run/secrets/webhook_secret")


def log(msg):
    print(time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), "hook", msg, flush=True)


try:
    with open(SECRET_FILE, "rb") as f:
        SECRET = f.read().strip()            # "echo secret > file" adds a newline
except OSError as e:
    sys.exit(f"hook: cannot read webhook secret: {e}")
if len(SECRET) < 16:
    sys.exit("hook: webhook secret is empty or shorter than 16 characters")


def ring_doorbell(after, delivery):
    """Atomically replace the request file - sync.sh never reads half a file."""
    tmp = f"{REQUEST_FILE}.{os.getpid()}.{threading.get_ident()}.tmp"
    with open(tmp, "w", encoding="utf-8") as f:
        json.dump({"after": after, "delivery": delivery, "at": int(time.time())}, f)
    os.replace(tmp, REQUEST_FILE)


class Hook(BaseHTTPRequestHandler):
    server_version = "hook"
    sys_version = ""
    timeout = 10                             # slow or stuck clients are dropped

    def log_message(self, *args):            # one line per delivery is enough
        pass

    def reply(self, code, text):
        body = (text + "\n").encode()
        self.send_response(code)
        self.send_header("Content-Type", "text/plain; charset=utf-8")
        self.send_header("Content-Length", str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def do_GET(self):
        if self.path == "/healthz":
            self.reply(200, "ok")
        else:
            self.reply(404, "not found")

    def do_POST(self):
        delivery = self.headers.get("X-GitHub-Delivery", "-")[:64]
        event = self.headers.get("X-GitHub-Event", "-")[:64]
        try:
            code, text = self.handle_delivery(event, delivery)
        except Exception as e:               # never leak a traceback to the caller
            log(f"internal error: {e!r}")
            code, text = 500, "internal error"
        log(f"{code} event={event} delivery={delivery} {text}")
        self.reply(code, text)

    def handle_delivery(self, event, delivery):
        if self.path.split("?", 1)[0] != HOOK_PATH:
            return 404, "not found"
        try:
            length = int(self.headers.get("Content-Length", ""))
        except ValueError:
            return 411, "length required"
        if length < 0 or length > MAX_BODY:
            return 413, "payload too large"
        body = self.rfile.read(length)

        # 1. Authenticity: HMAC-SHA256 over the RAW body, compared in constant time
        expected = "sha256=" + hmac.new(SECRET, body, hashlib.sha256).hexdigest()
        received = self.headers.get("X-Hub-Signature-256", "")
        if not hmac.compare_digest(expected.encode(), received.encode("utf-8", "replace")):
            return 401, "bad signature"

        # 2. Only push events for the deploy branch
        if event == "ping":
            return 200, "pong"
        if event != "push":
            return 202, f"ignored event {event}"
        try:
            if self.headers.get("Content-Type", "").startswith("application/x-www-form-urlencoded"):
                body = parse_qs(body.decode("utf-8"))["payload"][0]
            data = json.loads(body)
        except (ValueError, KeyError, IndexError):
            return 400, "invalid payload"
        if not isinstance(data, dict):
            return 400, "invalid payload"
        repo = (data.get("repository") or {}).get("full_name", "")
        if REPO and repo != REPO:
            return 403, f"wrong repository {repo}"
        ref = str(data.get("ref", ""))
        if ref != f"refs/heads/{BRANCH}":
            return 202, f"ignored {ref[:80]}"
        if data.get("deleted"):
            return 202, "ignored branch deletion"

        # 3. Ring the doorbell - the actual deploy runs in sync.sh
        after = str(data.get("after", ""))[:40]
        ring_doorbell(after, delivery)
        return 202, f"deploy queued for {after[:7]}"


def main():
    signal.signal(signal.SIGTERM, lambda *_: sys.exit(0))
    os.umask(0o022)
    server = ThreadingHTTPServer(("0.0.0.0", PORT), Hook)
    server.daemon_threads = True
    log(f"listening on :{PORT}{HOOK_PATH} - branch {BRANCH}"
        + (f", repository {REPO}" if REPO else ""))
    server.serve_forever()


if __name__ == "__main__":
    main()
Dockerfilehook/Dockerfilecomplete · 7 lines
# Internet-facing listener: Python standard library only - no git, no ssh, no shell tools needed
FROM python:3.12-alpine
RUN adduser -D -H -u 1001 hook && mkdir /signal && chown hook:hook /signal
COPY --chmod=755 hook.py /usr/local/bin/hook.py
USER hook
EXPOSE 9000
ENTRYPOINT ["hook.py"]
Bashtools/send-test-webhook.shcomplete · 18 lines
#!/usr/bin/env bash
# Send a push event signed exactly like GitHub signs it (X-Hub-Signature-256).
#   usage: send-test-webhook.sh URL SECRET_FILE [branch] [event]
#   extra curl options via CURL_OPTS, e.g. CURL_OPTS="--cacert root.crt"
set -euo pipefail
url=$1 secret_file=$2 branch=${3:-main} event=${4:-push}

payload=$(printf '{"ref":"refs/heads/%s","after":"%s","deleted":false,"repository":{"full_name":"%s"}}' \
  "$branch" "$(openssl rand -hex 20)" "${REPO_FULL_NAME:-your-name/your-site}")
sig=$(printf '%s' "$payload" | openssl dgst -sha256 -hmac "$(cat "$secret_file")" | sed 's/^.*= //')

curl -sS ${CURL_OPTS:-} -X POST "$url" -w '  [HTTP %{http_code}]\n' \
  -H 'Content-Type: application/json' \
  -H 'User-Agent: GitHub-Hookshot/send-test-webhook' \
  -H "X-GitHub-Event: $event" \
  -H "X-GitHub-Delivery: $(cat /proc/sys/kernel/random/uuid)" \
  -H "X-Hub-Signature-256: sha256=$sig" \
  --data-binary "$payload"
Bashon the Docker host, in ~/next-app
cp .env.example .env && nano .env        # SITE_DOMAIN, REPO_URL, REPO_FULL_NAME
cp app.env.example app.env               # optional: runtime variables of your app
docker compose up -d --build
docker compose logs -f deployer builder  # first clone + build, then: deploy live: <commit> ...
Outputfirst startreal run · lab adds Gitea
$ docker compose up -d --build 2>/dev/null
$ docker compose logs -f --no-log-prefix deployer builder | sed "/deploy live:/q"
2026-10-07T20:28:50Z build waiting for build requests in /srv/source/request
2026-10-07T20:28:50Z deploy watching /signal/request - branch main, fallback sync every 120s
2026-10-07T20:28:50Z deploy first run: cloning branch main
2026-10-07T20:28:50Z deploy new commit f384989 "Version 1: first release" - building
2026-10-07T20:28:51Z build building f384989 "Version 1: first release"
2026-10-07T20:29:08Z build built f384989 in 17s (npm ci 10s, next build 6s, release 67.1M)
2026-10-07T20:29:08Z deploy build ok in 17s (built)
2026-10-07T20:29:08Z deploy starting f384989 in slot blue, waiting for /api/health
2026-10-07T20:29:10Z deploy live: f384989 "Version 1: first release" on blue in 20s (build 17s)
$ docker compose ps --format "table {{.Service}}\t{{.Status}}\t{{.Ports}}"
SERVICE     STATUS                             PORTS
app-blue    Up 21 seconds (health: starting)   3000/tcp
app-green   Up 21 seconds (health: starting)   3000/tcp
builder     Up 21 seconds                      
deployer    Up 21 seconds (health: starting)   
gitea       Up 24 seconds                      22/tcp, 127.0.0.1:3300->3000/tcp
hook        Up 21 seconds (health: starting)   9000/tcp
proxy       Up 20 seconds                      127.0.0.1:80->80/tcp, 127.0.0.1:443->443/tcp, 443/udp, 2019/tcp

The first deploy has no cache: 17 seconds from clone to live, 10 of them npm ci. Only Caddy publishes ports; the slots show 3000/tcp on the internal network.

08
Step eight

The First Push

git commit · git push · measured: 9,657 ms

Version 1 is live in slot blue. Commit, build time and slot come from the deployment itself:

Chromium showing https://www.example.com: the APEX Note Demo App, Version 1, commit f384989 'Version 1: first release', built at 2026-10-07 20:28:59 UTC, served by slot blue, Node.js 22.23.3, Next.js standalone.
Real runFigure 4 — Version 1 of the APEX Note Demo App, served by slot blue. Commit and build time were inlined by the builder; the slot is read at runtime.

Now change the intro text, commit, push — and measure until the page shows the new version:

Terminal: git commit 'Version 2: new intro text', git push and a curl loop that reports the new version live 9657 ms after git push; the logs show hook 202 deploy queued, deploy doorbell, build built 03ebd6e in 3s with npm ci skipped, slot green starting the release and deploy live on green in 7s.
Real runFigure 5 — The real push: live after 9,657 ms. Below, one delivery through all containers — hook, deployer, builder (npm ci skipped, next build 3 s), slot green, switch. Log cropped to this push.

9.7 seconds from git push to the new page. About 2–3 seconds of that is the webhook delivery itself; the build took 3 seconds because the dependencies had not changed. The new version runs in slot green — and slot blue still runs Version 1, ready for a rollback:

Chromium showing https://www.example.com after the push: Version 2 badge, new intro text, commit 03ebd6e 'Version 2: new intro text', served by slot green.
Real runFigure 6 — After the push: Version 2 from slot green. Slot blue still runs Version 1 — that is the instant rollback.

Prove It: Downtime, Broken Build, Rollback

"Zero downtime" is a claim; I wanted a number. watch-downtime.mjs runs N parallel clients with keep-alive connections (the harder case — a reused connection that gets closed fails, a new one would just retry) and counts every request that does not return 200 with a complete page:

JavaScripttools/watch-downtime.mjscomplete · 62 lines
#!/usr/bin/env node
// watch-downtime.mjs - hammer a URL with N parallel keep-alive clients and count
// every request that does not return a complete page. Ctrl+C or the time limit ends it.
//   usage: node watch-downtime.mjs URL SECONDS [CLIENTS]
// A request counts as "ok" only with HTTP 200 and a complete HTML page (</html>).
import http from "node:http";
import https from "node:https";

const [url, seconds = "30", clients = "4"] = process.argv.slice(2);
if (!url) { console.error("usage: node watch-downtime.mjs URL SECONDS [CLIENTS]"); process.exit(2); }
const lib = url.startsWith("https") ? https : http;
const agent = new lib.Agent({ keepAlive: true, maxSockets: Number(clients) });
const start = Date.now(), end = start + Number(seconds) * 1000;
const t = () => ((Date.now() - start) / 1000).toFixed(1).padStart(5) + "s";
let ok = 0, failed = 0, lastOk = start, maxGap = 0, firstFail = null, lastFail = null, last = "";
const versions = new Map(), errors = new Map();

function once() {
  return new Promise((done) => {
    let settled = false;
    const resolve = (r) => { if (!settled) { settled = true; done(r); } };
    const req = lib.get(url, { agent, timeout: 5000 }, (res) => {
      let body = "";
      res.setEncoding("utf8");
      res.on("data", (c) => (body += c));
      res.on("error", (e) => resolve({ error: e.code || e.message }));
      res.on("close", () => resolve({ error: "response cut off" }));   // only if "end" never came
      res.on("end", () => resolve(res.statusCode === 200 && body.includes("</html>")
        ? { version: (body.match(/Version (\d+)/) || [, "?"])[1], slot: (body.match(/slot (blue|green)/) || [, "?"])[1] }
        : { error: `HTTP ${res.statusCode}` }));
    });
    req.on("timeout", () => req.destroy(new Error("timeout")));
    req.on("error", (e) => resolve({ error: e.code || e.message }));
  });
}

async function client() {
  while (Date.now() < end) {
    const r = await once();
    const now = Date.now();
    if (r.error) {
      failed++; firstFail ??= now; lastFail = now;
      errors.set(r.error, (errors.get(r.error) || 0) + 1);
      if (last !== r.error) { console.log(`${t()}  FAIL ${r.error}`); last = r.error; }
      await new Promise((s) => setTimeout(s, 50));
    } else {
      ok++; maxGap = Math.max(maxGap, now - lastOk); lastOk = now;
      const key = `Version ${r.version} (slot ${r.slot})`;
      versions.set(key, (versions.get(key) || 0) + 1);
      if (last !== key) { console.log(`${t()}  ok   ${key}`); last = key; }
    }
  }
}

await Promise.all(Array.from({ length: Number(clients) }, client));
maxGap = Math.max(maxGap, Date.now() - lastOk);          // a site that never answers counts too
console.log(`---\nrequests: ${ok + failed}   ok: ${ok}   failed: ${failed}`);
for (const [k, v] of versions) console.log(`  ${k}: ${v}`);
for (const [k, v] of errors) console.log(`  ${k}: ${v}`);
if (failed) console.log(`site down from +${((firstFail - start) / 1000).toFixed(1)}s to +${((lastFail - start) / 1000).toFixed(1)}s`);
console.log(`longest time without a good answer: ${maxGap} ms`);
agent.destroy();

Deploy and switch under load

Four clients, one real deploy (Version 3) and then three rollbacks and three un-pins in a row — seven switches within seven seconds:

Terminal: watch-downtime.mjs with four keep-alive clients in the background, git push of Version 3, then three rollbacks and three unpins; the result shows the answers switching between Version 2 on slot green and Version 3 on slot blue: 7,543 requests, 7,543 ok, 0 failed, longest time without a good answer 113 ms.
Real runFigure 7 — Seven switches under load — one deploy, three rollbacks, three un-pins: 7,543 requests, 0 failed.

7,543 requests, 0 failed. The longest pause between two good answers was 113 ms — the CPU was busy with next build at that moment. A final regression run with the published code from an empty lab (cold build, one deploy, six switches, one broken build) gave 15,463 requests, 0 failed.

A broken commit

The same JSX typo that killed my draft, pushed to the new setup while the site is under load:

Terminal: a commit with a JSX error is pushed; watch-downtime shows 5,843 requests, 0 failed, all Version 3; the builder log shows FAILED in step next build with Turbopack's error Expected corresponding JSX closing tag for b at app/page.js:14:25; the deployer logs ERROR build of f61d6fa failed in step next build - 8f6cb02 on blue stays live.
Real runFigure 8 — A broken commit: the build fails in 3 seconds with Turbopack's error, the deployer keeps 8f6cb02 live — 5,843 requests, 0 failed. (npm's update notice cropped.)

The builder names the step and shows the compiler error, the deployer says in one line what happened and what stays live, and the clients did not notice anything. The full log stays in /srv/releases/.logs/<commit>.log. The deployer remembers the failed commit, so the 5-minute fallback sync does not rebuild it again and again; the next push (or deploy.mjs unpin) tries again.

Rollback

A release that builds and passes the health check can still be wrong. Because the previous release keeps running in the other slot, rollback is a file write — 0.3 seconds including docker compose exec — and it pins the version so the next fallback sync does not undo it:

Outputstatus · rollback · unpinreal run
$ docker compose exec deployer deploy.mjs status
live:    f3487a4  slot green  2026-10-07T20:33:19.734Z  Version 5: fix JSX
standby: 8f6cb02  slot blue  2026-10-07T20:30:10.188Z  Version 3: deploy under load
releases: 03ebd6e 8f6cb02 f3487a4 f384989
$ time docker compose exec deployer deploy.mjs rollback
2026-10-07T20:33:25Z deploy rolled back to 8f6cb02 on blue and pinned - run 'deploy.mjs unpin' to deploy main again

real	0m0.292s
user	0m0.044s
sys	0m0.037s
$ curl -s https://www.example.com/api/health; echo
{"status":"ok","commit":"8f6cb02","slot":"blue","uptime":202}
$ docker compose exec deployer deploy.mjs status
live:    8f6cb02  slot blue  2026-10-07T20:30:10.188Z  Version 3: deploy under load
standby: f3487a4  slot green  2026-10-07T20:33:19.734Z  Version 5: fix JSX
pinned:  yes (deploy.mjs unpin)
releases: 03ebd6e 8f6cb02 f3487a4 f384989
$ docker compose exec deployer deploy.mjs unpin
2026-10-07T20:33:25Z deploy new commit f3487a4 "Version 5: fix JSX" - building
2026-10-07T20:33:26Z deploy build ok in 0s (cached)
2026-10-07T20:33:26Z deploy starting f3487a4 in slot green, waiting for /api/health
2026-10-07T20:33:26Z deploy live: f3487a4 "Version 5: fix JSX" on green in 1s (build 0s)

How It Works Inside

1. The switch is a file, not a reload — because my first switch dropped requests

My first version switched Caddy through its admin API: post the Caddyfile with the new upstream to /load over a unix socket. Caddy calls that a graceful reload, and for new connections it is. But the reload closes idle keep-alive connections, and a client that sends its next request on such a connection at that exact moment gets a reset. My load test caught it on the first run:

Outputfirst version: switch via Caddy admin APIreal run · console copy
# Console output of my first zero-downtime test (20:21 UTC, 7 Oct 2026).
# At that time the deployer switched Caddy via its admin API: POST /load with the
# Caddyfile (upstream replaced) over a unix socket. Copied from the terminal - this
# run was not recorded with `script`. Command:
#   node tools/watch-downtime.mjs https://www.example.com/ 12 4 &
#   docker compose exec -T deployer deploy.mjs rollback; ...; deploy.mjs unpin
  0.1s  ok   Version 2 (slot green)
  3.2s  FAIL ECONNRESET
2026-10-07T20:21:57Z deploy rolled back to 79c2258 on blue in 55 ms and pinned - run 'deploy.mjs unpin' to deploy main again
  3.2s  ok   Version 2 (slot green)
  3.3s  ok   Version 1 (slot blue)
live:    79c2258  slot blue  2026-10-07T20:20:22.974Z  Version 1: first release
standby: 4aa4dd4  slot green  2026-10-07T20:21:27.554Z  Version 2: new intro text
pinned:  yes (deploy.mjs unpin)
releases: 4aa4dd4 79c2258
2026-10-07T20:22:00Z deploy new commit 4aa4dd4 "Version 2: new intro text" - building
2026-10-07T20:22:01Z deploy build ok in 0s (cached)
2026-10-07T20:22:01Z deploy starting 4aa4dd4 in slot green, waiting for /api/health
2026-10-07T20:22:02Z deploy live: 4aa4dd4 "Version 2: new intro text" on green in 1s (build 0s, switch 53 ms)
  8.2s  FAIL ECONNRESET
  8.2s  ok   Version 1 (slot blue)
  8.3s  ok   Version 2 (slot green)
---
requests: 1880   ok: 1878   failed: 2
  Version 2 (slot green): 1096
  Version 1 (slot blue): 782
  ECONNRESET: 2
site down from +3.2s to +8.2s
longest time without a good answer: 109 ms
# Note: "site down from ... to ..." is first/last failure - there were exactly two
# failed requests, one per config reload. Same test with the file switch: 7,313 ok, 0 failed.

Two failures in 1,880 requests — one per reload. Browsers usually retry a GET on a reused connection silently, but a POST (a form, a server action) is not retried. The fix was simpler than the original: Caddy's {file.*} placeholder reads /srv/control/live on every request, and deploy.mjs replaces that file with writeFileSync + renameSync. Nothing is reloaded, no connection is closed, the admin API is not needed at all, and a restarted Caddy reads the same file — no state to restore. The cost is one cached file read per request.

2. npm runs in a container that has nothing to steal

npm ci installs hundreds of packages, and any of them can run an install script. In my draft that script would have run next to the GitHub token. Here the builder has no deploy key, no webhook secret, no access to the running app or to Caddy — its only network goes to the npm registry, and its only output is a folder that has to pass the health check before it serves anything. The deployer, which holds the key, never runs npm.

3. The health check knows which build it talks to

"Port 3000 answers" is not enough: during the switch the slot may still run the old process for a second. The deployer waits until /api/health reports its commit, then requests GET / — because an app can be healthy and still crash on its main page. I tested exactly that with a commit that throws when STORE_API_URL is missing. It built fine, /api/health said ok, the page answered 500:

Outputbuilds fine, crashes on GET /real run
$ git diff -U0 app/page.js | tail -1; git commit -qam "Version 6: needs STORE_API_URL" && git push -q origin main 2>/dev/null
+  if (!process.env.STORE_API_URL) throw new Error("STORE_API_URL is not set");
$ docker compose logs -f --no-log-prefix deployer | sed -n "/Version 6/,/standby for rollback/p"
2026-10-07T20:33:49Z deploy new commit 6687188 "Version 6: needs STORE_API_URL" - building
2026-10-07T20:33:53Z deploy build ok in 3s (built)
2026-10-07T20:33:53Z deploy starting 6687188 in slot blue, waiting for /api/health
2026-10-07T20:34:54Z deploy slot blue not healthy after 60s (GET / answered 500)
2026-10-07T20:34:54Z deploy ERROR: 6687188 did not become healthy in slot blue - f3487a4 on green stays live
2026-10-07T20:34:54Z deploy slot blue goes back to 8f6cb02 (standby for rollback)
$ curl -s https://www.example.com/api/health; echo
{"status":"ok","commit":"f3487a4","slot":"green","uptime":159}
$ docker compose logs --no-log-prefix --since 5m app-blue | grep -E "slot-blue|STORE_API_URL" | head -4
2026-10-07T20:33:54Z slot-blue starting release 6687188
⨯ Error: STORE_API_URL is not set
⨯ Error: STORE_API_URL is not set
⨯ Error: STORE_API_URL is not set

After 60 seconds the deployer gives up, the live slot is untouched, and the idle slot goes back to the previous release so a rollback target still exists.

4. node_modules, npm cache and .next/cache

The builder keeps three things between builds in volumes: the npm cache, Next.js' .next/cache (linked into each build), and the last node_modules together with a hash of package.json, package-lock.json and node -v. Measured on the demo app:

Situationnpm cinext buildBuild total
First build, empty caches10–11 s6 s17 s
Dependencies changed, npm cache warm8 s2 s12 s
Dependencies unchanged (most pushes)skipped2–3 s3 s
Outputnpm ci with warm cache + peak memoryreal run
$ docker compose logs --no-log-prefix --since 2m builder | grep "Rebuild\|built"
2026-10-07T20:39:06Z build built b787fc8 in 3s (npm ci skipped, next build 2s, release 67.1M)
2026-10-07T20:39:31Z build building 85b5c70 "Rebuild: npm ci with warm cache"
$ docker compose exec builder rm -f /build/deps   # forget node_modules again, keep the npm cache
$ (while :; do docker stats --no-stream --format "{{.MemUsage}}" blog3-builder; done) > /tmp/mem.log &
$ git commit -q --allow-empty -m "Rebuild 2: measure memory" && git push -q origin main 2>/dev/null
$ timeout 120 docker compose logs -f --no-log-prefix --since 1s builder | sed "/built/q"; kill %1; sort -h /tmp/mem.log | tail -1
2026-10-07T20:39:43Z build built 85b5c70 in 12s (npm ci 8s, next build 2s, release 67.1M)
608.2MiB / 2GiB

Most of npm ci is unpacking Next.js' native SWC binaries, not downloading — that is why even a warm cache costs 8 seconds. Skipping it when nothing changed is safe because the hash includes the Node.js version (native modules) and the lock file.

5. NEXT_PUBLIC_* is frozen at build time

I started the live release a second time inside its slot, on port 3001, with both variables "changed at runtime":

Outputsame release, variables changed at runtimereal run
$ docker compose exec app-green sh -c 'cd /srv/releases/$(cat /srv/control/slots/green) && NEXT_PUBLIC_COMMIT=changed-at-runtime APP_SLOT=changed-at-runtime HOSTNAME=127.0.0.1 PORT=3001 timeout 4 node server.js >/dev/null & sleep 2; wget -qO- 127.0.0.1:3001 | grep -oE "<th>(Commit|Served by)</th><td>[^<]*" | sed "s/<[^>]*>/ /g"; wait'
 Commit  f3487a4 · Version 5: fix JSX
 Served by  slot changed-at-runtime
$ docker compose exec app-green sh -c 'cd /srv/releases/$(cat /srv/control/slots/green) && grep -rl "f3487a4" .next | head -3'
.next/server/chunks/ssr/[root-of-the-server]__02u33qrikb078._.js
.next/server/chunks/[root-of-the-server]__0sp20yik1olc-._.js

The commit did not change — the value is a string in the compiled server chunks (second command). APP_SLOT did. Rule of thumb: everything with NEXT_PUBLIC_ goes into build.env (builder) and ends up in every visitor's browser, so never put a secret there; everything else goes into app.env (slots) and is read per request. After changing app.env, recreate the slots — thanks to the dial retry in Caddy the live slot was recreated with 0 failed requests in my second test (longest wait 1 second). In the first attempt one request that was in flight when the old container stopped was cut off — so recreate the live slot in a quiet minute:

Outputrecreate the live slot under loadreal run
$ cat app.env; curl -s https://www.example.com/api/health; echo
STORE_API_URL=https://api.example.com/stores
{"status":"ok","commit":"da42f32","slot":"green","uptime":19}
$ node tools/watch-downtime.mjs https://www.example.com/ 15 4 > /tmp/env.log & sleep 3; docker compose up -d --force-recreate app-green 2>&1 | tail -1; wait; tail -4 /tmp/env.log
 Container blog3-app-green Started 
---
requests: 2514   ok: 2514   failed: 0
  Version 9 (slot green): 2514
longest time without a good answer: 999 ms
$ docker compose exec app-green printenv STORE_API_URL
https://api.example.com/stores
^@

6. Memory: the build needs more than the app

The running app needs about 40–70 MB. The build peaked at about 600 MiB in docker stats for this tiny app; real apps need 1–3 GB. On a small VPS the kernel kills the build — I simulated that with a 256 MB limit:

Outputbuilder limited to 256 MB, then back to 2 GBreal run
$ docker update --memory 256m --memory-swap 256m blog3-builder >/dev/null   # simulate a small VPS
$ git commit -qam "Version 7: STORE_API_URL optional" && git push -q origin main 2>/dev/null
$ timeout 120 docker compose logs -f --no-log-prefix --since 5s deployer builder | sed "/stays live/q"
2026-10-07T20:36:52Z deploy doorbell: {"after": "b787fc8acbc5ff11bcd85fa3b195feb05451a9ef", "delivery": "2f4947b5-fdd8-47fc-9efb-293446ad3e2d", "at": 1791405412}
2026-10-07T20:36:52Z deploy new commit b787fc8 "Version 7: STORE_API_URL optional" - building
2026-10-07T20:36:53Z build building b787fc8 "Version 7: STORE_API_URL optional"
2026-10-07T20:36:55Z build FAILED in step "next build" after 2s - last lines of /srv/releases/.logs/b787fc8acbc5ff11bcd85fa3b195feb05451a9ef.log:
    | 
    | ### copy source
    | 
    | ### npm ci skipped - dependencies unchanged (ef0e40a1b74993f6)
    | 
    | ### next build
    | 
    | > apexnote-demo-app@1.0.0 build
    | > next build
    | 
    | ▲ Next.js 16.4.0 (Turbopack)
    | ✓ Running next.config.mjs took 22ms
    | 
    |   Creating an optimized production build ...
    | Killed
2026-10-07T20:36:55Z deploy ERROR: build of b787fc8 failed in step "next build" after 2s - f3487a4 on green stays live
$ docker inspect blog3-builder --format "OOMKilled={{.State.OOMKilled}} RestartCount={{.RestartCount}}"; dmesg 2>/dev/null | grep -i "killed process" | tail -1
OOMKilled=true RestartCount=0
[ 4512.248865] Memory cgroup out of memory: Killed process 13441 (next-build (v16) total-vm:22931848kB, anon-rss:239392kB, file-rss:125656kB, shmem-rss:0kB, UID:1000 pgtables:2228kB oom_score_adj:0
$ docker update --memory 2g --memory-swap 2g blog3-builder >/dev/null
$ docker compose exec deployer deploy.mjs unpin
2026-10-07T20:39:03Z deploy new commit b787fc8 "Version 7: STORE_API_URL optional" - building
2026-10-07T20:39:06Z deploy build ok in 3s (built)
2026-10-07T20:39:06Z deploy starting b787fc8 in slot blue, waiting for /api/health
2026-10-07T20:39:07Z deploy live: b787fc8 "Version 7: STORE_API_URL optional" on blue in 5s (build 3s)
$ docker compose logs --no-log-prefix --since 30s builder | tail -1; docker stats --no-stream --format "table {{.Name}}\t{{.MemUsage}}" blog3-builder blog3-app-blue blog3-app-green
2026-10-07T20:39:06Z build built b787fc8 in 3s (npm ci skipped, next build 2s, release 67.1M)
NAME              MEM USAGE / LIMIT
blog3-builder     16.61MiB / 2GiB
blog3-app-blue    37.51MiB / 512MiB
blog3-app-green   69.81MiB / 512MiB

Because build and app live in different containers, the out-of-memory kill hit only the builder; the site kept serving. In my draft the same kill would have taken the site down with it. Give the builder its own mem_limit (2 GB here) and keep --max-old-space-size below it, add swap on small servers, or build in CI (see below).

7. Why releases are folders, not Docker images

The textbook answer is "build an image per commit, tag it, roll back to the previous tag". On a single server without a CI platform that needs the Docker API — docker build and docker run — which means docker.sock, which means root on the host for whoever controls the deployer. A standalone release is already self-contained (server, traced node_modules, static files), so an immutable folder per commit gives the same properties: a release never changes, rollback is "run the previous one", the last five are kept. If you build images in GitHub Actions anyway, push them to a registry and let a pull-based agent deploy them — but that is a different post, and I have not tested it here.


Security & Edge Cases

RiskWhat this setup does
Anyone triggers deploysHMAC-SHA256 over the raw body, constant-time compare, 401 without a valid signature (tested below)
Forged payload deploys foreign codeDoorbell pattern: the deployer always fetches main of REPO_URL itself; a payload with a made-up commit only causes already live
docker.sockNot mounted anywhere — the slots restart their own process
Malicious npm packageRuns in the builder: no secrets, no route to app or proxy; its output must pass the health check
Compromised appRead-only filesystem, read-only releases, internal network without internet, no capabilities, non-root
Secrets in the browser bundleOnly build.env reaches the build; runtime secrets live in app.env
Build eats the server's RAMmem_limit: 2g on the builder, 512 MB per slot
Outputsigned like GitHub signsreal run
$ tools/send-test-webhook.sh https://www.example.com/_hooks/github secrets/webhook_secret main
deploy queued for fdbb9e7
  [HTTP 202]
$ openssl rand -hex 32 > /tmp/wrong; tools/send-test-webhook.sh https://www.example.com/_hooks/github /tmp/wrong main
bad signature
  [HTTP 401]
$ curl -s -X POST https://www.example.com/_hooks/github -H "X-GitHub-Event: push" -d "{}" -w "  [HTTP %{http_code}]\n"
bad signature
  [HTTP 401]
$ tools/send-test-webhook.sh https://www.example.com/_hooks/github secrets/webhook_secret feature/x
ignored refs/heads/feature/x
  [HTTP 202]
$ curl -s -o /dev/null -w "GET /_hooks/github -> %{http_code}\n" https://www.example.com/_hooks/github
GET /_hooks/github -> 404
$ sleep 2; docker compose logs --no-log-prefix --since 10s hook deployer
2026-10-07T20:45:16Z hook 202 event=push delivery=d9f1c9a1-c64f-4525-8227-69c2d6729c18 deploy queued for fdbb9e7
2026-10-07T20:45:16Z hook 401 event=push delivery=8e600439-9328-4268-8e09-f97d0e61b929 bad signature
2026-10-07T20:45:16Z hook 401 event=push delivery=- bad signature
2026-10-07T20:45:17Z hook 202 event=push delivery=62af2342-e3e2-4cbf-ab18-0150259b3932 ignored refs/heads/feature/x
2026-10-07T20:45:17Z deploy doorbell: {"after": "fdbb9e7be5a3a1e3d698c7a918a147559ee76271", "delivery": "d9f1c9a1-c64f-4525-8227-69c2d6729c18", "at": 1791405916}
2026-10-07T20:45:17Z deploy already live: da42f32 on green

Edge cases you should know before you rely on it:

  • A crash of the live process — slot.sh restarts node within a second. Without the dial retry my test showed 48 × 502; with lb_try_duration 5s it was 0 failed requests, the slowest answer waited 1 second (output below)
  • Restarting Caddy itself is the one thing that drops connections (42 failed requests, 0.6 s in my test). The deploy never does it; do it in a quiet minute
  • Database migrations — during a switch, old and new code run at the same time. Make migrations backwards-compatible (expand first, contract in a later release)
  • WebSockets and long requests stay on the old slot until it is replaced by the next deploy — that is a feature, but plan for it
  • ISR and the image optimizer want to write into .next/cache; releases are read-only here. For a fully dynamic app like the demo that is fine — if you use ISR, mount a writable cache volume per slot or configure a cache handler
  • More than one server — this is a single-host design; with several hosts, use a registry and an orchestrator
Outputkill -9 the live node processreal run · before/after lb_try_duration
$ node tools/watch-downtime.mjs https://www.example.com/ 12 4 > /tmp/crash.log & sleep 3; docker compose exec -T app-green pkill -9 -f next-server; wait; cat /tmp/crash.log
  0.1s  ok   Version 9 (slot green)
  3.1s  FAIL HTTP 502
  4.1s  ok   Version 9 (slot green)
---
requests: 1977   ok: 1929   failed: 48
  Version 9 (slot green): 1929
  HTTP 502: 48
site down from +3.1s to +3.6s
longest time without a good answer: 1024 ms
$ docker compose logs --no-log-prefix --since 15s app-green | grep slot-
2026-10-07T20:47:15Z slot-green node stopped (exit 137) - starting it again
2026-10-07T20:47:15Z slot-green starting release da42f32
$ node tools/watch-downtime.mjs https://www.example.com/ 15 4 > /tmp/proxy.log & sleep 3; docker compose restart proxy 2>&1 | tail -1; wait; tail -6 /tmp/proxy.log
 Container blog3-proxy Started 
requests: 2822   ok: 2780   failed: 42
  Version 9 (slot green): 2780
  ECONNRESET: 7
  ECONNREFUSED: 35
site down from +3.2s to +3.7s
longest time without a good answer: 623 ms
$ grep -A4 "reverse_proxy {file" Caddyfile
		reverse_proxy {file./srv/control/live}:3000 {
			# node restarting in the live slot? retry the dial instead of answering 502
			lb_try_duration 5s
			lb_try_interval 250ms
		}
$ node tools/watch-downtime.mjs https://www.example.com/ 12 4 > /tmp/crash.log & sleep 3; docker compose exec -T app-green pkill -9 -f next-server; wait; cat /tmp/crash.log
  0.1s  ok   Version 9 (slot green)
---
requests: 1765   ok: 1765   failed: 0
  Version 9 (slot green): 1765
longest time without a good answer: 1052 ms

Troubleshooting: Real Errors and Their Fixes

Every message in this section came out of my lab while building this post.

Killed / OOMKilled=true during next build

The kernel killed the build for lack of memory (inside, point 6). Raise mem_limit of the builder, keep NODE_OPTIONS=--max-old-space-size about 25 % below it, add swap. The site is not affected — fix the limit and run docker compose exec deployer deploy.mjs unpin to retry the same commit.

Error: Turbopack build failed with 1 error

A real compile error in your code. The old version stays live; the builder log names file, line and column. Fix it, push again. If a commit was force-pushed away, the deployer simply builds the new head.

did not become healthy in slot … (GET / answered 500)

The build is fine, but the app fails at runtime — almost always a missing runtime variable. docker compose logs app-blue shows the exception (⨯ Error: STORE_API_URL is not set in my test). Add it to app.env, recreate the slots, deploy.mjs unpin.

npm error SELF_SIGNED_CERT_IN_CHAIN

My first npm view next version inside a container failed with this, because the lab's egress proxy re-signs TLS. If your server sits behind such a proxy, give the builder the proxy's CA (NODE_EXTRA_CA_CERTS, mounted read-only) — never strict-ssl=false.

mkdir: can't create directory '/srv/releases/.results': Permission denied

My very first start. A named volume gets the owner of the directory in the image of the first container that mounts it — and that was the deployer, whose image had /srv/releases owned by root. All images in this post now create every shared directory with the same owner (UID 1000), and the proxy, which cannot be changed, starts after the deployer (depends_on). If you hit it: docker compose down, docker volume rm the affected volume, start again.

dial tcp: lookup 127.0.0.1:9001:80: no such host

From my first {file.*} test. If the upstream is a placeholder, Caddy cannot see a port and adds :80. Write the port outside the placeholder ({file./srv/control/live}:3000) and only the host name into the file.

deploy ERROR: Internal Server Connection Error … Could not read from remote repository

The first clone ran while Gitea was still starting. The deployer retries after 30 seconds instead of waiting for the 5-minute fallback; with GitHub you see this only during outages or when the deploy key is missing (Permission denied (publickey) — see the PHP post).

📋 Copy the code
deploy.mjs + build.sh + slot.sh + docker-compose.yml

Complete code in this post — copy it from Steps 2–7; every code box has a Copy button. MIT licensed and free for personal and commercial projects. Originally developed by S&H Software Solutions – first published on APEX Note (October 2026).

Questions? Leave a comment below.


Final Thoughts

The PHP version of this idea was about security: who may trigger a deploy, and what a reachable .git folder gives away. The Next.js version is about time: as soon as a build sits between git push and a running server, "restart the container" means "take the site offline for as long as the build runs — or forever, if it fails". The answer is not a bigger CI platform; it is three small decisions: build next to the live version, prove the new one works, and make the switch so small that nobody can catch it halfway.

The most useful finding was the one I did not plan for: a "graceful" proxy reload that still dropped one request per switch. Without a load test I would have shipped that and called it zero-downtime. Measure it — watch-downtime.mjs is in the post for exactly that reason.

If you run this with a monorepo, pnpm, or an app that needs ISR, and something does not fit — leave a comment below. I read all of them.

SH
Sajjad Hanifa
Software Developer · S&H Software Solutions · Oracle APEX, PL/SQL, Docker
Docker Next.js Node.js GitHub Webhook Blue/Green Deployment Zero Downtime Caddy Security

 {fullWidth}

Please Select Embedded Mode To Show The Comment System.*

Previous Post Next Post

نموذج الاتصال