Zero-Downtime Blue-Green Release Guide
Deploy a new version beside the active version, verify the real public path, switch traffic safely, and keep a fast rollback path.
The core idea
1. Keep the old version serving traffic
2. Deploy the new version to the idle environment
3. Switch traffic only after health and smoke checks pass
4. Keep the old version for a rollback windowA typical path is DNS or CDN, then the reverse proxy, then one active upstream. The active slot may be blue while green is prepared, or the other way around. The proxy should switch the upstream without stopping the old container first.
Pre-release checks
- Record the branch, commit, image tag, and build artifact.
- Confirm no important uncommitted change is being released accidentally.
- Keep keys, environment files, certificates, and private documents out of Git.
- Confirm database migrations are backward compatible with both versions.
- Provide a real
/healthor/readyendpoint. - Verify that the proxy can reload without interrupting connections.
- Write down the rollback command before starting.
Release flow
1. Identify the active slot
2. Deploy the new version to the idle slot
3. Start the idle container
4. Wait for /ready to return 200
5. Run an internal smoke test
6. Point the proxy upstream to the idle slot
7. Reload the proxy
8. Verify /ready, /health, and core pages through the public hostname
9. Keep the old slot until the observation window endsA container being “running” is not release evidence. Verify through the public hostname because users traverse DNS, CDN, TLS, the proxy, static assets, authentication, and backend APIs.
Switching the proxy
Keep one explicit active-upstream configuration. Replace it with the idle slot and reload the proxy; do not stop the old container as part of the switch.
reverse_proxy app-green:3000docker ps --format "table {{.Names}}\t{{.Image}}\t{{.Status}}"
cat caddy-upstreams/app-active.conf
curl -fsS https://example.com/ready
curl -fsS https://example.com/healthFrontend-only updates
For copy, CSS, i18n, or page-only changes, publish the built static directory to a timestamped temporary path, verify its files and permissions, atomically swap it into place, and preserve the previous directory as a backup. Check the actual chunk files and browser console, not only the homepage status code.
Health-check design
Health checks:
- /health: process and dependency liveness
- /ready: instance is prepared for real traffic
- Public smoke test: proxy, TLS, static assets, auth, database, and core APIKeep /health lightweight. Make /ready stricter: it should mean the instance can accept real traffic and its required dependencies are usable.
Rollback
1. If traffic has not switched, stop the new slot
2. If traffic has switched, point the upstream back to the old slot
3. Reload the reverse proxy
4. Verify public /ready, /health, and core flows
5. Restore the previous static-resource directory when needed
6. Record the cause before attempting the same release againNever invent rollback steps during an incident. Keep the old slot and its image or static-resource backup until the observation window has ended.
High-risk cases
- Database migrations are not backward compatible.
- Old and new versions cannot read the same data safely.
- Queue consumers duplicate work or compete for locks.
- Storage layout changes destructively.
- The new process starts but no core smoke test exists.
- Proxy reload behaviour has not been tested under live traffic.
Post-release observation
- Public
/readyand/healthstatus codes. - 4xx and 5xx rates in proxy logs.
- Recent errors from the new slot.
- Core API latency and time to first byte.
- Whether
index.htmlreferences existing new chunks. - Login, payment, and API-call flows from the user perspective.
For API integration, see the complete integration guide; for agent failures, see Codex tool recovery.