[IRRIGOAPI-117] API survives Postgres not ready at boot; no watcher in live stack #15

Merged
chris merged 12 commits from bug/IRRIGOAPI-117 into main 2026-09-27 21:16:00 -06:00
Owner

Ticket

IRRIGOAPI-117 API silently dies if Postgres isn't ready at boot (host reboot race + bun --watch masks crash)

After the 2026-09-26 host reboot, the API hit Postgres 57P03 in verifyMigrations and threw uncaught. bun --watch kept the dead process alive, and with no healthcheck the container sat "Up" with nothing on 9753 for ~3.5h.

Summary

  • waitForDbReady (api/db/wait-for-db-ready.ts): bounded 60s select 1 probe before verifyMigrations. It retries transient SQLSTATE/socket errors (57P03, ECONNREFUSED, Docker DNS ESERVFAIL/ENOTFOUND, …) and throws on timeout or a non-transient error.
  • New shared findPgErrorCause (api/db/pg-error.ts) walks .cause, because Drizzle's db.execute wraps driver errors in DrizzleQueryError without copying .code. It also fixes a latent bug: verify-migrations' "not migrated" (42P01/3F000) branch was unreachable in production.
  • runMain (api/run-main.ts) is now the single startup exit point. Any startup error is logged and ends in exit(1). bootstrap no longer calls process.exit itself.
  • docker-compose: command: ${API_COMMAND:-bun start}. The live stack runs without the file watcher, so a crash exits and restart: unless-stopped recovers it. Worktrees opt into hot reload with API_COMMAND=bun run dev (documented in .env.example and CLAUDE.md).
  • docker-compose: a Bun-fetch healthcheck on 127.0.0.1:9753/health (the image has no curl/wget) with start_period 120s, plus the autoheal=true label.
  • Live-verified on the primary stack:
    • docker restart irrigo-db irrigo-api → the API recovers unaided and reaches healthy.
    • DB stopped → retries for ~60s, then "startup: fatal error, exiting", and the container cycles instead of sitting Up. After DB start it recovers.
    • Tests 1090/1090 and type-check pass.
  • Operational note: the primary stack no longer hot-reloads. Run docker restart irrigo-api after editing api/ code there.

🤖 Generated with Claude Code

## Ticket [IRRIGOAPI-117](http://192.168.2.100:7123/home/browse/IRRIGOAPI-117/) API silently dies if Postgres isn't ready at boot (host reboot race + bun --watch masks crash) After the 2026-09-26 host reboot, the API hit Postgres `57P03` in `verifyMigrations` and threw uncaught. `bun --watch` kept the dead process alive, and with no healthcheck the container sat "Up" with nothing on 9753 for ~3.5h. ## Summary - `waitForDbReady` (`api/db/wait-for-db-ready.ts`): bounded 60s `select 1` probe before `verifyMigrations`. It retries transient SQLSTATE/socket errors (`57P03`, `ECONNREFUSED`, Docker DNS `ESERVFAIL`/`ENOTFOUND`, …) and throws on timeout or a non-transient error. - New shared `findPgErrorCause` (`api/db/pg-error.ts`) walks `.cause`, because Drizzle's `db.execute` wraps driver errors in `DrizzleQueryError` without copying `.code`. It also fixes a latent bug: verify-migrations' "not migrated" (42P01/3F000) branch was unreachable in production. - `runMain` (`api/run-main.ts`) is now the single startup exit point. Any startup error is logged and ends in `exit(1)`. `bootstrap` no longer calls `process.exit` itself. - docker-compose: `command: ${API_COMMAND:-bun start}`. The live stack runs without the file watcher, so a crash exits and `restart: unless-stopped` recovers it. Worktrees opt into hot reload with `API_COMMAND=bun run dev` (documented in `.env.example` and `CLAUDE.md`). - docker-compose: a Bun-fetch healthcheck on `127.0.0.1:9753/health` (the image has no curl/wget) with `start_period` 120s, plus the `autoheal=true` label. - Live-verified on the primary stack: - `docker restart irrigo-db irrigo-api` → the API recovers unaided and reaches healthy. - DB stopped → retries for ~60s, then `"startup: fatal error, exiting"`, and the container cycles instead of sitting Up. After DB start it recovers. - Tests 1090/1090 and type-check pass. - Operational note: the primary stack no longer hot-reloads. Run `docker restart irrigo-api` after editing `api/` code there. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Docker's embedded DNS drops a stopped/not-yet-started service's hostname
from resolution, which surfaces to postgres.js as getaddrinfo ESERVFAIL
rather than ECONNREFUSED. Found via live verification: `docker stop
irrigo-db` + `docker restart irrigo-api` exited instantly instead of
retrying for up to 60s.
db.execute() always wraps the driver's real error in a DrizzleQueryError,
which puts the connection error's `code` on `.cause`, not on itself.
isTransientDbError was checking the wrong object, so every real db.execute
failure (57P03 included) was misclassified as non-transient and rethrown
immediately instead of retried. Found via live verification of Acceptance
criterion #2.
PORT is the host-side port mapping and differs per worktree (9853/9953),
but the app always listens on the hardcoded Config.port (9753). Reading
process.env.PORT in the healthcheck probed a closed port on any worktree
stack, which would cause autoheal to restart-loop the container.
Add api/db/pg-error.ts (findPgErrorCause), a shared helper that walks
err.cause to find the code-carrying node — db.execute()'s DrizzleQueryError
wrapper puts the driver's real code on .cause, not on itself. Use it from
both isTransientDbError (wait-for-db-ready.ts) and the pre-existing
isPostgresMissingObjectError (verify-migrations.ts), which had the same
bug: it checked err.code directly, so the "database not migrated" branch
was dead code against a real Postgres connection in production.

Also: log the code-carrying node's own message on retry (not the
top-level "Failed query: select 1" wrapper message), add
CONNECTION_DESTROYED to the transient socket codes, document why
ENOTFOUND/ESERVFAIL are treated as transient, note that a single in-flight
probe can overrun DB_BOOT_TIMEOUT_MS by postgres.js's connect_timeout, and
drop a stale doc-comment cross-reference. Expanded test coverage: sleep
call assertions, a sleep-clamp case, a non-transient error following
transient retries, the wrapped production error shape through the retry
loop, and pg-error.test.ts's own cyclic/deep-chain cases.
bootstrap.ts's failed-migration-verification branch and the app.listen
catch both called process.exit(1) directly, bypassing runMain's
console.error + exit; now both throw/rethrow so runMain is the only place
that decides how a startup failure is reported and exited.

run-main.ts: default exit to `code => process.exit(code)` instead of the
unbound `process.exit` method reference. Replaced run-main.test.ts's
duplicate "57P03-shaped" rejection case with a synchronous-throw case
(start() throws instead of returning a rejected promise), which exercises
a path `await start()` wasn't otherwise proven to catch.

index.ts: import bootstrap/run-main as siblings (./bootstrap, ./run-main)
per root CLAUDE.md's import convention, not the @/ alias.
State the no-hot-reload behavior once (in the worktree section, where the
opt-in lives) and point to it from Local development / .env.example
instead of restating the "no longer" narrative in three places. Note the
API_COMMAND opt-in on api/CLAUDE.md's `dev` script row.
IRRIGOAPI-117: Compress CLAUDE.md
All checks were successful
plane-sync / sync (pull_request) Successful in 1s
b37d69a2f2
chris merged commit 105dfea9e3 into main 2026-09-27 21:16:00 -06:00
chris deleted branch bug/IRRIGOAPI-117 2026-09-27 21:16:00 -06:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
chris/irrigo!15
No description provided.