[IRRIGOAPI-117] API survives Postgres not ready at boot; no watcher in live stack #15
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "bug/IRRIGOAPI-117"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Ticket
IRRIGOAPI-117 API silently dies if Postgres isn't ready at boot (host reboot race + bun --watch masks crash)
After the 2026-09-26 host reboot, the API hit Postgres
57P03inverifyMigrationsand threw uncaught.bun --watchkept the dead process alive, and with no healthcheck the container sat "Up" with nothing on 9753 for ~3.5h.Summary
waitForDbReady(api/db/wait-for-db-ready.ts): bounded 60sselect 1probe beforeverifyMigrations. It retries transient SQLSTATE/socket errors (57P03,ECONNREFUSED, Docker DNSESERVFAIL/ENOTFOUND, …) and throws on timeout or a non-transient error.findPgErrorCause(api/db/pg-error.ts) walks.cause, because Drizzle'sdb.executewraps driver errors inDrizzleQueryErrorwithout copying.code. It also fixes a latent bug: verify-migrations' "not migrated" (42P01/3F000) branch was unreachable in production.runMain(api/run-main.ts) is now the single startup exit point. Any startup error is logged and ends inexit(1).bootstrapno longer callsprocess.exititself.command: ${API_COMMAND:-bun start}. The live stack runs without the file watcher, so a crash exits andrestart: unless-stoppedrecovers it. Worktrees opt into hot reload withAPI_COMMAND=bun run dev(documented in.env.exampleandCLAUDE.md).127.0.0.1:9753/health(the image has no curl/wget) withstart_period120s, plus theautoheal=truelabel.docker restart irrigo-db irrigo-api→ the API recovers unaided and reaches healthy."startup: fatal error, exiting", and the container cycles instead of sitting Up. After DB start it recovers.docker restart irrigo-apiafter editingapi/code there.🤖 Generated with Claude Code