[IRRIGOAPI-115] Fix duplicate concurrent daemon tick loops #12
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "bug/IRRIGOAPI-115"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Ticket
IRRIGOAPI-115 — Daemon runs multiple concurrent tick loops, duplicate re-plans and duplicate
scheduling_decisionsrowsRoot cause — it is not
--watchThe ticket suspected hot-reload stacking under
bun --watch. That is not it:bun --watchrestarts the process, so old timers die with it. This reproduces identically underbun run start, so it is a production bug.The defect is a timer leak in
TimerRegistry.setRePlanHandle(api/service/daemon/runtime.ts), which overwrote the stored handle without cancelling the timer it replaced.scheduleNextTickcalls it at the end of every re-plan. A scheduled tick is net-neutral — its timer has already fired before it re-arms. But every operator-triggered re-plan (POST /replan, schedule enable/disable/skip/resume, system enable/disable) runs off-timer and arms an extra tick while the previously-armed one is still live. The pending-timer count therefore grew by exactly one per operator re-plan and reset only on container restart — matching the reported "2, then 7, then 3, cleared on restart" signature.Summary
setRePlanHandlenow takes theClockand clears the handle it replaces, so at most one re-plan timer is ever live.rePlan()go through a promise-chain queue, so an operator re-plan arriving mid-tick can no longer interleave with the running one onzonesRepo.advanceDepletion/scheduleEntriesRepo.replaceForZone. A failed re-plan can't wedge the queue, and the caller still sees its own rejection (soPOST /replankeeps its 502 semantics).shutdownran would re-arm a tick right aftercancelAllTimerscleared it, leaving a live timer holding the process open. Queued re-plans are now dropped andscheduleNextTickis a no-op once shutdown starts.scheduling_decisionsgains a unique index on(zone_id, date)and the repository upserts viaonConflictDoUpdate, so a duplicated re-plan can no longer fan out rows. The row now holds the latest replan's decision for that night — the authoritative answer to "why didn't zone X water on night Y?".0020_tense_iceman.sql: keeps the newest row per(zone_id, date), withidas a deterministic tie-break.HA open failed/Weather API stalestyle titles). These were already failing onmain; the suite is now fully green.Deploy note
verifyMigrationsexits the api container at startup when the schema is behind the codebase, sodocker compose run --rm api bun run db:migratemust be run before the new image boots. That migration is also what performs the duplicate cleanup.Verification
bun --cwd=./api run type-check— cleanbun --cwd=./api test— 1010 pass, 0 fail (was 996 pass / 12 fail onmain)From Claude: resolved merge conflicts — please re-review.
Both branches had independently created a migration numbered
0020. Resolution: kept main's0020_whole_wolfsbane(addsprecipitation_probability_maxtoweather_daily_snapshots) as index 0020, and renumbered our0020_tense_iceman(deduplicates scheduling_decisions rows + adds unique index on zone_id/date) to0021. Snapshot chain and journal updated accordingly.