fix: end four silent outages — infra, dashboard, duplicate sends, WhatsApp digest - #220
Conversation
Everything reported "running" while nothing worked. Two independent
supervision gaps, both invisible because the process stayed alive:
1. Infra was down for 13 days. dev.lantern.infra is RunAtLoad-only and the
containers had no restart policy, so when Docker Desktop lost the boot
race on Aug 3 the wrapper logged "docker NEVER became ready", exited 1,
and nothing ever retried. Postgres/Redis/MinIO stayed down while every
service looped on connection-refused.
- compose: restart: unless-stopped, so Docker's own supervisor brings
them back whenever the daemon does and launchd is out of the loop.
- infra.plist: KeepAlive{SuccessfulExit:false} + 300s throttle, so a
Docker-not-ready exit retries instead of giving up forever. A
successful one-shot run still exits 0 and is not restarted.
2. The dashboard served a build it never made. Next.js content-hashes chunk
filenames per build and the wrapper deliberately skips rebuilding when a
bundle exists, so a server left running across a rebuild answers HTML
referencing chunks it doesn't have: every /_next/static/chunks/*.js 400s
and the page dies with "Application error: a client-side exception has
occurred". The running process dated from Jul 19; the build from Jul 24.
- dashboard.plist: WatchPaths on .next/BUILD_ID, so a rebuild actually
takes effect.
Verified: infra containers healthy and carrying the restart policy; API
/healthz ok and serving agents/runs from Postgres; every dashboard chunk
200 both on localhost and over the tailnet.
Audit of the owner's real chat.db found the "spammy / repetitive" report is two concrete defects, not a tone problem. 1. Duplicate-send storms — 88 duplicate pairs sent <=300s apart. One fraud alert went out 22x, another 27x, and one reply 15x inside 32 seconds, several within the same second. Cause: autoActLifeEvent has always had a hasActed() guard, but the ping and digest routes that surface a life event to the owner had none, so a re-delivered or re-classified inbound re-sent the same DM with nothing to stop it. The fraud pair alone is 49 of ~1,899 self-chat messages. Fix: dedup where all three routes converge — right after proactiveDecision, reusing the idempotency key already computed for auto-act — rather than bolting a separate guard onto each route. Applied to both bridges. 2. "your code is undefined" — shipped 11x over 36 days. A codeless OTP interpolated f.code straight into the template. Fixed at the source in buildOwnerMessage so no caller can ship it: a codeless OTP now says a code arrived and to check the original message. Audited the sibling templates — f.time and f.merchant were already ternary-guarded; f.code was the only unguarded interpolation. Tests: 1489/1489 pass (2 new regression tests pinning the undefined case and that a real code still renders). Both bridges tsc --noEmit clean. Not fixed here: the self-narrated capability apologies (F3) and the WhatsApp sendSelf outage, both still under audit.
The agent-null fallback fired on every failed round-trip with no throttle, so it apologised on a loop: 11× in production, 8 of them inside 55 seconds. The eighth "give me another minute" tells the owner nothing the first didn't — it is just the noise they reported. Once per 5 minutes per chat now; suppressed repeats log instead of sending. Same fix on both bridges (WhatsApp uses a scalar rather than a per-chat map because that path only ever writes to the owner's own thread). Deliberately NOT changed: the sibling "couldn't generate a reply to <contact>" notice. It already dedups on a 5-minute window (DROP_NOTIFY_DEDUP_MS), so its 4 occurrences across a 36-day window are 4 genuine incidents, not a storm. Silencing real operational drops would trade noise for blindness. Both bridges tsc --noEmit clean; 1489/1489 bridge-core tests pass.
…sendSelf runNewsProactiveTick guarded on `!this.socket`, but sendSelf requires `!this.socket || !this.connected`. On `connection === "close"` the socket OBJECT survives while `connected` flips false, so every 12 minutes the tick cleared its own guard and then threw "not connected" inside sendSelf. Result: 3,238 of 3,246 sendSelf calls failed since Jul 24 — the WhatsApp daily digest and news alerts have been ~100% dead for 23 days while the bridge reported healthy. Contact replies use a different path and were unaffected, which is why it stayed invisible. Two changes: - The fix: make the tick's guard match what sendSelf actually requires. The timer intentionally keeps running, so the feed resumes on its own when the socket returns. - Hygiene: disconnect() cleared eleven timers but missed newsTimer. That alone would NOT have fixed this (the production path is the close handler, which never calls disconnect()), but a leaked interval on teardown is worth closing while here. Checked the siblings: runCommuteTick, runEnergyTick and runHealthCoachTick already test `connected` — news was the lone outlier, matching the stack traces. runFlywheelTick omits it too but fires every 8h with a conditional send, so it is 3/day worst case rather than 120/day; left alone deliberately. tsc --noEmit clean.
🔎 Codex cross-audit (agent:audit)Blocking
Major
Tests not run; this was a read-only audit. |
|
🔒 Reviewer verdict: BLOCKING — New ping/digest dedup reuses the auto-act idempotency key, which for OTP (and partially fraud_alert) is content-blind — it silently suppresses every OTP notification after the first, in both bridges. |
🔷 Gemini cross-audit (agent:audit-gemini)Audit Report:
|
🛡️ Vuln scan — ❌ vulnerable dependency found |
Review caught a real regression in the dedup I added: markActed() ran before
the send, and the send swallows errors (.catch(() => {})). A transient failure
would therefore mark the event surfaced forever — the retry sees hasActed()
and stays silent. That traded "fraud alert sent 27 times" for "fraud alert
possibly sent zero times", which is the worse failure.
Keeping the claim BEFORE the emit is deliberate: two concurrent
classifications of the same inbound must not both pass the check and both
send — that race is the storm. So claim up front, and RELEASE the claim on
every path where nothing actually reached the owner:
- owner DM send failed (both bridges)
- no owner self-chat target / no own JID (both bridges)
Digest keeps its claim: the enqueue is synchronous and cannot fail.
releaseLifeEventClaim() wraps the existing unmarkActed() store primitive,
whose round-trip is already pinned by "unmarkActed lets a key be re-acted
later" in auto-act-ladder.test.ts. The WhatsApp case is not hypothetical —
sendSelf throws whenever the socket is down, the exact failure that took the
digest out for 23 days, so without this the first outage would have
permanently silenced every event it touched.
1489/1489 bridge-core tests pass; both bridges tsc --noEmit clean.
No unit test at the session layer — session.ts has no test harness in this
repo; the store primitive it relies on is covered.
|
🔒 Reviewer verdict: BLOCKING — Digest-route life events now claim their idempotency key up front but never release it on non-delivery (in-memory queue lost on crash/restart), silently and permanently suppressing retries in both bridges — everything else in the diff is solid. |
🛡️ Vuln scan — ❌ vulnerable dependency found |
…no-op The WatchPaths I added to the dashboard plist does not do what I claimed. launchd's WatchPaths only LAUNCHES a job that is not running; it does not restart a running one. The dashboard sets KeepAlive=true, so it is always running and the key was dead weight. Verified rather than assumed: changed BUILD_ID, PID did not move. Replaced with a separate dev.lantern.dashboard-reload job — not KeepAlive, so a BUILD_ID change starts it, it kickstarts the dashboard, and it exits. Verified: touching BUILD_ID moved the dashboard PID 88764 -> 99422 and the tailnet page plus its chunks came back 200. RunAtLoad is false on purpose (otherwise it would kick the dashboard on every login) and ThrottleInterval 30 coalesces a build's writes into one restart. This closes a second symptom of the same stale-build defect, not just the first. NEXT_PUBLIC_API_URL is inlined into the client bundle at BUILD time, so a server left running across a rebuild also serves the OLD API host: the owner's phone was posting login to http://localhost:8080 — itself — and the page reported it as "invalid credentials". Rebuilding with the tailnet host and restarting fixed it; this job stops it recurring.
🛡️ Vuln scan — ❌ vulnerable dependency found |
govulncheck failed the PR on seven stdlib advisories, all reachable from real call paths in the control-plane (TLS handshake via grpc.Server.Serve and http.Server.ListenAndServe, bytes.Buffer.ReadFrom in EmbedText): GO-2026-6088/6089/6090/6091 and GO-2026-5026 and friends, spanning crypto/tls, net/http, net/url, html/template, encoding/asn1, encoding/xml Every one is "Fixed in: <pkg>@go1.26.6" — no dependency is implicated, so the whole fix is the toolchain directive. Bumped 1.26.5 -> 1.26.6 across all 13 modules so they stay consistent; CI resolves its Go via go-version-file: services/control-plane/go.mod and picks this up. Verified: all 13 modules build, and govulncheck on the control-plane (the module whose call paths were flagged) now reports "No vulnerabilities found". Pre-existing on master, not introduced by this branch — fixed here because the branch is what CI gates.
🛡️ Vuln scan — ❌ vulnerable dependency found |
…idden The toolchain bump in the previous commit did not take: the vuln job ran on that exact SHA and still reported "Found in: html/template@go1.26.5". Both workflows set GOTOOLCHAIN: go1.26.5 as job env, and that overrides the go.mod toolchain directive, so CI kept building with the vulnerable stdlib no matter what the modules asked for. Bumped both pins to go1.26.6 and refreshed the vuln comment, which still named the older advisory it was pinned for. Verified locally on the control-plane (the module whose call paths were flagged): govulncheck reports "No vulnerabilities found" on 1.26.6.
|
🔒 Reviewer verdict: BLOCKING — New dashboard-reload launchd job is dead on arrival: install.sh never installs it and never substitutes its UID placeholder. |
Review caught that the new job was dead on arrival for anyone but me: I
installed it by hand with manual placeholder substitution, so it works on this
machine while a fresh `install.sh` run would never create it — and if it
somehow did, the plist would still carry a literal __UID__ in its kickstart
target and fail.
Two one-line gaps, both real:
- __UID__ was not in the sed substitution list (only __NODE__,
__REPO_ROOT__, __HOME__)
- "dashboard-reload" was not in ALL_SERVICES
Nothing else was needed: SRC/DST are derived from the service name, and the
node_modules case-arm correctly skips a job that has no package.
Verified by running the installer's exact substitution into a temp file: the
result lints clean, has zero remaining placeholders, resolves to
gui/501/dev.lantern.dashboard, and diffs IDENTICAL to the hand-installed plist
already proven to restart the dashboard on a BUILD_ID change.
Audit of the owner's real
chat.dband service logs found four defects, each of which failed silently — every affected process reported healthy throughout.1. Infra down 13 days
dev.lantern.infraisRunAtLoad-only and the containers had no restart policy. When Docker Desktop lost the boot race on Aug 3 the wrapper loggeddocker NEVER became ready, exited 1, and nothing retried. Postgres/Redis/MinIO stayed down while every service looped on connection-refused.restart: unless-stopped— Docker's own supervisor owns liveness, launchd is out of the loopinfra.plist:KeepAlive{SuccessfulExit:false}+ 300s throttle so a not-ready exit retries; a successful one-shot run still exits 0 and is not restarted2. Dashboard served a build it never made
Next.js content-hashes chunk filenames per build, and the wrapper deliberately skips rebuilding when a bundle exists. The running process dated from Jul 19; the build from Jul 24. Every
/_next/static/chunks/*.jsreturned 400 and the page died withApplication error: a client-side exception has occurred. Not a Tailscale issue — it failed identically on localhost.dashboard.plist:WatchPathson.next/BUILD_ID3. Duplicate-send storms +
your code is undefined88 duplicate pairs sent ≤300s apart. One fraud alert 22×, another 27×, one reply 15× inside 32 seconds (several in the same second).
autoActLifeEventhad ahasActed()guard; the ping and digest routes had none, so a re-delivered inbound re-sent the same DM unbounded. The fraud pair alone is 49 of ~1,899 self-chat messages.Fixed where all three routes converge rather than per-route. Separately, a codeless OTP interpolated
f.codeand shipped🔑 your code is undefined11× over 36 days — fixed at source inbuildOwnerMessage; siblingsf.time/f.merchantwere already ternary-guarded.Also throttled the agent-null fallback ("give me another minute"), which fired 11×, 8 inside 55 seconds, to once per 5 min per chat.
4. WhatsApp digest/news dead 23 days
runNewsProactiveTickguarded on!this.socket, butsendSelfrequires!this.socket || !this.connected. Onconnection === "close"the socket object survives whileconnectedflips false, so every 12 minutes the tick cleared its own guard and threw insidesendSelf: 3,238 of 3,246 calls failed since Jul 24. Contact replies use a different path, which is why it stayed invisible.Worth noting: the obvious fix (clearing the leaked
newsTimerindisconnect()) would have changed nothing — the production path is the close handler, which never callsdisconnect(). The guard mismatch was the cause. The timer clear is included as teardown hygiene only.Deliberately not fixed
couldn't generate a reply to <contact>(4×) — already deduped on a 5-min window; 4 incidents over 36 days is real signal, and silencing it trades noise for blindnessrunFlywheelTickshares the missingconnectedcheck but fires every 8h with a conditional send — 3/day worst case, not 120/dayVerification
packages/bridge-core: 1489/1489 tests pass (2 new regression tests pinning theundefinedOTP and that a real code still renders)tsc --noEmitclean/healthzserving from Postgres; every dashboard chunk 200 on localhost and over the tailnetNote
These fixes are inert on the live bridges until they are restarted.