Skip to content

fix: bug 005 (pilot_report OOM, poller DB reconnect, DB-gated dead-man) + urllib3 2.8.0 - #47

Merged
chrisdpurcell merged 7 commits into
mainfrom
dev
Oct 3, 2026
Merged

chrisdpurcell merged 7 commits into
mainfrom
dev

Conversation

@chrisdpurcell

Copy link
Copy Markdown
Contributor

Summary

  • Bug 005 (production incident 2026-10-03): pilot_report joined a shared ~28 KB raw page payload per snapshot (565 MB through the join) and OOM-killed PostgreSQL in the 4 GiB CT; the poller then failed every job on dead connections until restarted.
    • 982eebe pilot_report reads narrow, streamed value rows; provider classified in SQL.
    • 8ac812a ConnectionHygieneExecutor recycles Django DB connections around every scheduled job.
    • 60a19b4 dead-man push withheld while the database is unreachable (spec §18.5, ADR-0017).
    • 017d8f9 test isolation fix for the shared sync_to_async thread.
  • 99b0942 urllib3 2.7.0 -> 2.8.0 (PYSEC-2026-4175/4176/4177; runtime via scrapy).
  • Docs: F5a proof env drained; eBay x CPU day-6 pilot review evidence; bug 005 record.

No migration, settings, or admission change. eBay x CPU stays the only admitted cell.

Verification

Local full gate (scripts.check): ruff format/check, basedpyright, 3255 passed / 2 opt-in skips, 96% coverage, pip-audit clean, Actor gate 75 passed / 99%.

The snapshot read joined raw_payload with select_related and materialized every row, so ~50 snapshots sharing one page payload each carried a copy of its body. A one-week production window joined 216 payloads into ~565 MB and the Python decode OOM-killed PostgreSQL's container on 2026-10-03.

Snapshots are now read as narrow value tuples, streamed with iterator(), and only the newest row per bucket key is kept. The provider is classified in SQL, which also lets the freshness query group by provider instead of by every distinct raw endpoint. Run reads defer ProviderRun.import_listing_ids, run_output and ScraperRun.error, and stream. Report output is unchanged.
Django recycles connections only at HTTP request boundaries. The poller is not a request handler, so after PostgreSQL restarted on 2026-10-03 the dead connection stayed cached in the sync_to_async thread and every job raised 'the connection is closed' until the service was restarted.

ConnectionHygieneExecutor, the scheduler's default executor, runs close_old_connections before and after each job on the thread that owns the job's ORM connection: through sync_to_async for coroutine jobs, inline in the worker for plain-function jobs. Regression tests sever the driver connection and require the following jobs to succeed, with CONN_MAX_AGE 0 and with persistent connections.
…able

Spec 18.5 defines the push as a dead-man's switch that alerts on the absence of success, and ADR-0017 relies on it to surface a stalled poller off the box. deadman_job pushed on bare process liveness, so on 2026-10-03 the Uptime Kuma monitor stayed green while every collection job failed.

The job now probes the database with SELECT 1 on the jobs' own sync_to_async thread and skips the push when the probe fails. The probe never raises, so pushing resumes as soon as the database is back.
…ery test

The persistent-connection variant left the sync_to_async thread holding a
connection with no close_at, so test_poller_deadman reused it and its
unreachable-port probe never dialed; the failure appeared only in suite order.
@chrisdpurcell
chrisdpurcell merged commit 1208526 into main Oct 3, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant