Skip to content

Harden evidence boundaries and freeze P0 reproducibility - #35

Draft
HERRY423 wants to merge 2 commits into
mainfrom
codex/p0-evidence-consistency-freeze
Draft

Harden evidence boundaries and freeze P0 reproducibility#35
HERRY423 wants to merge 2 commits into
mainfrom
codex/p0-evidence-consistency-freeze

Conversation

@HERRY423

@HERRY423 HERRY423 commented Sep 4, 2026

Copy link
Copy Markdown
Owner

Annotation routing could disagree with the direct evaluator and ABI, while source-availability booleans could be mistaken for measured support. Stable-clustering requests could proceed without any perturbation results. This change centralizes annotation intake/ceilings and adds a bounded Leiden resampling grid with cell-aligned ARI recomputation, explicit coverage, and caller-declared criteria.

The frozen branch integrates the previously uncommitted evidence-boundary work as well: untrusted tool-receipt factor suppression, connector and lineage contracts, stricter benchmark outcomes, source-bound reports, and the retained BNS interoperability files. Historical study results remain preserved; author-associated studies cannot count as independent evidence. Generated plugin mirrors and validation rollup provenance are synchronized. Rollup synchronization is maintenance, not a new scientific study.

Validation completed locally on Windows/Python 3.13.9: 1291 passed, 8 explicitly skipped across tests/; Ruff, registry/mirror checks, clean wheel test, and staged whitespace checks passed. The bounded benchmark is 84/84 gating + 14/14 frontier, including 5/5 planted L3 outcomes, explicitly excluding the public flagship suite. Three single-resolution ROBUST cases now expect the more conservative advisory; no claim ceiling was raised. The full suite also ran with the pinned public CITE-seq and extracted Xenium data present. Reports and source/fixture hashes are under review/p0-2026-09-04/.

Final hosted validation on cc923849783ac1d0b430f5e7ba494fa1cef1b0aa passed: 35 successful checks, with only the PR-disabled GitHub Pages deployment skipped. The supported Python 3.10–3.12 × Linux/macOS/Windows core and scientific matrices are both 9/9. The strict full benchmark, including the public-data track, is 106/106 with 13/13 L3, no failures or skips. See https://github.com/HERRY423/BioNexus/actions/runs/33871774845 . Three additional source-archive regression tests passed after the initial full local suite. Each core/scientific job retains an exact Git source archive and resolved environment snapshot, including failed jobs. Python 3.13 local results are supplemental and do not replace that matrix. Version snapshots do not bundle wheel hashes or ignored datasets.

No independent scientific validation, approved empirical calibration, or new clinical authority is claimed. This remains a reviewable draft; main is unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant