Harden evidence boundaries and freeze P0 reproducibility - #35
Draft
HERRY423 wants to merge 2 commits into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Annotation routing could disagree with the direct evaluator and ABI, while source-availability booleans could be mistaken for measured support. Stable-clustering requests could proceed without any perturbation results. This change centralizes annotation intake/ceilings and adds a bounded Leiden resampling grid with cell-aligned ARI recomputation, explicit coverage, and caller-declared criteria.
The frozen branch integrates the previously uncommitted evidence-boundary work as well: untrusted tool-receipt factor suppression, connector and lineage contracts, stricter benchmark outcomes, source-bound reports, and the retained BNS interoperability files. Historical study results remain preserved; author-associated studies cannot count as independent evidence. Generated plugin mirrors and validation rollup provenance are synchronized. Rollup synchronization is maintenance, not a new scientific study.
Validation completed locally on Windows/Python 3.13.9: 1291 passed, 8 explicitly skipped across
tests/; Ruff, registry/mirror checks, clean wheel test, and staged whitespace checks passed. The bounded benchmark is 84/84 gating + 14/14 frontier, including 5/5 planted L3 outcomes, explicitly excluding the public flagship suite. Three single-resolution ROBUST cases now expect the more conservative advisory; no claim ceiling was raised. The full suite also ran with the pinned public CITE-seq and extracted Xenium data present. Reports and source/fixture hashes are underreview/p0-2026-09-04/.Final hosted validation on
cc923849783ac1d0b430f5e7ba494fa1cef1b0aapassed: 35 successful checks, with only the PR-disabled GitHub Pages deployment skipped. The supported Python 3.10–3.12 × Linux/macOS/Windows core and scientific matrices are both 9/9. The strict full benchmark, including the public-data track, is 106/106 with 13/13 L3, no failures or skips. See https://github.com/HERRY423/BioNexus/actions/runs/33871774845 . Three additional source-archive regression tests passed after the initial full local suite. Each core/scientific job retains an exact Git source archive and resolved environment snapshot, including failed jobs. Python 3.13 local results are supplemental and do not replace that matrix. Version snapshots do not bundle wheel hashes or ignored datasets.No independent scientific validation, approved empirical calibration, or new clinical authority is claimed. This remains a reviewable draft; main is unchanged.