The weekend's two runs: CoactDetect holds under the budget, and the rest is inside the noise - #664
Merged
Merged
Conversation
…est is inside the noise Both runs finished 2026-09-19 with no errors on disjoint draws, and each report was murderboarded in three blind rounds plus a finding-driven pass. The rounds moved the headline twice, which is why this page waited for the final build rather than the first. Under the shared false-alarm budget the answer is clean and it replicated: CoactDetect is ahead of every net in every fold of both draws. That is the selection with a stated operating constraint, and it is the one that holds. Chosen on F1 alone the two are nearly tied and the sign is not settled. Counting every training, CoactDetect leads on average in both draws; but the best net moves ahead once both sides use the same merge gap, a setting only the coded side was allowed to tune, and in the replicate it leads in most folds with its average pulled down by two failed trainings. The matched-gap leads carry t of only -1.5 to -2.1, and margins this small are at the limit of what the scoring resolves. The replicate supplied the scale that makes any of this checkable: changing only the recordings moves a net by a median of 0.010 F1 and a coded detector by 0.003. The shakedown's leads of +0.011 and +0.016 were exactly that size. Three findings a single pass would have missed. The earlier +0.103 lead is gone even for the untuned nets, so it was not lost to tuning them. The two sides win in different places: the chorus nets take the faintest events when background firing is high, CoactDetect when it is low, which a pooled F1 hides. And the nets fit 10 of 72 training recordings against a coded setting scored on all 72. Binned SCE does beat CoactDetect under the budget in a few folds. Its F1-alone first place does not stand: a 30 s merge gap it moved to in 8 of 8 folds, failing the crowded check in 7 of 8.
syncytium2
force-pushed
the
goals/weekend-two-runs
branch
from
September 19, 2026 16:48
069dbaa to
83486b6
Compare
This was referenced Sep 19, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rewritten against the final report, which #665 landed while this PR was open. Three blind murderboard rounds moved the headline twice, and the first version of this page carried the round-1 reading. Corrected here before it reaches
main.What holds
Under the shared false-alarm budget — 1.6× CoactDetect's own rates — CoactDetect is ahead of every net in every fold of both draws. That is the selection with a stated operating constraint, and it is the one that replicated.
What does not
Chosen on F1 alone the two are nearly tied and the sign is not settled. Counting every training, CoactDetect leads on average in both draws. But the best net moves ahead once both sides use the same merge gap — a setting only the coded side was allowed to tune — and in the replicate it leads in most folds, its average pulled down by two failed trainings. The matched-gap leads carry t of only −1.5 to −2.1, and margins this small are at the limit of what the scoring resolves.
The replicate supplied the scale that makes any of this checkable: changing only the recordings moves a net by a median of 0.010 F1, a coded detector by 0.003. The shakedown's leads of +0.011 and +0.016 were exactly that size — a single run could never have separated them from the draw.
Three findings a single pass would have missed
Corrections to the earlier version of this PR
Still recorded: both workstations re-measured the bench on the de-pinned export on 2026-09-17, every value inside its bootstrap interval (
2120516, not onmain).chorus_norm's training collapse is under diagnosis on WSMIP065; the merge-gap rerun is with WSMIP064.Docs only, two files.
tests/test_index_resolves.py329 passed; sapper,check_quotesandcheck_pipelinesclear.🤖 Generated with Claude Code
https://claude.ai/code/session_017hQX6iESxBsQk7Jan875e3