The fair comparison's report: tuned nets against tuned coded detectors, reviewed in three rounds - #665
Merged
Conversation
…and SCE's lead is the merge-gap artifact Goal 2's run on WSMIP064 (2026-09-18 16:14 to 2026-09-19 05:58, 1,963 jobs, 0 errors). Its summaries are committed beside a report written for a reader new to the project, with eight inline-SVG figures that explain the problem, the bench, the four nets, nested cross-validation, the fold defect on this run's own draws, the two selections and the second draw before any result. The coded search had the context rule but not goal 1's crowded veto, so tools/crowded_check_fair_comparison.py applies it afterwards with goal 1's own scoring (search_all_settings._job over TAIL). CoactDetect's choices pass in all 8 cases; binned SCE's 30 s merge gap fails in 7 of 8, so its first place does not stand. locust has no admissible gated result: the budget refused every candidate in all 4 folds. WSMIP065 raised all three checks. Draft before the murderboard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, the merge-gap ablation, the replicate's numbers tools/fair_comparison_evidence.py, three subcommands, each writing into the run's summary folder: - fold-draws: Figure 4's data from the run's own fit records and, for "before the fix", from bugarach.learn.train.fold_maker. It reproduces the committed fold_draws.json exactly; three murderboard roles found that no committed code had produced it. - merge-gap: the coded detectors tuned their merge gap and the nets' was fixed at 2 s. This rescores every coded choice, and re-decodes every chosen net refit from its saved model, with ONLY the merge gap changed, after first reproducing the run's own scores at each side's own gap. WSMIP065 will run the same code on the second draw. - replicate: the replicate's headline numbers, with source and commit recorded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…rge gap reverses the headline All eleven roles ran against the first build. Their reports are archived verbatim in docs/reviews/fair-comparison-2026-09-19-roles/ (paths replaced; originals in the darkroom). The finding that changes the result is role 4's: the coded detectors tuned their merge gap and the nets' was fixed at 2 s. merge_gap.json measures both directions with only that setting changed, after reproducing the run's own scores (coded exactly, nets within 0.0015 F1). At the run's unequal gaps CoactDetect leads the best net by 0.007 F1; with the gap matched, at 2 s or at 8 s, chorus_norm leads. The page now says this run does not settle which side is better, and that a rerun tuning the nets' merge gap needs no retraining. Also from the review: the earlier comparison ran on a different, retired simulator; binned SCE's lead is its 30 s merge (it falls below CoactDetect at 8 s); one SCE choice passes the crowded check; the nets exceeded the budget on held-out data in 25 of 80 refits, now reported beside the coded side's misses; the coded side is scored on pooled training folds, not the inner rotation; the probe rate, the backgrounds (25th and 75th percentiles), the hit rule (spans), the tube description, the distractor construction and the coded detectors' origins are corrected; the contamination limit cites its record and the open participation item; t values carry the Nadeau-Bengio factor with its scope; tables are numbered; figures redrawn; the hand-drawn architecture figure is gone (draughtsman owns it). Glossary: merge gap, crowded-recording check, shared false-alarm budget, admissible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All eleven round-2 reports, archived as they arrived (paths scrubbed; the originals are in the darkroom). Three pieces of evidence they asked for: - The crowded-recording check now also scores goal 1's own reference, the shipped operating points. No verdict changes, and the page can now say so from a file rather than a sentence. - A breakdown of held-out recall by participation level and background. The pooled F1 hides where the two sides differ: the faintest events, 10% of the ROIs, on the busy background go to the nets, and on the quiet one mostly to CoactDetect. - The replicate's refits below 0.2 F1, read by the same rule as this run's. Two failed fits in one replicate fold carry that draw's whole margin against the nets on F1 alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ect leads; on F1 alone the sign is not settled Round 2 of the murderboard found the round-1 headline half right. The merge gap does reverse the order, but only for choices on F1 alone: under the shared false-alarm budget CoactDetect stays ahead of both chorus nets at every matched gap. And the replicate's "agreement" was two refits that failed to train; set aside by one rule in both draws, the replicate's best net leads on F1 alone in 4 of 4 folds while this run's trails in 4 of 4. The page now says that, and a sentence about the data is printed only while it is true: the builder's claim() stops the build otherwise. Also: - the argument reordered: terms first, the method with its rules (the context window, the crowded check, "admissible" as the glossary defines it) before any number, and the merge gap in its own section after the results it qualifies; - new figures: the whole bench recording, how a call is scored (one call, one event), binned against sliding counting, both draws per fold with failed refits drawn twice, matched gaps for both selections, where the two sides differ (the faintest events, by background), and tuning; - attributions corrected: the nets are the Deep Sets shape, binned SCE is not the 2003 rule, CICADA cited by its DOI, SPIKE-synch's root added; - the tube refit's mechanism corrected (its threshold came from the inner fits), the corrected t factor stated exactly, reproduction bounds stated as measured; - a sapper feedback note for SAP004's blindness to backslashes, and the first author of the CICADA framework paper corrected in the README. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, and the headline held to its own uncertainty Round 3 of the murderboard (eleven roles, blind) found no wrong number in the tables and a page that still said more than its evidence. Fixed here: - The answer leads with every refit counted: CoactDetect ahead on average in both draws and both selections. The best net leads only at a matched merge gap, or fold by fold in the replicate once failed refits are set aside, and that set-aside is now called what it is, a check made after seeing the held-out scores. - Under the budget, the matched-gap lead over chorus_norm is as weak as the F1-alone reversal (t -1.5 to -2.1), and the page now says so; the budget is named as CoactDetect's own rates times a declared 1.6. - The confound narrows: untuned chorus_norm is already within 0.02 of the tuned CoactDetect in both draws, so the earlier lead was not lost to tuning the nets. - What was not made equal gets its own section, and the limits add how the nets were trained (the probe and the distractors labeled negative), the three merge rules, normalization over the whole recording, and the tolerance measured for one side only. - The merge-gap concept moves to section 2 as its own figure; the gap curves get a logarithmic axis and a key; the margins figure drops the connecting lines that read as confidence intervals; new figures for the participation breakdown at all three levels and the crowded check; the failed refits of both draws become a table. - Attributions: CoactDetect's null credited to whole-train and per-cell shifts (Pipa 2008; Bocchio 2020; Dard 2022), Grün part II, the guard to Rohling 1983, SPIKE-synch's window to event synchronization and its detection step to Cecchini 2021, "designed here" read as "not found elsewhere yet". - Glossary: call, firing, background, empty recording, refit, the failed-training signature; merge gap and the budget corrected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…issible wins under the budget, and the last overreach removed A finding-driven pass over every round-3 finding, and role 10 re-run on the final render, both archived in docs/reviews/fair-comparison-2026-09-19-followup/. The pass caught one error the round-3 fixes introduced: section 9 said the crowded-recording check was never run on the replicate. It was. Read from the replicate's own crowded_check.json and selections, binned SCE under the budget is admissible and ahead of CoactDetect in 3 of 4 replicate folds (+0.011 to +0.017 F1), against 1 of 4 in this run. The page and its lede now say so, behind a claim() guard. Also: the contrast between the two selections is stated as consistency (under the budget every net is behind in every fold of both draws, guarded) rather than margin size, which the replicate contradicted; the scorer-resolution margins are computed; the lede's tuning and faint-event sentences are scoped; stale figure numbers are gone from the SVG labels; the ringed dot means one thing on every figure; Figure 8 says where its no-result marks sit; the axis-break bridge in Figure 11 no longer borrows LoCo's style. And a todo: goal 1's published crowded numbers compare two seed sets. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nal checks Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ng-driven check and role 10 on the final build Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e shipped build in its record Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…and says so Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
syncytium2
enabled auto-merge (squash)
September 19, 2026 15:05
… replacement field Figure 3's caption escaped its quotation marks inside an f-string expression, which 3.12 allows and 3.11, the declared floor, does not. tests/test_syntax_floor.py caught it on every leg; the caption now uses typographic quotes and says the same thing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…he new hash Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…fence it missed is fixed Figure 10's caption escaped quotation marks inside an f-string that was itself inside another's replacement field. The scan exempted all f-string literal text, including a nested one's, so it passed on 3.13 and 3.14 and the 3.11 leg alone failed. On 3.11 that text is expression, and a backslash there is a SyntaxError. The scan now flags a backslash in a nested f-string's text and a nested f-string reusing the enclosing quote; both new cases are proved to be 3.11 errors (checked on a real 3.11) and proved caught. Every tool and test in the tree parses on 3.11. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ecord the new hash Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The report on goal 2's weekend run (tuned nets against tuned coded detectors), written for a reader new to the project, and its review.
What the run shows. Held to a shared false-alarm budget (1.6 times CoactDetect's own rates), CoactDetect is ahead of every net in every fold of two independent draws of simulated recordings; binned SCE is the one coded detector that admissibly beats it under the budget in some folds (1 of 4 here, 3 of 4 in the replicate). Chosen on F1 alone the two are nearly tied: counting every refit, CoactDetect is ahead on average in both draws. The best net moves ahead only when both sides share a merge gap (a setting only the coded side tuned), and in the replicate fold by fold once two failed trainings are set aside. The earlier +0.103 lead is gone even for the untuned nets. The two sides win in different places: the chorus nets find more of the faintest events against a busy background, CoactDetect against a quiet one.
What is here
tools/build_fair_comparison_report.py: builds the page from the run's committed files. Every sentence about the data sits behind aclaim()that stops the build if it no longer holds. 14 numbered figures, 4 tables.tools/fair_comparison_evidence.pyandtools/crowded_check_fair_comparison.py: the regenerable evidence (fold draws, the merge-gap ablation, the participation breakdown, the replicate's refits, goal 1's crowded check against both references).docs/learned/tuned_vs_coact/fair_comparison_2026_09_18/: the run's summaries, the evidence files andreport.html.docs/reviews/fair-comparison-2026-09-19.md: the murderboard record, three blind rounds of all eleven roles, reports archived per round.Where the copies are:
<darkroom>/bugarach/2026-09-18-fair-comparison-run/report/index.html(the page),…/results/(the full run),…/report/review-roles*/(unedited role reports).Not edited:
docs/goals/, per the brief.🤖 Generated with Claude Code