Skip to content

The chorus-collapse page's blind round: three of six repairs do not survive it - #674

Merged
syncytium2 merged 4 commits into
mainfrom
chorus-verify
Sep 20, 2026
Merged

syncytium2 merged 4 commits into
mainfrom
chorus-verify

Conversation

@syncytium2

Copy link
Copy Markdown
Owner

Closes the heaviest of the four flags #667 landed with: the round-2 repairs had
never been blind-reviewed. This is that round -- 11 of 11 roles, blind, against
the built page on main.

Verdict: the flag cannot be retired. Three of the six repairs hold, three do not.

round-2 repair verdict
Figure 1 rebuilt as a matched pair holds
Interior census and replays (400-frame trim) holds as a measurement
Figure 5 merged holds
The #596 test replacing the separation measure does not hold
The selection split by rule (Table 2) does not hold for the budgeted column
The replicate report's reading "corrected" does not hold, and is new damage

The diagnosis itself survives, more firmly than before. collapse_table.json
was re-derived from the raw run archives without calling the builder (1,938 rows,
zero differences), every permutation test re-run with an independent generator at
up to 200,000 shuffles, and the nets rebuilt in torch to re-measure the untrained
controls. The page rebuilds byte-identically on Python 3.11, 3.13 and 3.14, so the
two cross-version determinism bugs that bit it before are closed, not dormant.

What failed, and why each is a different kind of failure

  • Figure 4 narrowed a defect instead of removing it. Round 2 replaced a test
    that could not fail. The replacement can fail, and still misses 6 chorus_norm and
    13 chorus_gain_norm fits whose votes are constant on a real recording. Separately,
    the untrained comparator it is set against is the lowest of the grid's four encoder
    shapes by 14x, and contradicts this project's own README at the commit the page cites.
  • Table 2's false-alarm-budget column is wrong under the page's own definition.
    Those two refits have n_detected = 0 at that rule's threshold -- they call
    nothing. The page quotes the sibling report's footnote verbatim, including the
    words "or called nothing", and then uses only the first half.
  • The correction of the sibling report is a straw man, and round 2 introduced it.
    That report already attributes the overlap to shared configurations and seeds.

Underneath them, the finding that is not a wording fix. census.json stores only
the minimum across the eight head layers and never the head's input, so the page cannot
tell a signal that died inside the head from one that never reached it -- and in 10 of
153 chorus_norm and 32 of 75 chorus_gain_norm collapsed fits, it never reached it. The
builder's own docstring promises both missing fields. Closing this needs a census re-run
on a GPU, not an edit.

Why nothing was repaired. Blocking went 3 -> 4 -> 6. The process says a flat or
rising blocking count means escalate rather than patch, because patching converts a
fixable draft into a long tail nobody has reviewed together. So the page on main is
untouched and the record states its fingerprint as unchanged.

Ranked for Tony -- wrong numbers ("1 working fit as dead" is 5; Table 2's column;
three bounds rounded inward and literally false); the straw man; Table 3's outcome
column clipped 211 px at every desktop width with no affordance; and the census re-run.
Two items outside this claim: docs/goals/learned-model-family.md states 112.6/40.4
against this page's 111.5/39.6 (#664 owns that file), and Kosson et al. 2024 is uncited
prior art sitting under the page's novelty claim.

Both gates green: roster 11/11 with reports, grants 11 ok / 0 mismatch. Record and all
11 role reports also delivered to the claimed darkroom folder; the claim is released here.

🤖 Generated with Claude Code

defazio2 and others added 4 commits September 20, 2026 16:00
#667 landed the diagnosis with four open flags, the heaviest being that its
round-2 repairs were never blind-reviewed. This re-claims the darkroom folder
that PR released, because a repair has to be rebuilt into the darkroom copy of
the page as well as the repo one, and starts the role archive for the round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Written verbatim per role, before synthesis, so a finding that synthesis later
merges or softens stays recoverable. Roles 1, 2 and 4 are still running.

Two findings are confirmed against the code rather than taken on the reviewer's
word. The sign rule the page describes as counting one working fit as dead
actually counts five (1 chorus_norm, 4 chorus_gain_norm) and misses three
collapsed chorus_gain_norm fits. And Table 3's outcome column is clipped by
211 px at every desktop width, with its scroll hint suppressed above 700 px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Roles 2, 4 and 7. Role 4 is the consequential one: it re-ran forward passes on
all 231 collapsed fits and made the page's own borrowed vote-response test fail,
and it shows the headline mechanism sentence is not supported by what the census
stores.

Verified here rather than taken on the reviewer's word: census.json keeps only
the minimum share of varying units across the eight head layers, with no
per-layer vector and no head-input field - while the builder's own docstring
promises both. So the page cannot tell a signal that died inside the head from
one that never reached it. Roles 2 and 4 also independently find that the page's
correction of the sibling replicate report addresses a sentence that report does
not contain.

SAP004 blocked the first attempt: role 2 wrote three MCP server names separated
by slashes and the rule matches 'Dropbox/' as a personal path. The rule is not
the thing to change - its own comment records the lowercase-home path that
slipped through a narrower version - so the separators became commas, the
archived report carries a note saying so because it is evidence, and the dispute
is filed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…survive it

#667 landed the diagnosis with four flags, the heaviest being that its round-2
repairs had never been blind-reviewed. This is that round, all 11 roles.

The diagnosis survives, and more firmly than before: collapse_table.json was
re-derived from the raw run archives without calling the builder - 1,938 rows,
zero differences - every permutation test was re-run with an independent
generator at up to 200,000 shuffles, and the nets were rebuilt in torch to
re-measure the untrained controls. The page also rebuilds byte-identically on
Python 3.11, 3.13 and 3.14, so the two cross-version bugs that bit it before are
closed rather than dormant.

Three of the six repairs fail. Figure 4's replacement test narrowed a defect
instead of removing it and still misses 6 chorus_norm and 13 chorus_gain_norm
fits whose votes are constant. Table 2's false-alarm-budget column counts two
refits as collapsed that call nothing at that rule's threshold - the page quotes
the half of the sibling report's footnote it does not use. And the correction of
that sibling report addresses a sentence the report does not contain, which
round 2 introduced.

Underneath them sits the one finding that is not a wording fix: census.json
stores only the minimum across the eight head layers and never the head's input,
so the page cannot tell a signal that died inside the head from one that never
reached it - and in 10 of 153 chorus_norm and 32 of 75 chorus_gain_norm
collapsed fits it never reached it.

Blocking went 3 -> 4 -> 6. Severity is rising, so the round escalated WITHOUT
repairing, per the process: patching here would lay a fourth layer of unreviewed
text over an artifact that needs a census re-run. The page on main is unchanged
and its fingerprint is stated as unchanged in the record.

Both gates green. The grants gate refused the first draft because role 7 named
its hand-off channel inside its GRANT line; the record now states what that
reviewer actually held rather than tidying the line away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@syncytium2
syncytium2 merged commit a201f41 into main Sep 20, 2026
3 checks passed
@syncytium2
syncytium2 deleted the chorus-verify branch September 20, 2026 21:19
syncytium2 added a commit that referenced this pull request Sep 20, 2026
… now agree with the collapse page (#678)

The blind verify round (#674) found docs/goals/ and the chorus-collapse page
disagreeing: 112.6/40.4 against 111.5/39.6. The goal page's numbers were mine
(#668) and they were the worse of the two.

Both estimate the same thing — how many of chorus_norm's inner fits would collapse
in BOTH draws if collapse were independent within a configuration. I pooled each
configuration's rate across the two draws and squared it. The page multiplies each
draw's own rate. Squaring an average is never below multiplying the two values, so
pooling biases the estimate upward, and it is the product that is unbiased when the
two draws may sit at different rates.

Recomputed from the same collapse_table.json: 111.5 against 111 observed, and 39.6
against 38. The reading does not change — the overlap is what chance gives, and the
failure is not shown to be deterministic — but the figure a reader checks is now
the defensible one and the two pages no longer contradict each other.


Claude-Session: https://claude.ai/code/session_017hQX6iESxBsQk7Jan875e3

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants