Skip to content

Kosson et al. is adjacent, not the same failure, so the novelty sentence comes down - #679

Merged
syncytium2 merged 1 commit into
mainfrom
kosson-prior-art
Sep 20, 2026
Merged

syncytium2 merged 1 commit into
mainfrom
kosson-prior-art

Conversation

@syncytium2

Copy link
Copy Markdown
Owner

Closes the last item the blind round (#674) left outside its own claim: role 2 flagged
arXiv:2410.23922 (Kosson et al., NeurIPS 2024) as uncited prior art sitting under the
page's closing sentence, "That it prevents this collapse is this page's result."

Verdict: adjacent, not the same. The novelty sentence is withdrawn. Both candidate
papers were read, not searched.

Kosson et al. 2024 — adjacent. It walks the same chain (too large an early step →
damage training does not undo → warm-up prevents it), and that is enough to retire an
unqualified novelty claim. It differs in two ways that matter:

Kosson et al. 2024 this page
where the dead-unit result lives §7 + Appendix A, on a small image network the whole page
direction units are alive, the update kills them units are already silent at the starting weights
when measured fraction of dead units at the end of training before the first step, in every replayed fit
activation ReLU — remedy tested is a leaky ReLU GELU — no exact zero region, so that remedy is unavailable and no layer is ever "dead" in his sense, only silent

Lu et al. 2020 (arXiv:1903.06733) is the closer neighbour, and was not on the list.
Born dead means dead before training; their theorem sends a gradient method (Adam
among those they name) to a constant function; and their depth-up / width-down ordering
is the order Table 1 finds. Recomputed here, it holds in all four comparisons at lr 0.03
— 69→90% and 46→81% with depth, 69→46% and 90→81% with width. But their theory is ReLU-only,
and it forbids what this page measures: a born-dead network cannot be recovered, while
these fits do train at a lower rate or behind a 200-step ramp.

What the page says now is narrower than what it said: that a head can start almost silent
in a network with no zero region to be stuck in, that whether training wakes it is what decides
the fit, and that the failure is recoverable — shown by replaying one fit bit for bit. It also
names what was searched and what was not, because this repo has already paid once for a novelty
claim that four web searches failed to check (detector_history.md).

Scope held. Related work and two references only. Nothing else on that page — the ranked
list from #674 is Tony's to rule on and the census re-run needs his GPU decision. No JSON
changed; only prose. I widened by one paper (Lu), because the novelty sentence could not be
rewritten honestly while a closer neighbour sat unread in my own review record.

Checks. tests/test_diagnose_chorus_collapse.py 3 passed; sapper clear; the page rebuilds
byte-identically on Python 3.11, 3.13 and 3.14; the darkroom copy moved with it
(0dfdc1e5… → 9083d574…, matching the repo copy) and the claim is released here.

🤖 Generated with Claude Code

…nce comes down

Role 2 of #674's blind round flagged arXiv:2410.23922 as uncited prior art under
the page's closing claim, "That it prevents this collapse is this page's result".
Both candidate papers were read, not searched.

Kosson et al. 2024 walks the same chain - too large an early step, damage
training does not undo, warm-up prevents it - but in the other direction and in a
different activation. His dead-unit result is an appendix on a small image
network where large updates kill units that were alive, counted at the end of
training, and his remedy is a leaky rectified-linear unit, which works because a
ReLU has an exact zero region to be stuck in. A GELU has none, so that repair is
unavailable here and no layer of this head is dead in his sense, only silent.
This page's head is already silent before the first step, in every replayed fit,
working and collapsed alike.

Lu et al. 2020 turned out to be the closer neighbour and was not on the list.
Born dead means dead before training, their theorem sends a gradient method (Adam
among those they name) to a constant function, and their depth-up / width-down
ordering matches Table 1 in all four comparisons at lr 0.03 - recomputed here as
69 to 90 percent and 46 to 81 percent with depth, 69 to 46 percent and 90 to 81
percent with width. But their theory is ReLU-only and forbids the recovery this
page demonstrates: these fits do train at a lower rate or behind a 200-step ramp.

So the page now cites both, states the difference, and says what is left, which
is narrower than the sentence it replaced. It also names what was searched and
what was not, because this repo has already paid once for a novelty claim that
four web searches failed to check.

Scope held: Related work and two references only. Nothing else on that page - the
ranked list from #674 is Tony's and the census re-run needs his GPU decision. The
data are untouched; only prose moved. The page rebuilds byte-identically on
Python 3.11, 3.13 and 3.14, and the darkroom copy was moved with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@syncytium2
syncytium2 merged commit 0d783f4 into main Sep 20, 2026
3 checks passed
@syncytium2
syncytium2 deleted the kosson-prior-art branch September 20, 2026 22:55
syncytium2 pushed a commit that referenced this pull request Sep 21, 2026
lit/optimization/ now holds Kosson et al. 2024 and Lu et al. 2020 with a README
entry each, in the shelf's own form: author, year, where it came from, and which
decision it bears on. Both PDFs were fetched by hand from arXiv and verified by
extracting the first page after the copy, not by file size.

The entries carry the distinction that made #679's ruling, so a later session
re-reading them does not have to redo it. Kosson's units are alive and the update
kills them, counted at the end of training, and his remedy is a leaky
rectified-linear unit - which is itself the evidence that the mechanism is ReLU's
exact zero region. Lu's born-dead networks are dead before training and his
theorem sends them to a constant function, and his depth-up / width-down ordering
matches Table 1 in all four comparisons at lr 0.03. Neither transfers to a GELU,
which has no zero region, so the shelf says plainly that the ordering match is a
match in direction and not evidence that his mechanism operates here.

The README also records what these papers do NOT settle, including the one search
nobody has run: whether a learned event detector collapsing to a single call
covering a whole recording has been reported in the imaging literature. That is
the remaining way the page's narrowed claim could still turn out to be old.

One correction carried forward: role 2's archived report cites Lu's
constant-function result as Theorem 3.4; in the text it is 3.7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants