Kosson et al. is adjacent, not the same failure, so the novelty sentence comes down - #679
Merged
Merged
Conversation
…nce comes down Role 2 of #674's blind round flagged arXiv:2410.23922 as uncited prior art under the page's closing claim, "That it prevents this collapse is this page's result". Both candidate papers were read, not searched. Kosson et al. 2024 walks the same chain - too large an early step, damage training does not undo, warm-up prevents it - but in the other direction and in a different activation. His dead-unit result is an appendix on a small image network where large updates kill units that were alive, counted at the end of training, and his remedy is a leaky rectified-linear unit, which works because a ReLU has an exact zero region to be stuck in. A GELU has none, so that repair is unavailable here and no layer of this head is dead in his sense, only silent. This page's head is already silent before the first step, in every replayed fit, working and collapsed alike. Lu et al. 2020 turned out to be the closer neighbour and was not on the list. Born dead means dead before training, their theorem sends a gradient method (Adam among those they name) to a constant function, and their depth-up / width-down ordering matches Table 1 in all four comparisons at lr 0.03 - recomputed here as 69 to 90 percent and 46 to 81 percent with depth, 69 to 46 percent and 90 to 81 percent with width. But their theory is ReLU-only and forbids the recovery this page demonstrates: these fits do train at a lower rate or behind a 200-step ramp. So the page now cites both, states the difference, and says what is left, which is narrower than the sentence it replaced. It also names what was searched and what was not, because this repo has already paid once for a novelty claim that four web searches failed to check. Scope held: Related work and two references only. Nothing else on that page - the ranked list from #674 is Tony's and the census re-run needs his GPU decision. The data are untouched; only prose moved. The page rebuilds byte-identically on Python 3.11, 3.13 and 3.14, and the darkroom copy was moved with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
syncytium2
pushed a commit
that referenced
this pull request
Sep 21, 2026
lit/optimization/ now holds Kosson et al. 2024 and Lu et al. 2020 with a README entry each, in the shelf's own form: author, year, where it came from, and which decision it bears on. Both PDFs were fetched by hand from arXiv and verified by extracting the first page after the copy, not by file size. The entries carry the distinction that made #679's ruling, so a later session re-reading them does not have to redo it. Kosson's units are alive and the update kills them, counted at the end of training, and his remedy is a leaky rectified-linear unit - which is itself the evidence that the mechanism is ReLU's exact zero region. Lu's born-dead networks are dead before training and his theorem sends them to a constant function, and his depth-up / width-down ordering matches Table 1 in all four comparisons at lr 0.03. Neither transfers to a GELU, which has no zero region, so the shelf says plainly that the ordering match is a match in direction and not evidence that his mechanism operates here. The README also records what these papers do NOT settle, including the one search nobody has run: whether a learned event detector collapsing to a single call covering a whole recording has been reported in the imaging literature. That is the remaining way the page's narrowed claim could still turn out to be old. One correction carried forward: role 2's archived report cites Lu's constant-function result as Theorem 3.4; in the text it is 3.7. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes the last item the blind round (#674) left outside its own claim: role 2 flagged
arXiv:2410.23922 (Kosson et al., NeurIPS 2024) as uncited prior art sitting under the
page's closing sentence, "That it prevents this collapse is this page's result."
Verdict: adjacent, not the same. The novelty sentence is withdrawn. Both candidate
papers were read, not searched.
Kosson et al. 2024 — adjacent. It walks the same chain (too large an early step →
damage training does not undo → warm-up prevents it), and that is enough to retire an
unqualified novelty claim. It differs in two ways that matter:
Lu et al. 2020 (arXiv:1903.06733) is the closer neighbour, and was not on the list.
Born dead means dead before training; their theorem sends a gradient method (Adam
among those they name) to a constant function; and their depth-up / width-down ordering
is the order Table 1 finds. Recomputed here, it holds in all four comparisons at lr 0.03
— 69→90% and 46→81% with depth, 69→46% and 90→81% with width. But their theory is ReLU-only,
and it forbids what this page measures: a born-dead network cannot be recovered, while
these fits do train at a lower rate or behind a 200-step ramp.
What the page says now is narrower than what it said: that a head can start almost silent
in a network with no zero region to be stuck in, that whether training wakes it is what decides
the fit, and that the failure is recoverable — shown by replaying one fit bit for bit. It also
names what was searched and what was not, because this repo has already paid once for a novelty
claim that four web searches failed to check (
detector_history.md).Scope held. Related work and two references only. Nothing else on that page — the ranked
list from #674 is Tony's to rule on and the census re-run needs his GPU decision. No JSON
changed; only prose. I widened by one paper (Lu), because the novelty sentence could not be
rewritten honestly while a closer neighbour sat unread in my own review record.
Checks.
tests/test_diagnose_chorus_collapse.py3 passed; sapper clear; the page rebuildsbyte-identically on Python 3.11, 3.13 and 3.14; the darkroom copy moved with it
(
0dfdc1e5…→9083d574…, matching the repo copy) and the claim is released here.🤖 Generated with Claude Code