diff --git a/docs/SESSIONS.md b/docs/SESSIONS.md index ce001488..84c09fc0 100644 --- a/docs/SESSIONS.md +++ b/docs/SESSIONS.md @@ -10,6 +10,24 @@ cannot travel (live process ids, that box's free disk, local scratch paths). --- +### 065/kosson-prior-art — DARKROOM claim RELEASED 2026-09-20: `bugarach/2026-09-19-chorus-collapse/` +- **Status:** **RELEASED 2026-09-20** with this PR. The write was made and verified: the folder's + `index.html` moved from `0dfdc1e5…` to `9083d574…`, matching the repo copy byte for byte. Was: + ACTIVE (WSMIP065), claimed before writing. The one literature question left open by the + blind round: is Kosson et al. 2024 the failure this page claims as its own? Ruled **adjacent, not + the same**; the novelty sentence is withdrawn and both Kosson and Lu et al. 2020 are now cited with + the difference stated. The page is therefore **rebuilt**, so the darkroom copy has to move with it. +- **Writes:** `index.html` in the folder below, replacing the copy the page's own §8 points at. + Nothing else in the darkroom, and no JSON changes — the data are untouched, only the prose. +- **Touches:** this block, `tools/diagnose_chorus_collapse.py` (Related work + two references only), + the rebuilt `docs/learned/chorus_collapse/index.html`, and one residual in + `docs/reviews/chorus-collapse-verify_2026-09-20.md`. **Nothing else on that page** — the ranked + list from #674 is Tony's to rule on and the census re-run needs his GPU decision. +- **Holds:** the folder below, until this lands. +- **Simulation only.** +- **Goal:** learned-model-family (goal 2). +- **Released when:** the rebuilt page and its verdict are on `main`. + ### 065/chorus-verify — DARKROOM claim RELEASED 2026-09-20: `bugarach/2026-09-19-chorus-collapse/` - **Status:** **RELEASED 2026-09-20.** The blind verify round ran, 11 of 11 roles, and escalated without repairing (blocking 3 → 4 → 6), so **the page and its JSON were never rewritten** — the diff --git a/docs/learned/chorus_collapse/index.html b/docs/learned/chorus_collapse/index.html index 0aa79699..14d6fe4e 100644 --- a/docs/learned/chorus_collapse/index.html +++ b/docs/learned/chorus_collapse/index.html @@ -243,16 +243,40 @@
Table 3. The 200-step linear warm-up at lr 0.03, fit by fit.
| fit | encoder | top m | steps | collapsed fits in its configuration (both draws) | final loss | outcome |
|---|---|---|---|---|---|---|
| 2736f584, seed 1 (Figure 1's configuration) | 4 × 6 | 4 | 1,800 | 33 of 36 | 0.07 | trains |
| 33622c24, seed 0 | 4 × 6 | 2 | 3,600 | 33 of 36 | 0.11 | trains |
| 4a12a8bb, seed 1 | 4 × 4 | 2 | 900 | 24 of 36 | 0.18 | trains |
| 5d2026d1, seed 0 | 8 × 6 | 8 | 3,600 | 28 of 36 | 0.09 | trains |
| 9ad792aa, seed 1 | 4 × 6 | 4 | 900 | 31 of 36 | 0.14 | trains (the same run as a longer twin's, stopped earlier) |
| a039a6a7, seed 0 | 4 × 4 | 4 | 1,800 | 26 of 36 | 0.11 | trains |
| e86433df, seed 2 | 8 × 6 | 2 | 3,600 | 31 of 36 | 1.70 | does not train |
| e87e294e, seed 1 | 4 × 4 | 8 | 900 | 25 of 36 | 0.09 | trains |
Final loss: the mean of the last 5 logged steps, on training batches. A fit trains when that is below 0.5; the collapsed fits sit near 1.5, about what a constant output scores on these batches.
-Related work. Dead ReLU units, which output zero for every input, have been measured to grow -with the learning rate (Gulcehre et al. 2022, in offline reinforcement learning), and Sokar et al. -(2023) call a unit dormant when its mean absolute activation, relative to its layer's, falls below a -threshold. This page thresholds variation instead, and here every head starts mostly silent, working or not: -what fails is training's waking it, so these are related observations, not this mechanism. A +
Related work. The share of dead rectified-linear units (ReLU units, which output zero for +every input) has been measured to grow with the learning rate (Gulcehre et al. 2022, in offline +reinforcement learning), and Sokar et al. (2023) call a unit dormant when its mean absolute activation, +relative to its layer's, falls below a threshold. This page thresholds variation instead, and here every +head starts mostly silent, working or not: what fails is training's waking it, so these are related +observations, not this mechanism. A learning-rate warm-up is the standard remedy for unstable early training at a large step size (He et al. 2016; Goyal et al. 2017); why it helps Adam is disputed (Liu et al. 2020; Ma & Yarats 2021), and Ma & -Yarats's rule of thumb, 2/(1 − 0.999) = 2,000 steps, is ten times the ramp tried here. That it -prevents this collapse is this page's result.
+Yarats's rule of thumb, 2/(1 − 0.999) = 2,000 steps, is ten times the ramp tried here. + +Two papers describe something close enough that the difference is worth stating. +Kosson et al. (2024) argue a warm-up works by holding down the size of the early update, and report +that large initial updates leave a small image network with a high share of permanently dead ReLU +units and a lasting loss of accuracy — the same chain this page walks, from too large an early step to +damage training does not undo. It differs in direction and in the activation. There the units are alive +and the update kills them, counted at the end of training; here 5 to 8 of the head's 8 layers already +pass nothing that varies at the starting weights, before the first step, in every replayed fit, and +what separates a collapsed fit from a working one is only whether training wakes them. Their remedy is +a leaky ReLU, which works because a ReLU has an exact zero region to be stuck in. A GELU has none, so +that repair is not available here and a layer of this head is never dead in their sense, only silent. +Lu et al. (2020) come closer: they call a network born dead when it is dead before training, +prove that a gradient method — Adam among those they name — then optimizes it to a constant function, +and show the probability rises with depth and falls with width — the order in Table 1, collapse by +shape, where at lr 0.03 the narrow deep 4 × 6 encoder collapses most and the wide shallow 8 × 4 least, +and each of the four comparisons runs that way. Their +theory is ReLU-only, and their result forbids what this page measures: a born-dead network cannot be +recovered, while these fits do train at a lower rate or behind a 200-step ramp. +So what is left here is narrower than "a warm-up prevents a collapse", which is published. It is +that a head can start almost silent in a network with no zero region to be stuck in, that whether +training wakes it is what decides the fit, and that the failure is recoverable — shown by replaying one +fit bit for bit and changing one thing at a time. The warm-up, dying-ReLU and dormant-unit literatures +were searched for this. Whether a learned event detector has been reported collapsing to one call +covering a whole recording was not.
Related work. Dead ReLU units, which output zero for every input, have been measured to grow -with the learning rate (Gulcehre et al. 2022, in offline reinforcement learning), and Sokar et al. -(2023) call a unit dormant when its mean absolute activation, relative to its layer's, falls below a -threshold. This page thresholds variation instead, and here every head starts mostly silent, working or not: -what fails is training's waking it, so these are related observations, not this mechanism. A +
Related work. The share of dead rectified-linear units (ReLU units, which output zero for +every input) has been measured to grow with the learning rate (Gulcehre et al. 2022, in offline +reinforcement learning), and Sokar et al. (2023) call a unit dormant when its mean absolute activation, +relative to its layer's, falls below a threshold. This page thresholds variation instead, and here every +head starts mostly silent, working or not: what fails is training's waking it, so these are related +observations, not this mechanism. A learning-rate warm-up is the standard remedy for unstable early training at a large step size (He et al. 2016; Goyal et al. 2017); why it helps Adam is disputed (Liu et al. 2020; Ma & Yarats 2021), and Ma & -Yarats's rule of thumb, 2/(1 − 0.999) = 2,000 steps, is ten times the ramp tried here. That it -prevents this collapse is this page's result.
+Yarats's rule of thumb, 2/(1 − 0.999) = 2,000 steps, is ten times the ramp tried here. + +Two papers describe something close enough that the difference is worth stating. +Kosson et al. (2024) argue a warm-up works by holding down the size of the early update, and report +that large initial updates leave a small image network with a high share of permanently dead ReLU +units and a lasting loss of accuracy — the same chain this page walks, from too large an early step to +damage training does not undo. It differs in direction and in the activation. There the units are alive +and the update kills them, counted at the end of training; here 5 to 8 of the head's 8 layers already +pass nothing that varies at the starting weights, before the first step, in every replayed fit, and +what separates a collapsed fit from a working one is only whether training wakes them. Their remedy is +a leaky ReLU, which works because a ReLU has an exact zero region to be stuck in. A GELU has none, so +that repair is not available here and a layer of this head is never dead in their sense, only silent. +Lu et al. (2020) come closer: they call a network born dead when it is dead before training, +prove that a gradient method — Adam among those they name — then optimizes it to a constant function, +and show the probability rises with depth and falls with width — the order in Table 1, collapse by +shape, where at lr 0.03 the narrow deep 4 × 6 encoder collapses most and the wide shallow 8 × 4 least, +and each of the four comparisons runs that way. Their +theory is ReLU-only, and their result forbids what this page measures: a born-dead network cannot be +recovered, while these fits do train at a lower rate or behind a 200-step ramp. +So what is left here is narrower than "a warm-up prevents a collapse", which is published. It is +that a head can start almost silent in a network with no zero region to be stuck in, that whether +training wakes it is what decides the fit, and that the failure is recoverable — shown by replaying one +fit bit for bit and changing one thing at a time. The warm-up, dying-ReLU and dormant-unit literatures +were searched for this. Whether a learned event detector has been reported collapsing to one call +covering a whole recording was not.