diff --git a/docs/SESSIONS.md b/docs/SESSIONS.md index ce001488..84c09fc0 100644 --- a/docs/SESSIONS.md +++ b/docs/SESSIONS.md @@ -10,6 +10,24 @@ cannot travel (live process ids, that box's free disk, local scratch paths). --- +### 065/kosson-prior-art — DARKROOM claim RELEASED 2026-09-20: `bugarach/2026-09-19-chorus-collapse/` +- **Status:** **RELEASED 2026-09-20** with this PR. The write was made and verified: the folder's + `index.html` moved from `0dfdc1e5…` to `9083d574…`, matching the repo copy byte for byte. Was: + ACTIVE (WSMIP065), claimed before writing. The one literature question left open by the + blind round: is Kosson et al. 2024 the failure this page claims as its own? Ruled **adjacent, not + the same**; the novelty sentence is withdrawn and both Kosson and Lu et al. 2020 are now cited with + the difference stated. The page is therefore **rebuilt**, so the darkroom copy has to move with it. +- **Writes:** `index.html` in the folder below, replacing the copy the page's own §8 points at. + Nothing else in the darkroom, and no JSON changes — the data are untouched, only the prose. +- **Touches:** this block, `tools/diagnose_chorus_collapse.py` (Related work + two references only), + the rebuilt `docs/learned/chorus_collapse/index.html`, and one residual in + `docs/reviews/chorus-collapse-verify_2026-09-20.md`. **Nothing else on that page** — the ranked + list from #674 is Tony's to rule on and the census re-run needs his GPU decision. +- **Holds:** the folder below, until this lands. +- **Simulation only.** +- **Goal:** learned-model-family (goal 2). +- **Released when:** the rebuilt page and its verdict are on `main`. + ### 065/chorus-verify — DARKROOM claim RELEASED 2026-09-20: `bugarach/2026-09-19-chorus-collapse/` - **Status:** **RELEASED 2026-09-20.** The blind verify round ran, 11 of 11 roles, and escalated without repairing (blocking 3 → 4 → 6), so **the page and its JSON were never rewritten** — the diff --git a/docs/learned/chorus_collapse/index.html b/docs/learned/chorus_collapse/index.html index 0aa79699..14d6fe4e 100644 --- a/docs/learned/chorus_collapse/index.html +++ b/docs/learned/chorus_collapse/index.html @@ -243,16 +243,40 @@

5. When it happens, and what prevents it

train each as often as its configuration's fits train, 1.4 of 7. These runs were judged by training loss, not by calls on held-out recordings.

Scroll sideways to see the whole figure.
06001200180024003000training step0.00.51.01.52.02.5training loss (mean of 5 logged steps)2736f584 seed 133622c24 seed 04a12a8bb seed 15d2026d1 seed 09ad792aa seed 1a039a6a7 seed 0e86433df seed 2e87e294e seed 1trainsdoes not trainloss 0.5, the line between them
Figure 7. A 200-step linear warm-up at lr 0.03 lets most of the collapsed fits tried train. Training loss, the mean of 5 logged steps, for each collapsed fit in Table 3, replayed with the warm-up. Dashed: loss 0.5, the line this page uses between training and not. The highest single logged loss, 4.0, is in an early spike the smoothing shortens.

Table 3. The 200-step linear warm-up at lr 0.03, fit by fit.

Scroll sideways to see the whole table.
fitencodertop mstepscollapsed fits in its configuration (both draws)final lossoutcome
2736f584, seed 1 (Figure 1's configuration)4 × 641,80033 of 360.07trains
33622c24, seed 04 × 623,60033 of 360.11trains
4a12a8bb, seed 14 × 4290024 of 360.18trains
5d2026d1, seed 08 × 683,60028 of 360.09trains
9ad792aa, seed 14 × 6490031 of 360.14trains (the same run as a longer twin's, stopped earlier)
a039a6a7, seed 04 × 441,80026 of 360.11trains
e86433df, seed 28 × 623,60031 of 361.70does not train
e87e294e, seed 14 × 4890025 of 360.09trains

Final loss: the mean of the last 5 logged steps, on training batches. A fit trains when that is below 0.5; the collapsed fits sit near 1.5, about what a constant output scores on these batches.

-

Related work. Dead ReLU units, which output zero for every input, have been measured to grow -with the learning rate (Gulcehre et al. 2022, in offline reinforcement learning), and Sokar et al. -(2023) call a unit dormant when its mean absolute activation, relative to its layer's, falls below a -threshold. This page thresholds variation instead, and here every head starts mostly silent, working or not: -what fails is training's waking it, so these are related observations, not this mechanism. A +

Related work. The share of dead rectified-linear units (ReLU units, which output zero for +every input) has been measured to grow with the learning rate (Gulcehre et al. 2022, in offline +reinforcement learning), and Sokar et al. (2023) call a unit dormant when its mean absolute activation, +relative to its layer's, falls below a threshold. This page thresholds variation instead, and here every +head starts mostly silent, working or not: what fails is training's waking it, so these are related +observations, not this mechanism. A learning-rate warm-up is the standard remedy for unstable early training at a large step size (He et al. 2016; Goyal et al. 2017); why it helps Adam is disputed (Liu et al. 2020; Ma & Yarats 2021), and Ma & -Yarats's rule of thumb, 2/(1 − 0.999) = 2,000 steps, is ten times the ramp tried here. That it -prevents this collapse is this page's result.

+Yarats's rule of thumb, 2/(1 − 0.999) = 2,000 steps, is ten times the ramp tried here.

+ +

Two papers describe something close enough that the difference is worth stating. +Kosson et al. (2024) argue a warm-up works by holding down the size of the early update, and report +that large initial updates leave a small image network with a high share of permanently dead ReLU +units and a lasting loss of accuracy — the same chain this page walks, from too large an early step to +damage training does not undo. It differs in direction and in the activation. There the units are alive +and the update kills them, counted at the end of training; here 5 to 8 of the head's 8 layers already +pass nothing that varies at the starting weights, before the first step, in every replayed fit, and +what separates a collapsed fit from a working one is only whether training wakes them. Their remedy is +a leaky ReLU, which works because a ReLU has an exact zero region to be stuck in. A GELU has none, so +that repair is not available here and a layer of this head is never dead in their sense, only silent. +Lu et al. (2020) come closer: they call a network born dead when it is dead before training, +prove that a gradient method — Adam among those they name — then optimizes it to a constant function, +and show the probability rises with depth and falls with width — the order in Table 1, collapse by +shape, where at lr 0.03 the narrow deep 4 × 6 encoder collapses most and the wide shallow 8 × 4 least, +and each of the four comparisons runs that way. Their +theory is ReLU-only, and their result forbids what this page measures: a born-dead network cannot be +recovered, while these fits do train at a lower rate or behind a 200-step ramp. +So what is left here is narrower than "a warm-up prevents a collapse", which is published. It is +that a head can start almost silent in a network with no zero region to be stuck in, that whether +training wakes it is what decides the fit, and that the failure is recoverable — shown by replaying one +fit bit for bit and changing one thing at a time. The warm-up, dying-ReLU and dormant-unit literatures +were searched for this. Whether a learned event detector has been reported collapsing to one call +covering a whole recording was not.

6. The choice it leaves