Skip to content

The goal page's expected-overlap figures use the right estimator, and now agree with the collapse page - #678

Merged
syncytium2 merged 1 commit into
mainfrom
goals/overlap-estimator
Sep 20, 2026
Merged

syncytium2 merged 1 commit into
mainfrom
goals/overlap-estimator

Conversation

@syncytium2

Copy link
Copy Markdown
Owner

The blind verify round (#674) flagged docs/goals/learned-model-family.md stating 112.6 / 40.4 where the chorus-collapse page states 111.5 / 39.6. The goal page's numbers were mine, from #668, and they were the worse of the two.

Both estimate the same quantity: how many of chorus_norm's inner fits would collapse in both draws if collapse were independent within a configuration.

estimator chorus_norm chorus_gain_norm
mine (#668) pool the two draws' rates per configuration, square 112.6 40.4
the collapse page multiply each draw's own rate 111.5 39.6
observed 111 38

Squaring an average is never below multiplying the two values, so pooling biases the estimate upward; the product is what's unbiased when the two draws may sit at different per-configuration rates. Recomputed from the same collapse_table.json.

The reading is unchanged — the overlap is what chance gives, and the failure is not shown to be deterministic. What changes is that the figure a reader checks is now the defensible one, and the two pages no longer contradict each other. That contradiction is exactly what #674's cross-check exists to find, and it found mine.

Docs only, one line. Sapper and check_quotes clear.

🤖 Generated with Claude Code

https://claude.ai/code/session_017hQX6iESxBsQk7Jan875e3


Generated by Claude Code

… now agree with the collapse page

The blind verify round (#674) found docs/goals/ and the chorus-collapse page
disagreeing: 112.6/40.4 against 111.5/39.6. The goal page's numbers were mine
(#668) and they were the worse of the two.

Both estimate the same thing — how many of chorus_norm's inner fits would collapse
in BOTH draws if collapse were independent within a configuration. I pooled each
configuration's rate across the two draws and squared it. The page multiplies each
draw's own rate. Squaring an average is never below multiplying the two values, so
pooling biases the estimate upward, and it is the product that is unbiased when the
two draws may sit at different rates.

Recomputed from the same collapse_table.json: 111.5 against 111 observed, and 39.6
against 38. The reading does not change — the overlap is what chance gives, and the
failure is not shown to be deterministic — but the figure a reader checks is now
the defensible one and the two pages no longer contradict each other.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017hQX6iESxBsQk7Jan875e3
@syncytium2
syncytium2 merged commit 63510f4 into main Sep 20, 2026
3 checks passed
@syncytium2
syncytium2 deleted the goals/overlap-estimator branch September 20, 2026 22:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants