Repository navigation
The slow bench's handoff says which fast results a slow run would overwrite - #706
Merged
Merged
Conversation
…rwrite Six places where building the slow stream's bench overwrites a fast result, and what to do about each: the one measured-values record `remeasure_bench.py` writes unconditionally, the single fast operating point every detector ships, the search's dated darkroom folder that two searches share on one day, the tuning tool's hardwired bench, the leaderboard's fixed filename, and the nets' fast-sized decode gap. Read off the tree rather than remembered — every row carries the file and line that makes the claim checkable. Two of the six have no seam at all: the search binds `bugarach.bench` at module scope, and the tuning tool records `bugarach.bench.make_recording` in the run's own declaration, so a slow run there trains on fast recordings and stamps them slow. That is step 2's trap, one tool further on. Also records Tony's ruling on how the slow stream gets measured: a separate `tools/measure_slow_bench.py` importing the shared tool's measurement rather than a `--bench` flag through it. The module exists because there was no time to expand what `bench.py` can do, and a stream threaded through the shared measurement tool is that same expansion in a second file; the tool is claimed on WSMIP065 besides, with an uncommitted change in its worktree. The one-stream-aware todo now deletes that tool along with the module, so the stopgap has an end. Nothing is measured or run here: no export folder, no venv and no darkroom on this machine, so this is a read of the tree and a doc change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015od7jewS4rLNXHchXQHFeb
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tony asked where WSMIP064 is on
HANDOFF-slow-bench.md, and for an assurance that building the slow stream's bench cannot clobber the fast fits and results. The first answer is "not started": nobench_slow.py, no branch, no PR, no commit on any ref, and #688 still records 064 as running the crowded-allowance sweep. The second answer is this file.Six places where a slow run overwrites a fast one, read off the tree at
a4db11dwith the file and line that makes each row checkable:docs/learned/bench_measured.json—tools/remeasure_bench.py:224writes that one path unconditionally. Caught after the fact bytests/test_bench_is_measured_on_the_declared_folder.py, which compares the record'sstreamwithbench.MEASURED_STREAM, so the loss is loud and recoverable but still a loss.bench.OPERATING_POINTS— one operating point per detector and it is a fast one; the viewer, real-data detection, the bake-off and the figure tools all read it. The seam for slow settings already exists elsewhere:detect_folder.load_settingskeys rows by(detector, stream), and detect --settings takes any parameter the detector takes, so tuned-but-unshipped settings reach real data #700 madedetect --settingstake any parameter a detector takes.tools/search_all_settings.py:809— a fast and a slow search on the same day write the same dated darkroom folder.tools/tune_learned_vs_coact.py—--outis required so no folder collides, but the bench is hardwired, including in whatmeta.jsonrecords as the run's own declaration. Without a seam a slow run trains on fast recordings and stamps them slow.tools/leaderboard.py— a fixed--runsdefault and a fixed output filename.learn/encode.py:165— the nets decode with a merge gap sized for fast.Two of those have no bench seam at all, which is step 2's trap one tool further on.
Also records the ruling on how the slow stream gets measured: a separate
tools/measure_slow_bench.pythat imports the shared tool's measurement, rather than a--benchflag threaded throughtools/remeasure_bench.py. The separate module exists because there was no time to expand whatbench.pycan do, and a stream flag through the shared measurement tool is that same expansion in a second file; that tool is claimed on WSMIP065 with an uncommitted change in its worktree besides.docs/todo/2026-09-21-one-stream-aware-bench.mdnow deletes the new tool along with the module, so the stopgap has a scheduled end.Documentation only — no code changes, and nothing was measured or run: this session has no export folder, no venv and no darkroom. The file keeps its "not murderboarded" banner for the same reason #699 landed with it: working material for a session in this tree, not for an outside reader.
🤖 Generated with Claude Code
https://claude.ai/code/session_015od7jewS4rLNXHchXQHFeb
Generated by Claude Code