Add offline RQ2 to RQ4 analysis: OOS, uncertainty, LLM fallback, oracle - #17
Merged
Merged
Conversation
make analysis reads the 75 logits archives and the 8,600 stored Haiku
predictions, trains nothing and calls no API, and writes
results/analysis/{summary,curves,haiku}.json.
Temperature, 8-way aggregation, deferral threshold, hybrid signal and
operating-curve thresholds are chosen on validation only; every choosing
function raises LeakageError on any other split. Inputs are checked
before anything runs (curve indexes, archive count, Haiku rows, gold
labels equal across archives and Haiku), and outputs are read back and
checked before the completion line is printed.
Two mutants survived the first mutation pass: a two-sided z for the threshold bound, and a temperature fit that skips the split check while a later guard still catches test data. Each now has a direct test.
…ostics Threshold choosers take a Scored built from one routed split, so a split name and another split's arrays cannot be mixed. RQ2 OOS detection uses 1 - max in-scope probability at T; the ablation compares both models with that score and one aggregation. A diagnostics block explains why validation thresholds miss the target on test (OOS share, harder test OOS) and reports the reweighted-validation sensitivity as diagnosis only. Tests cover OOS score direction, noise test logits, signal choice, entropy direction, temperature-scaled summed argmax, ties at tau, single-seed std, and a shared archive across indexes.
The diagnostics hand case had no OOS row routed to an agent and deferred, so dropping the kept condition went unnoticed.
… and summed OOS score The weighted threshold (used only by the prior-shift sensitivity) now bounds the weighted error rate with n_eff = (sum w)^2 / sum w^2 of the kept rows instead of the nominal weighted count. Taus of seeds that chose different signals keep per-seed values and signal names instead of a mean. New tests: equal coverage goes to the lower AURC even when that signal is listed later, and the summed OOS score is high when the oos agent carries the mass.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Step 4, second half (docs/PLAN.md sections 3 and 4, AC3, AC4, AC6): RQ2 to RQ4 from the stored logits and Haiku predictions. Nothing is trained and no API is called. No figures here; README and figures are step 5.
What it does
make analysisreads the 75 logits archives (BERT 18, of which k=100 are the three AC2 runs; ModernBERT 18; OOS ablation 3; baselines 36) andresults/llm/haiku-8way.jsonl, and writesresults/analysis/summary.json,curves.jsonandhaiku.json. Two runs give byte-identical files: there is no randomness anywhere.Before anything is computed: each curve index is re-verified with
completeness.verify_index; the archive count (75) and the Haiku row count (8,600) are literals; within each split every archive holds the same gold labels, and Haiku'sgold_intentequals them row by row. After: each output is written to a temporary file, read back and checked (25 groups, 3 seeds each, both Haiku splits) before it is moved into place, and only then iscompleted analysis (75/75 archives, 8600/8600 llm rows, 25 groups)printed.Chosen on validation only (AC4)
Every choosing function raises
LeakageErroron any other split, and a test shuffles the test labels and checks that no choice moves.fit_temperatureon validation. Not fitted for majority (constant scores).feasible: false.Definitions
mspandentropy(negative entropy) at T = 1,msp_tat the fitted T,marginthe gap of the top two log-probabilities at the fitted T (for argmax that is (z1 - z2)/T, the same ranking as the raw score margin). TF-IDF getsmsp_tandmarginonly and no uncalibrated ECE; majority gets none (PLAN section 4.1).AGENTSorder; replies with two or more labels are counted as order-dependent), parse failures, and accuracy with parse failures counted wrong.Main test numbers (8-way, %)
Haiku: 0 parse failures, 0 rows where the two parsers disagree.
Thresholds chosen on validation miss the target on test. At a 2% target, ModernBERT k=100 has a test selective risk of 7.34 ± 0.86 against 1.50 ± 0.13 on validation. The main result stays validation-only and is reported as is (option a); see the review section for the diagnosis.
Review fixes (round 1)
The review found the computation correct and asked for these, all done on this branch:
select_thresholdandcoverage_thresholdstake aScored, built byRouted.scored(signal), so the name they check and the scores they use come from one routed split. A new test replaces the test logits with noise and checks that nothing chosen on validation moves: T, aggregation, every tau, the hybrid's signal, the operating-curve taus and the (b) taus.1 - max in-scope probabilityat the fitted T (151 intents for argmax, 8 agents for summed). The four confidence signals stay for RQ3.summary.jsonhas a generatedablation_comparison: detection with one fixed aggregation (argmax) for both models, router behavior from the final router.diagnostics.pycovers the final hybrid of every encoder point. It reports validation and test selective risk, in-scope risk, the error rate of kept OOS rows, the test risk reweighted to the validation OOS share, and how much of the gap that share explains. It also reports option (b) as a sensitivity: validation reweighted to an 18.2% OOS share, a share known only from test, so diagnosis only.select_signal. F6 entropy direction. F7 a case where T changes the summed argmax. F8 a test row exactly at tau is kept (deferred_below). F9 std is null (not 0) when fewer than two seeds have a value, withnkept. F10 two indexes pointing at one archive fail. F11 noassert isinstancein production code.Ablation, corrected (test, %, generated in
ablation_comparison)The earlier "higher AUROC without OOS training" came from ranking OOS by low confidence, which counts a confident
oosprediction as the least OOS-like row.Why thresholds miss on test (target 2%, test, %, from
diagnostics)Validation is 3.2% OOS and test is 18.2%. Most of the gap comes from that share; the rest is test OOS being harder to catch. Reweighting to the deployment OOS share still does not fully close it, so thresholds should be recalibrated on data close to real traffic before deployment.
After the refactor,
curves.jsonis byte-identical to the previous commit.summary.jsononly gains fields, plus 48 std values that change from 0 to null (F9). Two runs ofmake analysisgive identical files.Review fixes (round 2)
msp_t) is listed after the worse one (msp) inSIGNALS. The old test's better signal also came first, so removing the AURC tie-break passed.oos_scoreis close to 1 and equals 1 minus the largest in-scope agent probability.select_threshold(used only by the (b) sensitivity) now uses Kish's effective sample size. At each cut, p is the weighted error rate and n_eff = (sum w)^2 / sum w^2 over the kept rows; the bound uses k = p * n_eff. With unit weights this gives the unweighted result (tested), and with unequal weights the bound is wider than the nominal weighted count (tested). ModernBERT k=100 (b) at 2%: test risk 2.60 ± 0.48 (2.17, 3.12, 2.50), coverage 87.33 ± 1.85. It was 2.93 ± 0.63 and 88.55 ± 1.76.taukeeps the per-seed values and signal names instead of a mean and std. This covers the final hybrid, each aggregation's hybrid, the diagnostics and (b). Taus underby_signalshare one signal and are still averaged.Field-by-field comparison of
summary.jsonagainst the previous commit: the only changes are the (b) fields (diagnostics.*.sensitivity_reweighted_validation) and howtauis written (R4). Every other number is unchanged, andcurves.jsonis byte-identical.Mutation check for this round (commit first, back up, mutate, restore from the backup). The control passes, and a known mutant fails. All of these fail: removing the AURC tie-break (R1); a summed score that keeps the oos agent (R2); the nominal count in place of Kish, two variants (R3); averaging taus across different signals, and
by_signalinheriting the hybrid's signal (R4).Tests
98 tests for the analysis modules (524 in total, all offline). They cover every metric against hand-computed values, the validation-only guards (shuffled test labels, noise test logits, split-name bypass), gold cross-checks, completion checks, and the diagnostics by hand (prior weights, kept rates, reweighting, gap explained, the (b) choice).
Mutation check (commit first, back up, mutate, restore from the backup):
Scoredshape check. One survived the first pass (the diagnostics' high-confidence misroute ignoring whether the row was kept); a hand case now covers it and it fails.make lint,make test,make smoke,make verify-logits,make verify-llmandmake analysispass locally.