Skip to content

Handle mixed per-seed aggregations in figures; add AC1b attempt ledger - #20

Merged
drewOrc merged 3 commits into
mainfrom
fix/mixed-aggregation
Sep 30, 2026
Merged

drewOrc merged 3 commits into
mainfrom
fix/mixed-aggregation

Conversation

@drewOrc

@drewOrc drewOrc commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Why

AC1b attempt 2 (resumed at 3992840ddb3f) finished Haiku (8,600/8,600, US$3.178751, 0 parse failures), analysis, latency and cost, then make figures stopped: ModernBERT k=100 seeds chose argmax, summed, summed on validation and figures.selected_aggregation assumed one aggregation for all seeds. Under the frozen criteria the attempt is a FAIL (program defect). Drew decided to fix it and rerun AC1b in full at the new commit, as a new attempt with its own US$5 cap.

What

  • Mixed per-seed aggregations. Every reader of the aggregation choice was checked. analysis already builds each seed's final router and diagnostics with its own validation choice; report already lists per-seed choices; cost, comparison, the learning curves, router and threshold figures read final or diagnostics. Only the risk-coverage figure assumed agreement. curves.json keeps each aggregation's curve as a three-seed mean and std (no per-seed curves, and adding them would change the committed bytes), so a k whose seeds disagree now gets one panel per chosen aggregation, titled "mean of 3 seeds; final router of seed(s) ...", and the README caption adds a sentence explaining this. Nothing is chosen on test and neither aggregation is dropped.
  • AC1b attempt ledger. docs/ac1b/attempts.json (validated by tinyrouter.ac1b_ledger) plus docs/ac1b/README.md and each attempt's comparison.md and INCIDENT.md. Attempt 1: INFRASTRUCTURE INTERRUPTED (Drew's decision, relayed by the coordinator, 2026-09-29: an external authentication incident, not an AC1b failure), US$0; its comparison verdict FAIL is kept as comparison_verdict and the evidence is not rewritten. Attempt 2: FAIL, program defect, US$3.178751.
  • Ledger result rules. failure_kind is one of infrastructure, program defect, ac2. INFRASTRUCTURE INTERRUPTED is allowed only for infrastructure, with incident evidence (incident, time, source, error type) and the classifying decision. The same infrastructure error type on the next attempt must be INFRASTRUCTURE BLOCKED (AC1b paused). A program defect can only be FAIL; PASS needs a PASS comparison. The final AC1b outcome is still only PASS or FAIL: INFRASTRUCTURE INTERRUPTED is neither a failure nor a pass, and only a PASS completes AC1b (stated in the budget blocks and docs/ac1b/README.md).
  • Budget in three parts in comparison.md/.json and the README: original AC6 cost US$3.19 (fixed), each attempt's Haiku spend with result and reason, and the reproduction-validation total. The spend is never folded into the original cost. The README block is generated from the ledger and a test fails when it is stale.
  • Aggregation choices, listed not judged. A new comparison section lists, for every group with a final router (all 25, including every encoder curve point and the OOS ablation), the original and rerun 8-way aggregation each seed chose on validation, whether they match and which seeds differ. Status is fixed at LISTED, NOT JUDGED: no REVIEW REQUIRED, no effect on the verdict. PLAN 5.1 (pre-run clarifications) gains one line fixing this before AC1b attempt 3. On attempt 2's outputs, 7 of 25 groups differ, including ModernBERT k=100 (argmax/argmax/argmax to argmax/summed/summed), which explains its small-only REVIEW row; the REVIEW count stays 192.
  • PLAN 5.1 budget rules: one added line for the per-attempt US$5 cap and the ledger location. No other criterion changed.
  • README: AC1b status is now "has not passed yet" (Tier 1 acceptance not complete).

Verification

  • Attempt 2's real rerun outputs, copied under reproduction/3992840ddb3f/results: make figures and make report with RESULTS_ROOT and README_OUT print all six completion lines.
  • Committed results: make figures and make report twice leave results/ byte-identical to HEAD; the generated README block is unchanged.
  • make lint, make test (684 passed), make smoke pass. The README budget block changes only because attempt 1's result is now INFRASTRUCTURE INTERRUPTED and the final-outcome rule is added; the generated results block is unchanged.
  • Mutation checks: restoring the old raise, keeping only the majority aggregation, dropping the caption note, adding the original cost into the total, not counting a pending attempt, sharing one Haiku journal across ids, not validating the per-attempt cap, and hand-editing the README budget each turn a test red. For the aggregation section: judging a differing row as REVIEW, counting the section in the REVIEW total, comparing only seed 42, dropping the OOS ablation group, collapsing per-seed lists to one seed, and leaving the section out of the markdown each turn a test red. For the ledger rules: allowing INTERRUPTED for any kind, dropping the evidence or decision requirement, dropping the kind enumeration, removing or weakening the repeat-to-BLOCKED rule, dropping the final-outcome sentence, accepting PASS with a FAIL comparison, widening the kind list, and dropping "mean of 3 seeds" from the panel title each turn a test red.

…t ledger

AC1b attempt 2 stopped in make figures: ModernBERT k=100 seeds chose
argmax, summed, summed on validation and risk_coverage assumed one
aggregation for all three seeds. A k with mixed choices now gets one
risk-coverage panel per chosen aggregation, titled with the seeds whose
final router uses it, and the README caption says so. Nothing is chosen
on test. Unanimous results render byte-identically.

docs/ac1b/attempts.json records every AC1b attempt with its result,
reason and Haiku spend. Comparison and README report the budget in three
parts: the fixed original AC6 cost, each attempt, and the
reproduction-validation total. PLAN 5.1 gains the per-attempt US$5 cap.
Every group with a final router gets a row with the original and rerun
8-way aggregation each seed chose on validation, whether they match and
which seeds differ. Status is fixed at LISTED, NOT JUDGED: it adds no
REVIEW REQUIRED and never changes the verdict. It explains router
differences such as ModernBERT k=100 small-only, where seeds 43 and 44
switched to summed in AC1b attempt 2. Scope fixed in PLAN 5.1 before
attempt 3.
…ults

Per Drew's decision of 2026-09-29, attempt 1 (HTTP 503 during an
external authentication incident) is INFRASTRUCTURE INTERRUPTED, not an
AC1b failure; its comparison verdict FAIL is kept as comparison_verdict
and the evidence files are unchanged. The ledger now requires
failure_kind from a fixed list, allows INTERRUPTED only for
infrastructure with incident evidence and the classifying decision, and
requires INFRASTRUCTURE BLOCKED when the same infrastructure error recurs
on the next attempt. AC1b is completed only by a PASS; the budget blocks
say so.

Mixed-aggregation risk-coverage titles now read "mean of 3 seeds; final
router of seed ...". Docs note the attempt 1 start time source, the
relayed decision wording and when the aggregation listing was added.
@drewOrc
drewOrc merged commit 9ee4eb2 into main Sep 30, 2026
5 of 6 checks passed
@drewOrc
drewOrc deleted the fix/mixed-aggregation branch September 30, 2026 02:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant