Re-scoped 2026-08-31. The post-1.5 measurement window was overtaken by 1.6 and 1.7
before it opened, and 1.7 moved the surfaces F3/F4/F8 measure (+22.9% guidance bytes, two
new subcommands). The measurement is now scoped to a post-1.7 window opening ~2026-09-11.
The scope, dates, and baselines below are superseded by
this comment;
read it first.
Outcome
A re-run of the session-corpus efficiency review's method, scoped to a post-1.5-deployment window comparable to the original review's post-repair era, exists and reports whether package version 1.5's guidance restructure changed agent behaviour on the four live findings it targeted — with a keep/iterate decision recorded against D-012 and D-014's reopen triggers.
Context
The session-corpus efficiency review measured six live findings against the corpus as of 2026-08-26. Spec rev 1.9 (docs/specs/2026-08-06-github-workflow-package-spec.md, package version 1.5, commit eb9812e on testing) restructures SKILL.md and the managed block specifically to answer F3 (whole-file guidance loads), F4 (field-vocabulary.md size), F8 (help/-h rediscovery), and F10 (delegated workers bypassing the load path). D-012 and D-014 each record an explicit reopen trigger conditioned on a post-1.5 corpus window, and OQ-004 in the spec is open precisely on this question: "Did the 1.5 guidance restructure actually change agent behaviour? ... measurable only after 1.5 has been deployed for a comparable window."
The original review's post-repair era (v5.20.0, 2026-08-15) was 58 sessions and 389 invocations over eleven days — explicitly flagged in the review as "small and recent."
Scope
- Re-run the review's method (as documented in its "Method" and "Counting rules" sections) over a post-1.5-deployment window that is comparable in duration and session count to the original post-repair era (~11 days / ~58 sessions), across the same six repositories that install
github-workflow.
- Re-derive these four metrics for the new window and compare them against the review's pre- and post-repair figures:
- guidance bytes loaded per session (the F3 "bytes of guidance per post-repair session" measure, and whether loads remain whole-file);
field-vocabulary.md (or its 1.5 orientation-subset successor) load count (F4);
gh-workflow help / <subcommand> -h call count (F8);
- raw-
gh routing violations for actions the routing table now covers, including delegated-worker sessions where recoverable (F10).
- Record the keep/iterate decision this data feeds against D-012's and D-014's reopen triggers in the spec's Deviations Log or a linked follow-up entry.
Out of scope
- Re-litigating F1, F2, F5 (routing violations already at zero), F6 (CI waiting — no subcommand exists to route to), F7, or F9, which the original review already resolved or found not live.
- Any package change; this issue is measurement only. A finding that a metric has not improved is itself a valid outcome and becomes its own follow-up issue rather than being fixed inline here.
Acceptance criteria
- A dated report (new
docs/reviews/github-workflow/ entry or an amendment to the existing review) exists that:
- states the comparison window (start date, end date, session count, invocation count) and confirms it is comparable in scale to the original ~11-day / ~58-session post-repair era;
- re-derives the four metrics named in Scope, each compared numerically against the review's pre-repair and post-repair figures;
- cites the review's own "Method" and "Counting rules" sections as the method reference, noting any deviation;
- states explicitly whether each of D-012's and D-014's reopen triggers fired, and records the resulting keep-as-is or iterate decision.
- The report is linked from OQ-004's row in
docs/specs/2026-08-06-github-workflow-package-spec.md (or OQ-004 is otherwise marked answered with a pointer to it).
Evidence / references
docs/reviews/github-workflow/2026-08-26-0847-github-workflow-session-corpus-efficiency-review.md — F3, F4, F8, F10, and the corpus-scale table.
docs/specs/2026-08-06-github-workflow-package-spec.md — revision 1.9 row; D-012 and D-014 reopen triggers; OQ-004.
- Package version 1.5 deployment (commit eb9812e,
testing branch) is the measurement's starting point; the window cannot begin before that deployment reaches the corpus's six repositories.
Outcome
A re-run of the session-corpus efficiency review's method, scoped to a post-1.5-deployment window comparable to the original review's post-repair era, exists and reports whether package version 1.5's guidance restructure changed agent behaviour on the four live findings it targeted — with a keep/iterate decision recorded against D-012 and D-014's reopen triggers.
Context
The session-corpus efficiency review measured six live findings against the corpus as of 2026-08-26. Spec rev 1.9 (
docs/specs/2026-08-06-github-workflow-package-spec.md, package version 1.5, commit eb9812e ontesting) restructuresSKILL.mdand the managed block specifically to answer F3 (whole-file guidance loads), F4 (field-vocabulary.mdsize), F8 (help/-hrediscovery), and F10 (delegated workers bypassing the load path). D-012 and D-014 each record an explicit reopen trigger conditioned on a post-1.5 corpus window, and OQ-004 in the spec is open precisely on this question: "Did the 1.5 guidance restructure actually change agent behaviour? ... measurable only after 1.5 has been deployed for a comparable window."The original review's post-repair era (v5.20.0, 2026-08-15) was 58 sessions and 389 invocations over eleven days — explicitly flagged in the review as "small and recent."
Scope
github-workflow.field-vocabulary.md(or its 1.5 orientation-subset successor) load count (F4);gh-workflow help/<subcommand> -hcall count (F8);ghrouting violations for actions the routing table now covers, including delegated-worker sessions where recoverable (F10).Out of scope
Acceptance criteria
docs/reviews/github-workflow/entry or an amendment to the existing review) exists that:docs/specs/2026-08-06-github-workflow-package-spec.md(or OQ-004 is otherwise marked answered with a pointer to it).Evidence / references
docs/reviews/github-workflow/2026-08-26-0847-github-workflow-session-corpus-efficiency-review.md— F3, F4, F8, F10, and the corpus-scale table.docs/specs/2026-08-06-github-workflow-package-spec.md— revision 1.9 row; D-012 and D-014 reopen triggers; OQ-004.testingbranch) is the measurement's starting point; the window cannot begin before that deployment reaches the corpus's six repositories.