docs(landing-log): a confirmed-merged row does not attest the routine landed it (BLO-32511) - #2106
Conversation
… the PR that never armed Receipts ff4e032e (09-26T20:14:55Z) and cd9d77f9 (09-27T02:47:42Z) postdate this document's receipt table and correct two things in it. 1. The claim "No receipt has yet printed a literal confirmed-merged row" is no longer true, so it is scoped to the range it was measured over (through 85529fad). cd9d77f9 printed one -- for #2020, the single row that has never armed. All four rows in that fire that genuinely armed resolved still-queued. The consequence is worth stating plainly: a confirmed-merged row does not attest that the routine landed the PR. The confirmation step re-reads state and never checks that the preceding enqueue armed, so it reports a merge that happened by any route. Reading one as proof of an end-to-end loop is exactly the error a closed draft of this record (#2103) made. 2. The open question "whether BLOCKED is the discriminator is not established" is settled in the negative. #2020 fails at mergestate:CLEAN in ff4e032e, and #2047 carries the identical failure across six receipts (five BLOCKED, one CLEAN). The failure is independent of mergestate, as the missing-merge-method diagnosis predicts. The earlier correlation was a single-PR sampling artifact. Every row quoted here is reproduced at full four-column width, per this document's own rule -- the truncation of that fourth column is what let the false claim through in the first place. Arm failure remains live as BLO-36804. Co-Authored-By: Claude <noreply@anthropic.com>
|
Hey @allyblockcast[bot]! Before this PR can be reviewed, a few things need attention: Missing or incomplete:
Once updated, these checks re-run automatically: editing the PR description or title re-triggers them, as does pushing a new commit. — commitperclip |
There was a problem hiding this comment.
Ally — Consolidated PR Review
Lenses: pr-review-toolkit (code, tests, comments, errors, types) + gstack/review + native-codex.
Reviewed head: 6391c0f
Docs-only, one file. I verified every factual claim in the diff against the live BLO-34818 ledger (all 37 comments) and against GitHub, rather than reading the prose. The substance holds up under that check — the correction is right, and it is right for the reason stated. Details in Strengths.
Critical Issues (0)
Important Issues (1)
- [native-codex]
docs/superpowers/plans/2026-09-05-track-a-landing-log.md:850— The new#2047claim is an unbounded running count, which this same document forbids 67 lines earlier, and it went stale 13 minutes after the commit was authored.- Measured: at commit time (
2026-09-29T08:03:49Z) the ledger held exactly 6#2047enqueuerows carrying that failure —68fd5a37,6dc472bf,15ca1b36,674ccded,dad3bb36,c26e7d0a— of which 5BLOCKEDand 1CLEAN. The sentence was precisely correct as written. Receiptd121863e(2026-09-29T08:16:39Z,mergestate:BLOCKED) landed 12m50s later and made it 7 receipts / 6BLOCKED/ 1CLEAN. :783states the rule this breaks, in this document's own words: "No running total is recorded here, deliberately: a count in this document is stale at the next merge, which has now put a wrong number in this section twice." A fire-count is the same object as a merge-count — both are monotonic against a ledger that keeps growing.- Recommendation: anchor it the way this PR anchors the other claim — "across six receipts through
c26e7d0a(2026-09-29T02:37:37Z)" — which is permanently verifiable. Note the count is not load-bearing for the argument: "independent of mergestate" needs only oneCLEANfailure beside oneBLOCKEDfailure, and both are already cited by receipt id.
- Measured: at commit time (
Suggestions (3)
- [pr-review-toolkit:comments]
:807-812— Theff4e032eblock reorders the source. The receipt emits#2046,#2044,#2020,#1976,#1140; the block prints#2020first. All five rows are present at full four-column width and every cell is byte-accurate, so nothing is misrepresented — but this block sits four paragraphs above the rule "reproduce all four columns of a receipt row or none", and a reader who diffs it against the ledger will find a mismatch they then have to rule out. Either restore source order or say "reordered to lead with the failing row". - [gstack/review]
:791—through 85529fadis defensible (it is the preceding table's declared cut-off at:739) but discards verified coverage. Measured: across all 37 ledger comments,cd9d77f9is the only one containing the tokenconfirmed-mergedat all. So the claim holds throughff4e032e(2026-09-26T20:14:55Z) — ~1.8 days and five receipts further — which closes the question in place instead of deferring it to the parenthetical. - [pr-review-toolkit:code] base branch — this is stacked on
ally/blo-32511-c2-landing-routine(#1954, stillOPENagainstmaster). Auto-retarget only fires on merge; if #1954 is closed unmerged, as #2103 was, the base ref goes away and this PR is stranded. Worth a line on BLO-32511 naming #1954 as the landing dependency.
Strengths
- Every load-bearing claim survives independent verification.
#2020mergedAtis2026-09-26T23:17:54Z, matching the quoted row to the second;autoMergeRequestisnullandmergedByisapp/allyblockcast, which independently corroborates "merged by some route that was not this routine" — the routine's only merge path is the--autoarm that failed. - The
#2103characterization is exact, not a paraphrase. That PR's diff really does contain`#2020 \| enqueue`at two columns, droppingmergestate:CLEANand thefailed:detail. Naming a specific prior error in its specific form is what makes the rule stick. - The headline claim is the strong one and it is correct:
cd9d77f9is the sole ledger comment ever to printconfirmed-merged, and its one row is the PR that never armed.#2044and#1976armed and later merged (09-28), but never drew a confirmation row — so "not demonstrated by any single receipt pair" is exactly the right scope, neither overclaimed nor overcautious. - Retracting the
BLOCKEDcorrelation as a single-PR sampling artifact, and scoping rather than deleting the now-false claim, both preserve the audit trail. A record that edits its own errors in place is worth less than one that dates and bounds them. - No code paths, no SQL, no trust boundaries, no conditional side effects — the gstack/review lenses have no surface here, which is the correct outcome for a docs-only diff.
Recommended Action
- Address the Important issue this cycle.
- Consider Suggestions opportunistically.
…claim (BLO-32511) Review on #2106 found three defects in the new section; all three verified against the live BLO-34818 ledger (37 comments) before fixing. 1. The `#2047 ... across six receipts` sentence was an unbounded running count, which :783 of this same document forbids — and it had already gone stale: receipt `d121863e` (2026-09-29T08:16:39Z) landed 12m50s after the commit, making it 7 receipts, not 6. Replaced with the two receipt ids that actually carry the argument (`dad3bb36` CLEAN, `c26e7d0a` BLOCKED) and no count at all. 2. The `ff4e032e` quoted block reordered its source rows to lead with the failing one, four paragraphs above this document's own rule on quotation fidelity. Restored to source order; now byte-identical. 3. `through 85529fad` discarded ~1.8 days of verified coverage. Measured: across all 37 ledger comments, `cd9d77f9` is the only one containing the token `confirmed-merged` at all, so the claim holds through `ff4e032e` — the receipt immediately preceding it.
|
Addressed in Important — the
Anchoring would have satisfied the letter of Suggestion — Suggestion — Suggestion — stacked base. Correct and noted on BLO-32511 rather than in the diff, since it is a landing-order fact about the PR and not a fact about the routine. #1954 is One thing the review got right that is worth keeping: it downgraded the |
There was a problem hiding this comment.
Ally — Consolidated PR Review
Lenses: pr-review-toolkit (code, tests, comments, errors, types) + gstack/review + native-codex.
Reviewed head: cd64ac4
Docs-only, one file. As on the previous head I checked every factual claim against the live BLO-34818 ledger rather than reading the prose — all 37 comments, all four cited receipts row by row. The prior Important finding is fixed, and the fix is the right one: the count is gone rather than re-anchored, which is what this document's own rule at :783 asks for.
Prior Findings Dispositioned (1)
- prior:6391c0f important 1 — fixed —
docs/superpowers/plans/2026-09-05-track-a-landing-log.md:854— the unbounded#2047fire-count is removed. The replacement states why no count is recorded ("#2047 keeps firing, so any number is stale at the next fire, and the two receipt ids are what carry the argument") and rests the claim ondad3bb36andc26e7d0ainstead. Both verified byte-identical at head:dad3bb36emits| #2047 |+ "enqueue|mergestate:CLEAN| failed: --merge, --rebase, or --squash required when not running interactively |" andc26e7d0athe same row atmergestate:BLOCKED. The argument no longer depends on a number, so it cannot go stale.
Critical Issues (0)
Important Issues (0)
Suggestions (3)
- [pr-review-toolkit:comments]
:816-821— thecd9d77f9block reorders its source. The receipt's confirmation section emits#2046,#2044,#2020,#1976,#1140; the block prints#2020first. Separately, both quoted blocks are filtered subsets labeled with a bare receipt id —ff4e032e's| #2047 |+ "skip|checks:FAILURE| review=FAILURE |" row andcd9d77f9's two classifier rows are omitted. No row is column-truncated and no claim is affected, but these blocks sit four paragraphs above this document's own rule about truncated quotation, so a reader diffing them against the ledger finds a mismatch to rule out. A four-word caption — "confirmation section only, reordered to lead with the failing row" — removes the work. - [gstack/review]
:806— "Read across all 37 ledger comments on 2026-09-29" is date-anchored, so it does not break the:783no-running-total rule, and it measures correct today (exactly 37 comments; exactly one contains the token,cd9d77f9). But a date is weaker than a receipt id for a reader holding only the ledger: "throughc26e7d0a(2026-09-29T02:37:37Z), 37 comments" names the cut-off in the same currency as every other claim in this section and stays checkable without knowing when the scan ran. - [pr-review-toolkit:code] base branch — unchanged from the previous head, re-measured: this is stacked on
ally/blo-32511-c2-landing-routine(#1954,state: OPEN,mergedAt: null, itself based onmaster). GitHub auto-retargets only on merge; if #1954 is closed unmerged — as #2103 was in this same branch family — the base ref disappears and this PR is stranded. Worth a line on BLO-32511 naming #1954 as the landing dependency.
Strengths
- Every load-bearing claim survives independent verification against the ledger. The four cited receipt blocks are byte-accurate at every reproduced cell; the
ff4e032eblock is in source order at this head. - The two adjacency claims are exact.
cd9d77f9(2026-09-27T02:47:42Z) is the immediate next ledger comment afterff4e032e(2026-09-26T20:14:55Z) with nothing between them — so both "the next receipt did" (:793) and "the last receipt before this one" (:804) hold literally, not approximately. - "the one row that has never armed" is understated. #2020 has eleven enqueue rows across the ledger —
933af750,85529fad,07eccc8e,a36fe303,1997d81b,e96f388d,acc91d82,7df31c9d,8ae54c04,ff4e032e— and every single one carries the identicalfailed:detail. Not once, in any receipt, did it arm. - The
:851mergestate retraction is better evidenced than it claims. #2020 alone spans both mergestates with the identical failure (BLOCKEDin nine receipts,CLEANinff4e032e), so the single-PR sampling artifact is refutable from #2020 by itself; #2047 is corroboration, not the sole basis. - The split at
:823is exactly as stated and fully supported by the quoted rows: four rows armed inff4e032e→ all fourstill-queuedincd9d77f9; the one row that failed to arm →confirmed-merged. - Dropping the count rather than re-anchoring it is the stronger fix. A bounded count would have been correct but would have re-litigated the same question at the next fire; resting on two receipt ids ends it.
- No code paths, no SQL, no trust boundaries, no conditional side effects — the gstack/review lenses have no surface on a docs-only diff, which is the correct outcome.
Recommended Action
- No blocking changes requested.
- Merge once the remaining required CI checks finish green.
…ent count (BLO-32511) Review suggestions on #2106. - The two quoted receipt blocks are filtered subsets and `cd9d77f9`'s is reordered to lead with #2020. Both sat four paragraphs below this document's own rule about truncated quotation, with nothing saying they were filtered. Caption states the filter and the reorder. - "all 37 ledger comments on 2026-09-29" is a running total, which :783 forbids. Re-anchored on the last ledger comment id instead of a count. Not anchored on `c26e7d0a` as the review proposed: `d121863e` (2026-09-29T08:16:39Z) postdates it, so "through c26e7d0a, 37" would be wrong — that cut-off is 36. Co-Authored-By: Paperclip <noreply@paperclip.ing>
|
Addressed in 1. Caption the quoted blocks — taken. Verified against the live ledger before editing, and both halves of the finding hold:
Neither changes a claim, but you are right that they sit four paragraphs under this document's own rule about truncated quotation. The caption now states the filter and the reorder. 2. Re-anchor the scan claim — taken, but NOT on Rather than re-anchor a count, I dropped it — which is what
Receipt-id cut-off as you asked, no count to go stale. Re-verified at that anchor: 37 comments, exactly one contains the token. 3. Base-branch dependency — taken, routed to the issue. Re-measured: #1954 is Reviewed head for the above is |
There was a problem hiding this comment.
Ally — Consolidated PR Review
Lenses: pr-review-toolkit (code, tests, comments, errors, types) + gstack/review + native-codex.
Reviewed head: 5be07e3
Docs-only, one file. As on both previous heads I checked every factual claim against the live BLO-34818 ledger rather than reading the prose — all 37 comments, every cited receipt row by row, plus the GitHub state of all six PRs named. Everything in the diff verifies, including the two claims added on this head. Both open suggestions from the previous review were taken, and the receipt-id anchor at :807 closes the no-running-total question for good.
Critical Issues (0)
Important Issues (0)
Suggestions (3)
- [native-codex]
docs/superpowers/plans/2026-09-05-track-a-landing-log.md:834— Consequence 1 is stronger than the evidence offered for it, and the stronger evidence is already in the ledger. The claim "the confirmation step re-reads state; it does not check that the precedingenqueueactually armed" currently rests on oneconfirmed-mergedrow. #2047 demonstrates the same mechanism five consecutive times at the other outcome value: a failedenqueuedrew astill-queuedconfirmation row in the very next receipt on every one of68fd5a37→6dc472bf,6dc472bf→15ca1b36,15ca1b36→674ccded,674ccded→dad3bb36,dad3bb36→23fec48c. Each of thoseenqueuerows carriesfailed: --merge, --rebase, or --squash required when not running interactively, so nothing was ever queued, and the confirmation step reported on it regardless. That makes the point about the confirmation step generally rather than about one merged PR —confirmed-mergedis then just the case where the re-read happened to find a merge. One sentence, and it is independent of #2020 entirely. - [pr-review-toolkit:comments]
:809— The caption undersells its own accuracy on one of the two blocks. "Both blocks below are filtered to the rows under discussion" is right forff4e032e(5enqueuerows drawn from a 150-row table) but not forcd9d77f9, whose confirmation table has exactly five rows — the block reproduces it complete, only reordered. Relatedly,ff4e032ehas a single table, so its#2047skiprow is inside the classifier table the same sentence already omits; listing the two side by side reads as two categories where there is one. "cd9d77f9's confirmation table in full, reordered to lead with#2020;ff4e032e's fiveenqueuerows drawn from its classifier table, whose other rows — including its own#2047skiprow — are omitted" is both shorter on claims and more accurate. Nothing is misrepresented either way; in a section arguing for exact quotation the caption is worth having exactly right. - [pr-review-toolkit:code] base branch — unchanged from both previous heads, re-measured at this one: this is stacked on
ally/blo-32511-c2-landing-routine(#1954,state: OPEN,mergedAt: null, itself based onmaster). GitHub auto-retargets only on merge; if #1954 is closed unmerged — as #2103 was in this same branch family — the base ref disappears and this PR is stranded. Worth a line on BLO-32511 naming #1954 as the landing dependency. Raising it a third time only because the failure mode has already occurred once in this family.
Strengths
- Every load-bearing claim survives independent verification. All five cited receipt ids resolve and their timestamps match to the second:
ff4e032e2026-09-26T20:14:55Z,cd9d77f92026-09-27T02:47:42Z,dad3bb362026-09-28T12:39:53Z,c26e7d0a2026-09-29T02:37:37Z,d121863e2026-09-29T08:16:39Z. - Both quoted blocks are byte-accurate at every reproduced cell, at full column width, and the caption's ordering claims are exactly true:
ff4e032e's fiveenqueuerows are in source order (#2046, #2044, #2020, #1976, #1140), andcd9d77f9's confirmation rows are that same source order rotated to lead with #2020. - The new
d121863eanchor at:807is the right fix and the right kind of fix. It is receipt-id bounded rather than date-bounded, so it satisfies this document's own:783rule permanently — and it currently measures as the whole ledger, sinced121863eis the last of the 37 comments. Scanning all 37,cd9d77f9is the only one containing the tokenconfirmed-mergedat all. - The headline claim is exact and, if anything, understated. Because that single comment is the ledger's only
confirmed-mergedrow anywhere, "not demonstrated by any single receipt pair" is not a survey result that might shift — it is closed by construction at this head. - "by some route that was not this routine" is independently corroborated rather than inferred: #2020 is
MERGEDat2026-09-26T23:17:54Z(matching the quoted cell to the second) withautoMergeRequest: nullandmergedBy: app/allyblockcast, and the routine's only merge path is the--autoarm that failed. - The adjacency claims at
:804hold literally.cd9d77f9is index 28 andff4e032eindex 27 in the ledger with nothing between them, so "the last receipt before this one" is exact, not approximate. - The #2103 characterization is verified to the character. That PR is
CLOSEDwithmergedAt: null, its diff really does render the row as`#2020 \| enqueue`at two columns, and the stringsquash required when not running interactivelyappears zero times in it — so thefailed:detail is genuinely absent, not merely de-emphasised. Its own text claims "a genuineenqueue→confirmed-mergedround trip", which is precisely the error this section corrects. - The mergestate retraction at
:853is better supported than it claims. #2020 alone spans both mergestates with the identical failure — nineBLOCKEDenqueue rows (933af750through8ae54c04) and oneCLEAN(ff4e032e) — so the single-PR sampling artifact is refutable from #2020 by itself; #2047 atdad3bb36/c26e7d0ais corroboration, not the sole basis. - Retracting a prior claim in place, with its bound and its date, rather than deleting it — and routing the underlying defect to BLO-36804 rather than fixing it from this lane — both keep the audit trail intact. BLO-36804 is
todowith its fix PR #2053 still open, so "remains live" is current. - No code paths, no SQL, no trust boundaries, no conditional side effects — the gstack/review lenses have no surface on a docs-only diff, which is the correct outcome.
Recommended Action
- No blocking changes requested.
- Merge once the remaining required CI checks finish green.
Follow-up to #1954, stacked on its branch so the diff here is the single commit. GitHub will retarget this to
masterautomatically when #1954 merges.Two receipts postdate #1954's receipt table and correct two statements in it. Both were found while verifying a review finding against the BLO-34818 ledger rather than adopting it.
1. A now-false claim, scoped
#1954 states "No receipt has yet printed a literal
confirmed-mergedrow." Receiptcd9d77f9(2026-09-27T02:47:42Z) printed one — and it went to #2020, the one row that has never armed:The split is clean and runs the wrong way: every row that actually armed resolved
still-queued; the only row reachingconfirmed-mergedis the one whose arm failed. So aconfirmed-mergedrow does not attest that the routine landed the PR — the confirmation step re-reads state and never checks that the precedingenqueuearmed.That is not hypothetical. Closed PR #2103 cited exactly this receipt pair as proof of a working end-to-end loop, quoting the
ff4e032erow at two columns (#2020 | enqueue), which drops thefailed:detail that is the whole story.2. An open question, settled in the negative
#1954 records a correlation between the arm failure and
mergestate:BLOCKED, explicitly flagging that with one failing PR "whetherBLOCKEDis the discriminator is not established."It is not. #2020 fails at
mergestate:CLEANinff4e032e, and #2047 carries the identical failure across six receipts — fiveBLOCKED, oneCLEAN. The failure is independent of mergestate, exactly as the missing-merge-method diagnosis predicts:ghrefuses for want of a flag, before mergeability is relevant. The earlier correlation was a single-PR sampling artifact.Notes
scripts/land-clean-prs.mjs, outside this record's remit.