Repository navigation
Conversation
…t-Field#1072) A model reply whose whole text was a number and a full stop (32., 1024.) parsed as a Markdown ordered list with one empty item. renderer.list drew the marker column, found no body, and popped the item as undrawn, taking the marker with it, so the reply rendered to zero rows and the deck skipped it. Draw the marker as literal text when the item body produced nothing, which is the only case a bare <number>. reaches.
Owner
Author
|
Superseded by the upstream PR Agent-Field#1746 (same branch, targeting the real repo). |
7vignesh
pushed a commit
that referenced
this pull request
Oct 8, 2026
* chat: memory owners, one write door, and visible failures
The memory pipeline had four structural defects: the scope column was a
label rather than a permission, so a project's memory could surface in —
and be rewritten by — an unrelated project; the routed `remember…` command
wrote through AddMemory, the one door with no duplicate check; every write
failure (extraction, settle, import row) vanished without a record; and the
legacy memory.md import renamed the file whether or not the rows landed,
ignoring the scanner's own error on the way.
Every memory now carries an owner (`user`, `machine`, `project:<key>`)
derived at write time, and every read that can put a memory in front of a
model — the router's shortlist, the dedup search, a forget match, a person's
query — filters on the owners its session can see and refuses an empty list
rather than widening. The project key is proved from the workspace's git
origin (normalized so an SSH clone and an HTTPS clone are one project), or
the canonical path when there is no remote. Rows the old schema left
unattributable are quarantined as `project:legacy` — never re-attributed to
the person at large — and a process-open pass re-homes the ones whose source
session provably ran in a workspace, reading each session's own meta.json.
All five mouths that say "keep this" — /remember, the remember tool, the
routed command, the post-turn extraction and the legacy import — walk one
store door ([store.Store.Write]) with same-owner dedup and journaled
outcomes. Failures are journaled (`memory_write_failed`), said once per
five-minute window as a dim line, and fall back to
`v3/memory-failures.log` when the store itself is what failed. The decider
no longer gates an explicit save: an outage journals and the words still
land. The import reads scanner.Err, renames only after every row landed,
and retries idempotently through the same door. Memory off no longer closes
the conversation index: `Config.MemoryIndex` is its own handle and the
search verb survives the memory row. The FTS5 index is proved by a probe at
open instead of assumed, keyset pagination arrives with MemoryPage, and the
sync seam (ExportMemoryEvents / ApplyMemoryEvents with a receiver's owner
policy, idempotent re-delivery, tombstone precedence) is tested but wired to
no transport — PR #1738's pairing carries that later.
docs/memory-architecture.md is corrected in the same change to state what
shipped and what is deferred on purpose (vectors, edges, team owners, the
transport), with the early draft's estimates marked as estimates.
submitted a change that the project's own build or tests do not pass. senior-dev's model said: Memory architecture implementation delivered and verified: owner isolation with quarantine migration, one write door across all five mouths, visible/journaled failures, lossless idempotent import, memory/search decoupling, FTS5 probe, keyset pagination, tested sync seam, corrected architecture doc, manual page, changelog entry — with Spark full-suite/race/real-model E2E and a recorded demo; PR-side gates (push, CI, labels, attach) are outside this run's reach and documented as such. senior-dev observed: 1 of the project's 2 build and test commands failed. submitted candidate failed verification (1 failing entrypoint); shipping it anyway as the run's own answer: 5600299 bytes across 1121 file(s), tree 90058f231a51
Assisted-by: CodeAF (glm-5.3-flash, deepseek-v4-flash-0731)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* fix: memory ownership guards and dedup window
Hardened write-path owner enforcement at the store boundary with
scoped UpdateMemoryForOwners and SupersedeMemoryForOwners, making
the store the gate and not only the session's pre-check.
Fixed update guard fail-open: a store error on the visibility
check now refuses the write instead of silently proceeding.
Fixed supersede path: target visibility is now checked through
GetMemories (owner-filtered) rather than the unfiltered
MemoryRecord, and the scoped SupersedeMemoryForOwners door
enforces the owner inside the transaction.
Fixed ownerForReplay and memoryEventOwner: an unknown owner
(future team:..., unrecognised scope word) quarantines to
OwnerLegacyProject rather than widening to OwnerUser. A filter
that cannot be understood must never degrade to no filter.
Fixed rehome fold path: validates the target owner (refuses
invalid owners and rehome-to-quarantine no-ops) matching the
live door's contract.
Expanded dedup window: write neighbor lookups from 5 to 200
(FTS path) and scan limit from 1000 to 10000 (no-index path)
so aged exact duplicates are caught.
Added nine regression tests proving: foreign-ID update/supersede
refusal through scoped store doors, ownerForReplay/memoryEventOwner
quarantine for unknown owners, rehome fold refusal of invalid
target owner, aged-corpus dedup beyond old 5-row window, and
session-level post-turn extraction ownership pinning.
* session: tidy is owner-scoped and writes through guarded doors
The background tidy read its batch with ListMemories(nil, …) — an unfiltered
walk across every owner namespace and the project:legacy quarantine — and
applied its plan through the raw by-id doors (UpdateMemoryFromSession,
SupersedeMemory). Under the owner model that made it a cross-owner writer:
a quarantined row, which the architecture says is never injected and only
moves by the person's explicit re-home, could have its text silently
rewritten, and one project's tidy pass could refine another project's rows
with no visibility proof and no person in the room.
The pass now carries an explicit authorized owners list (user + machine, plus
the Config's project key when one is set; never quarantine) derived in
NewMemoryTidy, reads its batch through ListMemories(p.owners, …), refuses a
nil/empty list with an error instead of widening, and routes every refine and
supersede through the store's ForOwners doors, which re-prove the target
row's owner inside the write transaction — so a plan that names a foreign id
fails closed even when the session layer was bypassed. The supersede
replacement carries the retired row's owner as well as its type and scope.
consolidateHighWater follows the same owner scope.
Regression tests prove the quarantine is absent from the listing the model
is shown and untouched even when a plan names it; that a pass scoped to one
project cannot refine or supersede another project's row or the person's
own; that a row whose owner moved between read and write is refused; that a
missing owners list fails closed with no call and no write; and that
legitimate same-owner refinement still lands. The manual's what-i-remember
page states the tidy's owner scope in the same change.
submitted a change that the project's own build or tests do not pass. senior-dev's model said: Round2 P1 fixed: the background consolidation pass is now owner-scoped (user+machine, never quarantine/nil, fails closed on an empty owners list), every refine/supersede writes through the store's ForOwners doors that re-prove ownership in-transaction, and five deterministic negative regressions plus the manual note are added. Runtime gates are explicitly UNVERIFIED because the one permitted Spark probe timed out. senior-dev observed: 1 of the project's 2 build and test commands failed. submitted candidate failed verification (1 failing entrypoint); shipping it anyway as the run's own answer: 23972 bytes across 3 file(s), tree b75b06575b21
Assisted-by: CodeAF (qwen3.6-plus, deepseek-v4-pro, deepseek-v4-flash-0731)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* fix(memory): reject quarantine tidy and surface failed writes
Keep mutation authorization on the observed row owner and preserve the retry watermark on partial failures. Source-only change; Spark verification remains blocked.
Assisted-by: CodeAF (gpt-6.1-sol-pro)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* feat(store): journal contextual evidence and suppression
* test: add ordinary hosted contextual memory terminal recorder
* fix(store): reject suppressed sources and assistant observations
* fix(memory): make dedup and synchronized mutations durable and owner safe
* feat(store): suppress complete memory provenance and retain reasoning
* fix(memory): normalize pending tombstone ownership consistently
* feat(memory): journal bounded observed dependency impacts and dismissals
* feat(memory): ground ordinary continuity and pre-action contextual recall
* docs: record contextual memory behavior and boundaries
* fix(memory): preserve verified impacts through action and recall boundaries
* test: check carry-on evidence under explicit assistant provenance
* docs: state observed attention and receipt limitations
* test(memory): reconcile settlement fixtures with contextual ownership
* fix(memory): respect tool availability and prefix budgets
* feat(memory): durable observed-outcome learning with before-action recall
Capture failed and blocked tool outcomes at the executeTool boundary,
independent of model extraction, into the one canonical evidence journal
so a real failure outlives the turn that produced it. Surface a relevant
prior failure before the first request of a later turn, labelled with its
source circumstances and quoted as an observation.
- store: add EventContextualAttempt rows; failures need a receipt, refusals
are blocked/unknown, duplicate receipts are idempotent, and a forget
retires an attempt only through the provenance it shares with a claim.
- session: bound the source snapshot to HEAD plus a capped dirty/untracked
overlay; propagate Git errors as unknown, hash real content bytes, detect
listing and stat races, and never certify a truncated capture as current.
- document the new before-action recall on the memory page.
Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* feat(standing): deferred-delivery core — arming, honest fingerprints, tri-state judgment, durable say
The standing ticker could judge only yes/no, fingerprint a file by name/size/mtime,
arm a watch's baseline on its first wake rather than at ratification, and deliver an
ActionSay with nothing durable behind it. This lands the core contracts in
internal/standing/** only.
- File-watch hints now receive the changed-file listing as evidence; the sentinel is
no longer asked to judge a change it was handed an empty string for.
- Store.Arm captures a WhenFile baseline at ratification, under the item lock, reading
and writing the on-disk document so a concurrent edit is not overwritten. A scan it
could not read whole sets no baseline and refuses with errBaselineIncomplete. An
unarmed item keeps the silent first-reading semantics exactly.
- fingerprint hashes bounded CONTENT, not mtime: a bare touch is not a change and an
equal-size same-mtime edit is; rename and symlink target are seen; files/bytes/time
are capped and truncation is reported, never certified as "unchanged".
- Judgment is three-valued (Verdict via the optional SentinelVerdict seam; the legacy
Sentinel stays source-compatible). False and unknown are different facts: a refusal,
timeout or sentinel error is unknown, writes nothing, consumes no opportunity, and
leaves no negative in the item's history; a known cost is still charged.
- ActionSay writes a durable Pending identity to the item before the line is carried
out and clears it on acknowledgement, so a restart re-delivers under the SAME
identity instead of guessing. The optional Deliverer seam lets a Runner dedup on that
identity; the Runner.Say fallback is at-least-once and says so. ActionTask makes no
exactly-once promise: a task that fails part-way is left for the person.
- The pending list refuses rather than silently evicting authorized work at its cap.
- Inbox notes carry an identity and dedup by it across duplicate appends, restart and
drain, via a bounded drained-identity record; identity-less notes pass through.
- A guarded write (Store.saveActive) refuses to overwrite a pause, stop, edit or
pause-then-resume that landed during a probe, keyed on an Item.Revision generation
rather than status alone, and never resurrects a deleted item.
Focused tests in deferred_test.go cover hints/evidence, arm-before-first-tick,
rename/equal-size-mtime/touch, true/false/unknown/refusal/timeout, nonzero predicate
evidence, durable and replayed say identity, pending backpressure, inbox dedup across
drain/restart, failed delivery, pause/retire/pause-resume mid-probe, missing-item and
stale-save safety, and max-per-day rails.
Integration still required from the session lane: call Store.Arm at ratification, set
SentinelVerdict for a three-valued judge, and implement Deliverer (note identity) for
exactly-once say delivery. This commit does not touch session/schema/prompt/manual.
Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* Harden contextual outcome evidence and close delegated-evidence gap
Fix the confirmed high/medium review defects on the outcome candidate, still
inside the one-journal design and the fixed 4800-rune budget.
- Snapshot integrity: a stat/read metadata mismatch now returns a dedicated
race error and becomes unknown rather than certifying empty content; the
clean path double-checks HEAD and the listing; the git listing is read
through a byte-capped buffer instead of buffering it whole; an unknown
snapshot never matches another, while an empty user revision stays a
wildcard.
- Circumstances: pre- and post-action snapshots are captured and a tree that
moved is labelled unknown, not post-state-earned-current; tool evidence uses
its own receipt snapshot rather than the start-of-turn identity.
- Redaction: attempt Action and Goal are redacted before persistence and the
action is re-emitted quoted, so secrets and injected newlines cannot reach
the prompt or disk.
- Relevance: failed lookups, invalid-argument refusals and unknown-tool noise
are not stored; matching requires two distinctive terms or an action token,
with generic engineering words stopworded.
- Forget: an explicit forget now reaches a standalone attempt the extractor
never made a candidate for, by its own source key, scoped so unrelated and
fresh attempts survive and no owner is muted.
- Delegated evidence: a worker's true tool boundary feeds an observation-only
collector tied to the root session/owner/task, and a promoted bash job's
real failure is recorded when it settles. Bounded, idempotent, dropped on
close.
- Retention: reads stay one bounded window; an events_index covers the
suppression scan; physical pruning is refused by the journal's append-only
law, so retention is logical and tombstones are never erased.
Adds adversarial tests for forget-no-candidate, generic-word quiet, action
injection, bounded listing, worker/closed-collector/settled-job evidence,
unknown-vs-empty revision, and before-first-request ordering through runTurn.
Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* fix(standing): harden deferred delivery — settle-before-look, task in-flight, schema barrier, inbox recovery
Core blockers B1, B4, B7, B8, B9 from the independent review, in
internal/standing/** only. Session/CLI wiring (B2/B3/B5/B6) is the integrator's
lane; the exact contracts it must land are stated in the summary.
B1 — durable pending is reconciled independently of a fresh look. Tick settles
every active item's Pending BY IDENTITY at the top of the pass, before the rails
and before look(), so a quiet file watch, an advanced NextDue or a spent daily
rail can no longer strand an authorized line. The same identity is reused, so a
retry mints no new intent and adds no second ledger charge. A non-active item's
unresolved line is surfaced as NeedsPerson rather than forgotten; after
PendingGiveUp failed attempts an active one is too. fire no longer files a
second intent while one is unresolved (one() skips), and the pending cap refuses
rather than evicting.
B4 — a task has a durable in-flight marker. TaskInflight (run dir, started,
attempts) is written under the guarded save BEFORE Runner.Run and cleared only on
a terminal outcome; a failed or crashed attempt keeps it, NeedsPerson is set, and
one() refuses to run the task again while it is up. No exactly-once promise about
what a task did — the marker makes replay STOP, it does not make a run idempotent.
B7 — schema barrier. Schema is 3; SchemaOf writes 3 for any item carrying a
Pending or a TaskInflight marker, so a build that predates those fields (codeaf,
devaf, stageaf share one home) skips it instead of dropping the intent or
replaying the task. Ordinary and isolated orders keep versions 1 and 2, so
legacy reads and migrations are preserved. The revision guard is unchanged.
B8 — inbox crash/race/recovery/bounds. Deliver/Drain/seen mutations take a
per-inbox process flock (lock file in the system temp dir, keyed by the absolute
path, so session folders stay clean). Drain recovers leftover *.draining files
deterministically alongside a live inbox; a file it could not read whole (open,
read, or a line past inboxMaxLine) is KEPT and reported rather than deleted.
Deliver refuses a note larger than inboxMaxLine and an inbox past inboxMaxBytes
with visible backpressure. The guarantee is durable persistence and a recoverable
read handoff, not user display acknowledgement; readInbox reports scanner.Errors.
B9 — accounting and consent boundaries. A ledger/log append failure after a
successful delivery is an accountingError: Tick counts and logs it without
relabelling the item's outcome via noteFailure. Consent is checked before delivery
begins; a pause during an in-flight delivery is the person's act winning and is
never recorded as "could not check", and stale saves never undo it.
Fingerprint — boundedGlob bounds match ENUMERATION (directory entries capped,
collection capped) instead of materialising every match before the cap; a file
that vanished between listing and stat, or changed under the read, marks the
reading UNKNOWN (truncated) rather than "unchanged".
Tests: hardening_test.go adds failure-injection coverage — settle retries a
file-watch delivery under the same identity with no new look; a pending on a
paused/unexpired item is surfaced; an expired item's waiting line is honest; a
task that edits then errors is never replayed across a restart; a successful task
clears its marker; schema 3 excludes a ≤2 reader for pending items; orphan
.draining recovery; unreadable and oversized inboxes are kept and reported;
concurrent Deliver/Drain loses nothing; a delivery whose log cannot be written is
not relabelled; a pause mid-delivery is not claimed as cancelled; boundedGlob caps
enumeration; an unreadable file marks the fingerprint truncated. Focused suite
and -race green; go vet clean; session/cmd still build.
Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
* feat(standing): durable-first delivery, arm-at-yes baseline, tri-state tick, pinned model
Session and CLI half of the deferred-delivery core:
- standingRunner implements standing.Deliverer; new Deliver stamps the pending
identity into the inbox Note and the note is written and Sync'd BEFORE the line
is offered to an open window. Say/deliver return real address/inbox errors. The
live handoff is at least once and that limit is documented; the durable note is
what survives a window closing or a crash.
- Ratification arms a file watch at the yes (bounded retry); an incomplete scan
sets no baseline, leaves a visible NeedsBaselineLead flag and the first complete
look clears it.
- v3StandingTicker wires SentinelVerdict, and the parser reads whole tokens, so
'yesterday'/'nothing' are unknown rather than yes/no.
- The ratifying session's one-model policy is frozen on the item origin, fenced
at the schema barrier, and honoured by the reloaded background pass: the
sentinel and every child text seat resolve from the item's pin with role pins,
tiers, fallback chain and nearest-model guess withheld; the client cache is
keyed by model AND policy.
- A paid undecided check persists its cost on the item without consuming the
look (Store.NoteSpend), so the lifetime figure and the daily rail stay honest.
- Tests: real inbox write failure, identity dedup/restart, live-then-offline
durability, arm incomplete visible, strict tri-state scoring, pinned model on
the wire through a local fake provider (including a refusal that must not call
another model), schema fence, folder-addressed lock across TEMP and symlink,
bounded glob with a metacharacter segment, paid-unknown budget fence, real tick
constructor.
* fix(memory): delegated continuity and truthful forgetting
Forward a worker's successful trusted full reads to the root's dependency
observation, not only its failures, under an immutable origin frozen at
worker creation (root session, user turn, goal, task, owner, run id). Fail
closed without that origin rather than reattributing a late arrival to the
live turn.
Make the collector's idempotency a bounded evictable hint backed by the
canonical journal (seen only after a successful append; failed writes stay
retryable and are exposed via errorState and the memory-failure lane), bound
per-origin emissions and origin bookkeeping (LRU, pending half-read pairs
protected), and dedup redundant repeats within one frozen origin while
preserving new-turn observations.
Route a promoted job's real ending through the root-owned collector for task
workers as well as sessions, freeze the launch turn/goal, include the run
identity beside the job id in the source key, refuse a missing frozen turn,
distinguish cancellation from execution failure, and admit through the same
in-flight gate as Close. Only the collector's owner may seal it.
forget now answers partial/error truthfully per owner and exact source
instead of swallowing read/write failures; a fresh independent attempt
remains learnable.
* fix(standing): partial scans, transactional drain, pin precedence, skip diagnostics
Repair the concrete independent/root review failures at a6c2caf10:
- WhenFile: a truncated digest is never adopted as the baseline and is never
compared with a complete one; a stored partial digest reads as no baseline.
First complete look establishes the baseline and clears NeedsBaselineLead;
later complete real changes fire.
- Drain: read everything, persist the dedup record, then remove. A seen/write/
scanner/removal/read failure retains the payload and the dedup state; notes
are returned for the caller to use even when an error is joined. IDs dedup
across staged files in one drain. Deliver and Drain no longer read an
unreadable or corrupt dedup record as empty.
- standingRunConfig: the OneModel pin suppresses a contradictory item model;
unpinned items keep their explicit model.
- judge: a billed unknown whose ledger/NoteSpend write fails keeps the original
WhenProbe NextDue instead of consuming it via noteFailure.
- Remove the unused live HandedOver API; state the at-least-once fold limit.
- ListChecked reports skipped newer/corrupt/unreadable documents; Tick and the
wake log surface them (bounded) instead of passing by absence.
- standingArmBaseline arms once at ratification; the ordinary bounded Tick
establishes an incomplete baseline on its first complete scan.
* fix(attention): bound approval retrieval, precise dismissal, verified producer effects, read-only standing binding
- Approved rules and confirmed decisions get their own bounded authority
projection, so 128+ newer incidental observations cannot starve a live
rare constraint; latest correction/demotion, expiry, conditions and
suppression still win.
- Impact attention is driven by observed state, not a keyword list: a real
producer change with an unchanged re-read consumer assumption is offered
whatever the words were, and an unchanged file stays quiet.
- A delegated paired re-read of changed producer + unchanged consumer no
longer resets the recorded baseline; it moves only when the consumer's
own assumption changes, so the post-action check still sees the
consequence.
- Dismissal is precise: named target, unambiguous singleton, or explicit
batch; a bare ambiguous dismiss drops nothing and 'do not dismiss' is a
negation. Held-offer state is bounded and evictable so new material can
recur; dismiss identity persists as a canonical suppression.
- Derived receipts (history lookup, task summary, memory_evidence) can no
longer become independent observation; genuine raw tool receipts and
literal user intent are preserved. Suppressed history cannot re-enter as
fresh proof.
- Authority and global scope are judged against the claim's own supporting
span, not the whole utterance. Persisted/rendered evidence is passed
through the secret redactor while literal validation reads the original
words.
- A standing run authorized to act unattended is lent the owner-scoped
canonical brain READ-ONLY and bound before its first action.
- Home filing Save failure is surfaced truthfully instead of claiming an
unpersisted origin.
- Harness: optional --model/--one-model, explicit CODEAF_CALL_LOG_BODIES=1,
dedicated pinned profile, fleet-first isolation and a credential-free
metadata whitelist.
- Rename standing.Verdict to SentinelReading (one Verdict type law) and
route the sentinel client through Config.clientConfig (gated-config law).
- Manual, docs and change entry 1777 updated; events_no_delete law intact.
* fix(memory): gate machine scope on its own supporting span
Machine scope was kept for any supported user quote and its evidence was
pinned to the origin project, silently restricting an explicit machine-wide
rule. Machine scope now needs a self-contained machine-wide span, carries no
project condition (so it applies across projects on the same authorized
machine), and is never inferred from a tool observation. A project span beside
a global sentence and 'everywhere in this project' both stay project-local.
* fix(standing): fail closed when an action-capable run cannot be bound
withBindingMemory silently returned an unbound config when the brain could not
be opened, so a future authorized task could edit with none of its approved
rules in front of it. Memory-on binding is now a promise: an unprovable owner,
an unopenable brain, or a failed binding read returns an error before any
provider or tool call, and the ticker leaves the task's in-flight marker and a
visible needs-person line rather than replaying blind. Memory-off stays
intentionally absent, and an already-supplied brain is still bound read-only.
* fix(memory): span authority rejects questions and rejected quotations
A true substring inside a question or a rejected third-party quotation could
become an approved rule. The containing clause is now read as well as the
quote: a rule inside a question or a quotation of something the person rejected
is demoted to a proposal, while polite directives and ordinary literal
constraints still bind automatically. Textual and conservative, no new approval
pipeline.
* memory: redact-then-hash receipts, a read-only binding boundary, settled-line bounds
Fixes the two source-verified defects of the combined review and the failures
the gate and the live acquisition found after them.
F1 — an attempt and the evidence row for the SAME receipt are now the same
bytes. Every independent-attempt writer (recordMemoryAttempt, the delegated
collector's observeAttempt and settleJob) redacts the receipt clip BEFORE
hashing it, which is what the evidence path already stored and hashed. The
memory-forget provenance join matches key AND hash, so on a secret-bearing
receipt it never fired and a forgotten claim could leave the attempt readable;
now an explicit forget through the claim's own evidence retires it.
F2 — a binding-only run lends NO write through the collector either. The
delegated outcome collector is built beside the brain, and nested task workers
carry it with a frozen origin, so a worker's failing calls were journaled as
project attempt rows into a brain the run only borrowed to be bound. observe
and settleJob now refuse when the root is binding-only. Binding reads, prior
attempt history and an ordinary session's delegated continuity are unchanged.
Settled lines are bounded by the store's own caps. A live acquisition wrote
memory_write_failed twice with via=settle: the decider's merged line (638 and
683 runes) replaced a candidate that had already been clipped to 512, and the
store refused the write. applyCandidate now clips the decider's title and text,
and memoryDraft bounds the draft at the door. Applicability, rationale and
authority live on the evidence row and are untouched.
Gate fixtures: the hosted-standing test now models the required durable inbox
address with the hosted session's real transcript; the session-folder lock test
names the persistent inbox flock (now exported as standing.InboxLockName)
instead of forbidding it, and still refuses every other sidecar.
Regression tests: contextual_repair_test.go drives the production chokepoint,
the executeTool boundary, the worker spawn, the real binding function, the
promoted-job settlement and the explicit forget; plus the oversize-settle case.
* memory: carry a later observed successful alternative on a recorded failure
A fresh turn was shown the pandas-import failure but no note that the same
work had already succeeded with the stdlib path, so it repeated the dead
import to rediscover it. A demonstrated failure recorded at the tool
boundary now carries the one later true tool-boundary success of the same
turn and goal, same tool class and a shared meaningful token, as an
AttemptSucceeded row whose AlternativeOf names the failure's source key.
It is rendered beside its failure with its own receipt and source
circumstances, as observed history rather than a cause or a ban. Lookups,
metadata calls and model summaries are never alternatives; no other success
is retained; forgetting either source retires the pair through the existing
suppression join; binding-only runs write nothing at either boundary.
* memory: narrow the alternative match and hold one per failure
A1: the observed-alternative match ran as a raw substring of the failed
body, so the shared `cd /home/.../<project> &&` prefix of every ledger
bash action made an unrelated ls/cat/git metadata command eligible and
let it consume the one slot before the genuine standard-library
success. The match is now whole-token, excludes a token shared only
through an absolute path/cwd unless the turn goal named it, and refuses
a shell action whose every pipeline command is a lookup or metadata.
A2: both writers marked the one-alternative slot only after the journal
append, so a concurrent tool batch appended one alternative per eligible
sibling. The session writer and the delegated collector now reserve the
slot atomically before the append and reopen it on a failed write; the
collector reserves the emission bound under the same lock.
Also: the first-provider test asserts the actual request #1 and that no
tool ran before it; the misleading check-then-append coverage is replaced
by committed A1/A2 regressions driving the real boundary. The change
entries that claimed draft PRs 1778/1779 are consolidated into the real
draft entry 1777, with truthful limits kept and the obsolete
delegated-only dependency limitation dropped.
* memory: refuse metadata wrappers and plan status reads as alternatives
* harness: confirm terminal exit before marking stopped and stamp the binary's own version
stop sent Ctrl+C and immediately wrote stopped: Ctrl+C interrupts a running
turn and only quits codeaf when idle, so a hosted service UI stayed alive
through the 'stopped' notification. stop now polls the real tmux pane
(list-panes pane_dead/current_command, or an absent pane/session), retries
Ctrl+C on a bounded ~15s schedule, and records stop_verified_by plus a
stopped_confirmed marker only when the pane is proven gone. If it cannot
confirm closure it exits non-zero, writes diagnostics, and records no
stopped marker. No stop-all or unrelated kill; already-gone panes are
confirmed with zero keystrokes.
start recorded git HEAD beside the binary as if it were the tested build,
but a stale or dirty binary names a different revision in its own version
line. start now runs 'codeaf version', records that line, and refuses to
launch when the built revision does not match the source HEAD or the tree
or binary is dirty. Metadata stays backward compatible (unknown keys are
preserved); no key material is recorded.
* memory: follow the anchored workspace's project and bind workers read-only before their first request
An anchor moved the workspace but left Config.MemoryProjectKey on the scratch
folder, scoping project reads/writes and stamped attempts to a phantom project;
the chatv3_process seam kept the stale key on the boot config too. AnchorWorkspace
now mints the key from the resolved workspace through gitidentity.ProjectKey and
recomputes the binding note so the next action-capable request carries the
repository's approved rules. Task workers are handed the parent's brain as a
read-only Config.bindingStore plus the frozen project key, so a deterministic,
bounded, owner-filtered approved-binding read reaches the first provider request
even when the node's asynchronous semantic shortlist is late or the router is
down; no worker gains a memory write or verb and already-admitted jobs keep their
frozen owner. Comments that claimed a worker inherits a copied config are fixed.
* memory: freeze admitted scope across a root anchor and share one worker ceiling
An anchor moved the conversation's project while already-admitted work kept
running, and four owner-relative paths still read the LIVE key:
1. rebindAfterAnchor now invalidates the state earned under the old owner, in
the one place the owner changes mid-turn and under the held a.mu: the pending
demonstrated failure and its unspent alternative slot are cleared, and the
cached owner-relative impact block/cue/notices are dropped, so a scratch
consequence cannot ride into the repository.
2. contextual_delegated.go freezes the project at origin admission. A delegated
origin and a promoted job's origin each carry an explicit Project, derived
from the frozen owner, and the failure, alternative and settled-job rows take
their `project` condition from it instead of the root's live
Config.MemoryProjectKey. observeRead forwards the admitted project and owner
and refuses a forwarded read whose session is not the collector root's own,
so a colliding numeric turn cannot borrow another session's live reads.
3. A task node's optional recall is routed under the owners and project key
frozen when its reading is created (memoryBlockScoped), so a routing that
finishes after the conversation anchors cannot read the new project's memory.
4. A worker's approved-binding note and its asynchronous semantic note share ONE
memoryBlockRunes ceiling. The binding is reserved first; the semantic block
spends only the remainder, trimmed by whole records. The binding projection
selects whole records and omits an oversized rule rather than truncating it.
Read-only worker posture is unchanged: the lent store is never Config.Memory.
Co-authored-by: root integration note: delegated observeAlternative documents
the exact alternativeEligible(..., worker.config.Workspace) argument to add once
the outcome owner widens the signature.
* memory: pair a standalone failure with the goal-named artifact's success
A4: in the real session b050dd21f038f7cf the standalone
`.venv/bin/python -c 'import pandas; print(pandas.__version__)'` failed and
the same turn's goal then genuinely succeeded with a csv/Decimal calculation
over week.csv and the ledger utility, but the canonical journal held only the
failed receipt, no succeeded AlternativeOf. The association required the
success to share a meaningful token with the FAILED ACTION, and the genuine
replacement ran a different command and library and shared only the turn's
frozen GOAL and the artifact it named.
The rule now ties a same-tool-class, non-metadata, non-check-only success to
the work by a whole token shared with the failed action OR by a FILE OPERAND
shared with the turn's frozen goal. The file's own base name is the grounding
(week.csv counts however it is pathed; the shared cwd alone is not), so a
genuine calculation over the goal-named artifact pairs whether or not it prints
a friendly label, while a command that merely echoes the goal's prose
(`print("independent grand total")`) names no file and is refused. A successful
check-only probe (`python -c 'import pandas'`) is refused as a replacement; the
quote-aware segment reader keeps a `;` inside a `-c` one-liner from hiding it.
Root and delegated writers share the rule; the atomic single-alternative
reservation is untouched. Agent.invalidateOutcomePairing exposes the required
reset for an owner/project transition (the four root pairing fields live in
contextual_memory.go, owned by the lifecycle worker; the failing receipt in the
journal is immutable and never rewritten).
Regressions drive the exact production calls and the exact live prompt at the
real root and delegated boundaries, plus negative unrelated/metadata/check-only
controls, an echoed-prose control, a path-qualified goal-file control and
concurrent-batch race checks.
* memory: ground the outcome alternative in real operand use and true file identity
* memory: verify a named producer by bounded recheck, never a compound shell guess
The A5/f06 live failure left the canonical journal with zero
contextual_dependency rows: the root read its consumer (report.py) at the
real boundary, the rest of the turn was condensed to a quick task, and the
child's producer read arrived as a compound shell line whose global zero
exit proves no single cat ran and whose output substring proves current file
contents rather than a read the agent made.
Forward the root's own boundary full-read receipts into the frozen origin's
observation under the turn both sides share, so the condenser's stub never
stands in for them, and establish the link from the consumer source's exact
resolved literal producer reference plus a bounded framework re-read of that
file. The framework side is recorded with a framework-verify: label, distinct
from an actual tool receipt; refused, partial, prose, stub, unreadable,
same-owner and own-project symlink paths stay unlinked.
* memory: authorize the producer recheck by real consumption and frozen scope
The independent review of the receipt-forwarding fix reproduced five blockers,
all on the new framework/candidate path. This lands the safe, narrow fixes and
completes the two remaining cross-worker interfaces.
- A quoted path literal is a producer reference ONLY when the source actually
CONSUMES it: an argument of a file-reading/executing call or a reading
operand, never a comment, a print/log string, a bare assignment or a list of
strings. Line comments are stripped before discovery, and the code-level
pathlib chain is read FIRST (and only outside a quoted string), so eight quoted
decoys can no longer starve the genuine chain.
- The framework recheck is bounded to the consumer's AUTHORIZED SOURCE SCOPE:
the resolved producer must lie inside the consumer's own repository
neighbourhood. A path in an unrelated third project is refused before any read
and mints no owner, so a hash of an arbitrary directory is never treated as a
frozen access grant.
- contextualReadFile refuses a non-regular file at open time, opens
non-blocking (a FIFO can no longer stall the post-turn recheck) and re-checks
the same file identity after the read, so a path swapped under the read is
refused rather than returned.
- contextualTurnReceipts authenticates the source SESSION, the frozen OWNER and
the admitted PROJECT before merging the root's live receipts into a delegated
origin, so a wrong session or a post-anchor read of another repository can no
longer be borrowed at a colliding numeric turn. The source's own trusted
boundary receipts are never filtered.
- The delegated alternative-eligibility call now passes the worker's own
Config.Workspace, so the goal-file identity is grounded at the real delegated
boundary; the root already passed its config's workspace.
- One pairing reset: the anchor path calls invalidateOutcomePairingLocked, the
same rule the explicit owner transition uses, instead of clearing the fields
by hand.
Focused session/store/law slices, the new production-boundary regressions and
-race all pass; no full suite, push, merge or release. The A5 pathlib chain plus
real subprocess use is preserved.
* memory: refuse non-reader operands, masked substeps and unconsented producer reads
The A4 observed-alternative rule accepted a bare goal-file operand of ANY
non-metadata command (rm/mv/chmod/touch/an unknown tool) and an error-masked
command whose own receipt was a failure traceback, each stealing the single
alternative slot from the genuine calculation. The producer recheck also read
a sibling repository by directory neighbourhood alone, which is a discovery
bound and not an authorization grant.
- shellSegmentOperands now grounds a bare operand only for a recognised
reader (the interpreter taxonomy); file-maintenance hands named in
taskoutside.go and unknown tools fail closed.
- alternativeEligible refuses a shell action that masks a substep exit (||,
redirect to the null device); both writers refuse a success whose receipt
plainly carries a raised error.
- the framework producer read is authorized by the existing consent gate
(a.decide): allow only, ask/deny are silent refusals, no prompt or model;
a revocation lands on the next observation and the frozen owner/project
checks still stand.
- contextualCodeOnly also drops // comments and triple-quoted blocks, so a
mention in an unsupported comment/string form fails closed.
Focused session/store/approval/namelaw checks, -race on the slot and
forwarding paths, test-laws, vet and both builds pass.
* standing: bound an unchanged positive probe reading by its identity
A condition watch fired again when the sentinel answered yes on probe
evidence that had not moved since the person was already told about it,
so "notify me once when ready becomes true" put a second notice in front
of the person while the state was byte-for-byte unchanged.
The failure was the sentinel's to word and not the machinery's to trust:
the same observed bytes produced a line that read "same state already
reported" and a verdict of yes in the same breath. The fix makes the
recurrence the machinery's own fact.
Item.Positive now records the identity of the last affirmative probe
reading (a SHA-256 of the exact bytes the sentinel was shown, never the
bytes themselves). A later look whose reading hashes to that identity is
the same observed state already reported: the watch stays quiet however
the sentinel words its answer, and no negative is written. A decided no
clears the identity so a later true is a fresh rising edge; an unknown
writes nothing and leaves it alone; a reading the clip cut short carries
no identity, so a partial view is never passed off as unchanged. The
field rides the schema-3 deferred-delivery barrier so a baseline build
cannot drop it and re-report. The stand tool's probe description now
stays the person is told about the change, while a rhythm stays the
recurring kind.
Focused tests: unchanged positive reported once across at least two
later due checks with a varying sentinel line; a no on the same bytes
does not rearm; a changed reading fires; true-false-true re-arms; an
unknown writes nothing and keeps the identity; a clipped reading cannot
certify unchanged; the identity survives restart and a failed delivery
settled under the same pending id; a scheduled every still speaks every
due moment; a probe error keeps the opportunity and identity; the
schema fence.
Co-authored-by: Spark <spark@local>
* standing: make the probe identity atomic with its intent and retire one-shot conditions
The positive probe identity was written into the item by the look and only
then committed with the delivery intent. If that save failed the outer pass
could persist the identity with no pending line, no run and no delivery,
dropping the one notice the person was still waiting on. The identity a yes
earns now travels out of the look and is committed in the same guarded save
as the intent (and the task in-flight marker); a failed save rolls the
in-memory copy back, so the document the pass would write carries neither.
A document carrying the identity or a one-shot condition is fenced at schema
4, because the build that introduced the identity was itself a schema-3
reader and would silently drop it and re-report. The deferred fields keep
version 3, so an ordinary reminder, rhythm or pending line is exactly as
readable as it was.
"Notify me once when ready becomes true" is one fulfilled notification, not
an alert every time the condition is true. The model now compiles that intent
into a when.once field (shown on the card, stored on the item, never a word
match on a later pass), and a one-shot say delivers its one line and retires
after the durable queue and delivery -- including when a lost delivery is
repaired under the same identity. A recurring probe (once unset) re-arms as
before and a rhythm keeps speaking on its schedule; an ambiguous one-shot
task still stops for the person rather than replaying.
Focused tests: a real read-only-root fault at the existing write primitive
leaves the item unchanged and the retry delivers once; one-shot delivers and
retires; a lost one-shot delivery is not retired and settles once; a failed
one-shot task stays for the person; the schema-4 fence; the compiled field on
the card band with a cadence said back; the canonical words untouched. The
stand schema and description were compacted to pay for the new field and the
discharge-test clarification, and both frozen prefix caps pass with room.
* memory: match interpreters exactly and give a run's nodes the root's outcome bridge
The reader taxonomy treated any command whose basename started with
"python"/"pypy" as an inline interpreter, so an unknown python-tool,
python-config or pythonista could ground a bare goal-file operand and steal the
single CSV/Decimal alternative slot from the genuine calculation behind it.
shellInterpreterWord now matches the family word exactly against a precise
version suffix - empty, a 2/3 major, or a 2/3 major.minor - and the pip-metadata
reader uses the same primitive instead of its own prefix test. A real .venv
python and versioned names (python3.12, pypy3.10) still ground operands;
python-tool, python-config, pythonista, pypyhelper, pythonw and a module run do
not. The weaker masking shapes are closed conservatively too: a trailing
true/:, a pipeline into a bare cat, and a backgrounded & are refused, read from
the LAST substep ONLY, so the ledger utility's semicolon-separated runs and a
pipeline into a real consumer are untouched.
Adaptive orchestration nodes ran the ordinary worker loop but were built with
only the lent binding store: child.outcomes and child.origin were never set (the
only assignments are task_run.go's), so a node's failing tool calls and promoted
jobs were recorded nowhere. The task worker's wiring is extracted into the one
primitive adoptWorkerObservation and both constructors call it, so a run's node
inherits the root collector read-only and freezes the root's binding
circumstances before it runs. The bridge is lent, never rebuilt, and a node's
Close refuses to seal a collector it does not own - one settlement, never two -
while the worker's read-only memory posture is unchanged.
* standing: settle a fulfilled one-shot from its durable firing record
A one-shot whose line already reached the person and whose retire-save then
failed stayed active with its intent on disk, so the next pass re-delivered it.
The append-only ledger already records the identity of every line carried out,
so that record is now read before a pending is delivered again: a settled intent
is dropped and a fulfilled one-shot retires from the ledger evidence, without a
second delivery. A delivery that failed writes no ledger line and is still
repaired under the same Pending ID.
* session: correct the shellInterpreterVersion comment on python -m module runs
The comment claimed a python -m <tool> module run could not ground an
operand. That is false: a recognised interpreter's -m module execution
still keeps the interpreter's own operand shape and may ground a file
operand; only a pip metadata lookup is refused, by segmentPipMetadata
rather than by this taxonomy. Correct the comment only; no behaviour,
tests, docs, schema, standing, owners or caps change.
* standing: settle a completed task from its own durable firing record
A task whose Run returned and whose post-run item write was lost kept its
in-flight marker, so the next pass read a finished task as an ambiguous attempt
and wrote "a task is waiting for you before it can run again" for work that had
already finished — and a one-shot that had already spoken never retired. The
append-only ledger line a firing writes after it succeeded is now read back by
its native identity: a delivery identity for a say, and the run folder for a
task. A task attempt whose record is there is settled before the walk, so the
false waiting line is never written, a fulfilled one-shot retires once, and a
recurring task is not run again; a task with no record for its own run stays for
the person and is never replayed.
A settlement now books the firing from the record it left behind rather than
from a fresh clock: the ledger line carries the outcome kind, the check line and
the question the firing left, so a recovery restores an item's spend, outcome,
history, run folder and needs-person line, derives the clean-run streak the way
the firing does, and retires only as the firing itself would. A needs-you outcome
is a REACHED outcome, not a clean success: its question is preserved and its
streak is put back to nothing. No new scheduler or store: one append-only record,
one reader and one settlement shared by both action kinds.
* session: read the observed-alternative action at its own width
The association parsed a success through attemptAction's 240-rune DISPLAY
clip, so the exact live A4 pair (session 52a5c88b845f362b) recorded the
standalone pandas import failure but no succeeded AlternativeOf: the real
replacement's goal-named week.csv sat at rune 237 of the clip and the
file-operand link saw only 'wee'. Eligibility now reads the raw arguments
at full width; the stored 240-rune canonical Action is unchanged. A body
past the existing contextualFileBytes ceiling is refused WHOLE rather than
truncated into a false identity, and no new cap is added.
* session: read a compound action's segments at their own quotes
The hosted weather false pair (project 50d9421c…) stored an introspection-only
success as the failed `plandb done` bookkeeping call's AlternativeOf. The
wrapper `env | grep -i -E 'agent|task|plandb|codeaf' ; echo ---; plandb task
overview 2>&1 | head -40` is metadata and nothing else, but shellMetadataOnly
split on every `|`, `;` and newline in the raw text: the pipes INSIDE the
quoted grep pattern became phantom segments whose command words were the
pattern's own words ("task", "plandb"), the wrapper read as ordinary work, and
the shared pattern words then grounded the pairing.
The metadata reader now segments with the ONE quote-aware reader
[shellSegments], so a word inside a quoted argument is judged by the command
that carries the argument, and `&` as a file-descriptor duplication (`2>&1`)
stays part of its redirection instead of splitting a phantom `1` segment off.
A compound that wraps real work keeps only the work's tokens for the shared-
token link: a word contributed solely by a metadata, navigation or no-op
segment ([actionWorkTokens]) never carries the association.
No command list, cap, default or schema changed; no token denylist for this
case. The genuine CSV/Decimal + utility workflow, the go/test and remedy
cases and the one-alternative-per-failure slot are unchanged, and the failed
administrative call stays a stored observed receipt (just not a remedy).
* session: prefer an observed successful path before reconfirming a failure
* session: freeze the admission read ceiling for delegated producer reads
An independent review reproduced, at the real production boundary, that a
delegated worker's producer re-reads were judged only by the LIVE root policy.
After the root anchored a NEW repository and gained a NEW blanket allow, an OLD
worker's already-admitted consumer receipt was re-forwarded through
observeContextualDependencies and the framework then read, hashed and journaled
the OLD private sibling producer the admission policy had DENIED.
Freeze the admission-time consent policy on the frozen origin
(delegatedOrigin.ReadCeiling, captured once in rootOrigin beside Session/Owner/
Project) and judge every delegated producer re-read by the INTERSECTION of that
ceiling and the live policy (contextualProducerReadAllowedUnder). A read runs
only when both plainly allow it: an old deny can never be widened by a later
blanket allow, a later deny still revokes, and an admission both halves allow is
unchanged. Nested/orchestrate children inherit the ceiling through the one
origin-adoption seam. A nil ceiling is the configured-nothing admission and
constrains nothing. The root's own turn has no admission and still answers to
the live policy alone.
No new grant, prompt, dialog, registry or store is introduced: the ceiling is
the existing approval.Policy read by the existing pure Check. Immutable journal
events, the 4800/static caps, and existing team/chat/task ownership are
untouched.
Regression tests drive the real task-worker and nested-worker seams: old
denial/new allow forbidden, old allow/new deny forbidden, and an eligible
unchanged allow still forming the canonical edge.
* session: give delegated workers read-only prior outcomes and compose context by whole records
Task, orchestrate and audit workers are built with a LENT bindingStore and
the FROZEN MemoryProjectKey but no memory writer and no brain
(remembers() is false), so priorOutcomeContext never reached their first
request: prepareWorkerBinding rendered only the approved binding rules.
- priorOutcomeContext now delegates to a new priorOutcomeBlock(st, key, cue,
snapshot) that reads the SAME bounded, owner-filtered journal against an
explicit store and project key. A worker read borrows exactly the lent
store and frozen key it was handed, writes nothing, and is granted no
remember/forget/import verb or future-task authority. A read error on a
brain-less worker renders nothing instead of taking a nil brain.
- prepareWorkerBinding renders the approved rules FIRST and whole, then the
relevant prior outcome pairs as advisory context bounded by whole records
within what is left of the one shared memoryBlockRunes (4800) ceiling. The
asynchronous routed recall still spends only the remaining space and can
never replace the deterministic outcomes (they ride the binding note, not
the recalled block).
- Ordinary manager/chat context no longer clips a record mid-sentence. The
new composeBeforeRequestContext reserves the mandatory rules whole first
and fits the optional contextual impacts, the relevant prior outcomes and
the routed recall by WHOLE records under the single ceiling. The new
trimRenderedWholeRecords keeps the wrapper and preamble intact, drops only
trailing whole records, keeps a failure with its observed alternative on
one line, and omits a block WHOLE when the preamble does not fit or no
record would remain: an honest absence, never an instruction fragment.
withBindingContext, prepareBindingContext and the post-action impact
refresh all compose through it.
Worker snapshots stay labelled "circumstances unknown" (no root snapshot is
borrowed). No new store, schema, cap, default, owner, dataset text or static
prompt file; no change to memoryBlockRunes, ownership or the scope-ceiling
admission pointer propagation.
Focused checks on Spark (GOMAXPROCS=4, GOFLAGS=-p=2, -count=1):
go build ./internal/... ./cmd/... ; make build ; go vet ./internal/session/
go test ./internal/session/ -run
'Worker|ManagerContext|ManagerBeforeFirstRequest|TrimRendered|PriorOutcome|
ObservedAlternative|Binding|Lifecycle|Followthrough|Contextual|Delegated|
ScopeCeiling' -> ok
go test ./internal/namelaw/ -> ok
Seven new regressions in worker_outcome_context_test.go drive the REAL worker
first-request seam (newTestAgent + bindingStore) and the actual task/nested/
audit/orchestrate constructors, and pin: the whole failure+alternative pair in
the first request; read-only (no journal write, no verb); omission whole under
a full rules ceiling; no mid-record clip and closed wrappers for a manager;
and that a late routed block cannot replace the deterministic outcomes. A
negative control (overlay with the worker outcome read removed) fails the two
seam tests, so the new coverage is not vacuous.
Full serialized acceptance is left to root.
* session: worker own-source snapshot and success-first prior outcomes
* session: recognize a pathlib read_bytes calculation as meaningful work
The live acquisition (trace ee16d52197737e4a) ran .venv/bin/python
ledger.py vendor.csv (UnicodeDecodeError) and then independently computed
the category and grand totals with
.venv/bin/python -c "...pathlib.Path('vendor.csv').read_bytes()...".
The eligibility reader knew only a bare open(...) inside an inline
program, so the genuine calculation yielded no operand, was misread as a
check-only probe and could never be stored as the observed alternative.
Recognize the pathlib constructor chained to a real read
(read_bytes/read_text/open) as a file consumer, leaving a path that is
only built, printed or stat()ed as a probe. A token the failure carries
only as a file-name fragment now grounds the pairing only when the
success uses one of the failure's own files, so a different file with the
same base name cannot steal the one slot. Bound, interpreter taxonomy,
quote-aware segmentation, metadata/print/stat guards and the single
alternative slot are unchanged.
* session: read a literal assigned inside a stdin heredoc program
The live acquisition (trace f7635df7ba2eb60a, session 7f5a21a70563ff5c)
ran .venv/bin/python ledger.py vendor.csv (UnicodeDecodeError) and then
independently computed the category and grand totals with a Python STDIN
HEREDOC that assigned the literal `path = "vendor.csv"` and read it
(`open(path, "rb")`). The eligibility reader knew only literal call
arguments, so the heredoc body produced no operand, the genuine
calculation was misread as a check-only probe and no observed alternative
could be stored beside the failure.
Read a single, clearly delimited interpreter heredoc body and resolve a
bare-name call argument (`open(path, ...)`, `Path(path).read_bytes/
read_text`) only when the program assigned that name a simple string
constant EXACTLY ONCE at the top level. A dynamic alias, a reassignment or
a block-scoped literal fails closed and names no operand, as does a
cat-fed or printed body, a commented or fake call, and a command with
multiple or unterminated stdin delimiters. A bare `-` now keeps the
heredoc body as the program instead of clearing it. Bound, segmentation,
metadata/print/stat guards, interpreter taxonomy and the single
alternative slot are unchanged.
* session: honest authority and escaped records in the binding projection
The binding block is the one projection a conversation, a manager, a member
and a borrowed-store task worker open with, and the live receipts show a
worker being told its whole block was "approved rules and confirmed
decisions" over records the journal itself labels observation.
- bindingMemories labels a lexical-fallback record "History only" unless its
own latest journal row is an approved rule or a confirmed decision; the
approved window is still reserved first and whole.
- The shared preamble and the worker header now say approved rules come
first and the rest is provenance, not authority, and that Source words are
historical provenance binding only where a durable rule or an explicit
decision was stated.
- The record renderer emits each untrusted field as one JSON-quoted physical
line with < and > escaped, so a newline or a forged </memory> can no longer
split a record or forge a wrapper; the journal keeps the raw bytes.
- A clear one-turn task authorization mislabelled approved_rule or
confirmed_decision is demoted to an observation unless the span states a
durable decision; durable rules and explicit we-decided decisions still bind.
* session: close heredoc reader false-opens with coherent state
* session: fail closed one-line suites and unterminated heredoc bodies
Close two verified false-open lexical classes at the reader seam:
R-L1: an unindented one-line def/if/lambda body is no longer treated as
module top level, so a bare-name read inside a suite that may never run
names no operand.
R-L2: an unterminated heredoc keeps its whole remainder as the interpreter
stdin body, so a body line cannot be split into a phantom shell command
that falsely names a file.
Also refuse the explicit-empty -c/-e program's bare argv file arguments
(R-L4) while keeping ordinary command literal operands and the real -c
reader. The frozen vendor heredoc, its Path(...) readers, the delegated
pair and the race controls are unchanged.
* session: keep explicit lasting constraints authoritative past task verbs
F3 (verified over-demotion): the one-turn demotion read ordinary task verbs
('run the', 'check the', 'verify the', 'compare the') as a one-turn frame, so
genuine lasting constraints phrased with those verbs were demoted and stopped
binding. Split the cue list: an explicitly bounded one-off frame
(cross-check, independently, read-only, keep this, for now, for this task)
still demotes on its own, while an ordinary task verb only demotes a span that
does not also state a lasting quantifier ('always', 'never', 'must') or a
durable decision. One-turn task demotion and observation provenance are
preserved; 'for this task always print all columns' and 'run the check once; it
must not edit files' still demote.
memory_test.go: TestTheRoutedMemoriesAreRenderedIntoTheBlock asserted the
pre-quoting raw '- prefers tabs: ...'; update only that formatting expectation
to the intended safely quoted one-record format, keeping the routed-memory and
body-presence checks.
* session: scope one-turn demotion to time frames, not workstyle words
F3 follow-up (verified over-demotion): the unconditional one-turn demotion
read neutral workstyle words as a bounded time scope. 'cross-check',
'independently' and 'read-only' describe HOW work is done, not how long it
lasts, so 'Always independently verify the ledger before every release',
'The release artifacts must stay read-only' and 'We decided to cross-check the
amounts before every deployment' were demoted and stopped binding.
Split the two tiers cohere: only a genuinely explicit one-turn temporal frame
('for this task', 'for this turn', 'for this check', 'one-off', 'for now',
'just once') bounds a span on its own; the workstyle verbs now sit in the
ordinary-task-cue tier with 'run the', 'verify the' and 'check the' behind the
same durable-and-lasting guard. 'keep this' is no longer a one-off cue, since
it can introduce a lasting invariant. Lasting markers are read as fields
instead of ' must ' substrings, so punctuation ('always.', 'must,') still
counts while 'mustard'/'whenever' do not. Observation non-promotion, source
grounding and the earlier one-turn and lasting examples are preserved.
memory_authority_format_test.go: add focused cases for the three verified
spans, a punctuation-adjacent lasting marker, a 'keep this' lasting invariant
and explicitly bounded read-only tasks, keeping the existing one-turn cases.
* session: reserve approved rules before advisory history in the one budget
* session: refuse pure byte-count read previews as alternatives
* session: project historic pure byte-preview alternatives out of prior outcomes
* session: count full read dumps and straight decodes as inspection metadata
* session: disambiguate shared prior-outcome method guidance
The shared <prior_outcomes> preamble asked the model to START with an
observed successful path and not re-run a known-failed method, but said
nothing about what a person's words actually fix. A live ledger manager
read 'under the application interpreter' as a demand to run the exact
command already observed to fail, an…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
A model reply whose whole text is a number and a full stop (
32.,1024.) parsed as a Markdown ordered list with one empty item.renderer.listininternal/tui2/prosepushed the marker column, rendered the empty item, and popped it as undrawn, taking the marker with it. The reply rendered to zero rows and the deck skipped it: the fold showed only the thought row though the transcript held the text (Agent-Field#1072).The list renderer now draws the marker as the literal text the reader sent when the item body produced nothing.
32.stands under the chip;1. oneand multi-item lists still draw as lists.Diff is 9 lines in
render.goplus a guard test and a change entry.How it was checked
go test ./internal/tui2/prose -run TestABareNumberedReplyStillDraws -v— new test, fails on the base commit ("32." rendered to nothing), passes with the fix.go test ./internal/tui2/prose— full package green.go vet ./internal/tui2/prose— clean.go build ./...— whole tree builds.32.→[32.]at column 0;1. one→[1. one];1. one\n2. two→ two list rows;32→[32]. Markers for real lists are unchanged.go run ./cmd/codeaf-changes check—12 entries, all well formed.Note:
internal/namelaw's## The namesection test andinternal/tui3'sTestAConversationWithNoTitleIsNamedInWordsNotHexfail on a clean checkout of the base too, so they are pre-existing and unrelated to this change.Checklist
docs/changes/unreleased/1072-bare-number-reply-draws.md..github/known-red.txt.render.go, the new test, and the change entry.Closes Agent-Field#1072.