Skip to content

Contextual memory across chats, teams, and tasks - #1777

Merged
santoshkumarradha merged 101 commits into
devfrom
contextual-memory-20261005
Oct 6, 2026
Merged

santoshkumarradha merged 101 commits into
devfrom
contextual-memory-20261005

Conversation

@santoshkumarradha

@santoshkumarradha santoshkumarradha commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Carries constraints, decisions and observed outcomes into later chats, team members and tasks. An owner-scoped evidence journal preserves provenance. Approved rules and confirmed decisions load locally before the first request, separately from optional advisory recall. Recall is bounded to 4,800 runes; task workers read shared memory without writing it.

The built-in manual now explains workspace scope, manager/member/task recall, advisory failure reuse, forgetting, changed inputs, credential recovery and where closed-app notifications wait. Retrieval regressions check relevant sections and the distinctions they must explain.

Examples tested

  • Three fresh-chat trials reused or adapted a previously successful method after a failure. Changed input produced 40.00 instead of the old 37.35.
  • A fresh manager and newly started member received scoped lessons; a native task completed WORK and CHECK with exact-decimal 42.35. An unrelated successful calculation was not labelled as repairing a failed parser.
  • Forgetting a decision prevented a later history search from silently saving it again.
  • With the entire app closed, the real Linux user timer stayed quiet while readiness was false, queued one notification when true, and delivered it when the originating chat reopened, without replay on another reopen.
  • Forced UI termination did not prevent the next check. Missing credentials left a watch retryable; restoring them and restarting across a missed timer boundary produced one catch-up.

Validation

Runtime examples were tested on d708504cb using deepseek/deepseek-v4.1-flash (Spark gate 20261006-171029-001708; closed-app operator 20261006-180803-001716, witness 20261006-181518-001717).

Final head 3c8aaaee63ac2e36d521bc641c75b228d9acd376 adds only manual documentation and retrieval tests to that runtime. Full Spark pre-merge checks passed with no initial failures: 20261006-200539-001739. Nine fresh hosted help scenarios were checked on that exact head/model: 20261006-200550-001740.

Earlier answers exposed incorrect explanations about provider credentials, binding versus routed memory, failure recall, default scope, changed-input history and inbox destination. Those manual contradictions were corrected and the questions rerun; failed evidence is retained. One minor final wording qualification: explicit user scope also saves personal memory; /remember is its personal-by-default shortcut.

Limits

Notifications stay inside codeaf; no desktop toast or email. Ordinary chat items wait in their originating chat; Home “ask here” items use the project inbox. Producer-to-consumer notice delivery remains unimplemented. Arrival-order synchronization is not convergent, and journal/derivation costs grow. Host suspend/reboot was not tested; timer downtime was. Examples are finite and model-dependent.

@santoshkumarradha santoshkumarradha changed the title WIP: grounded contextual memory and ordinary-work acceptance WIP: contextual memory continuity, outcome learning and deferred delivery Oct 6, 2026
@santoshkumarradha
santoshkumarradha force-pushed the contextual-memory-20261005 branch from bbc86b8 to f06b5f7 Compare October 6, 2026 03:45
@santoshkumarradha santoshkumarradha changed the title WIP: contextual memory continuity, outcome learning and deferred delivery Contextual memory across chats, teams, and tasks Oct 6, 2026
agentfield-bot and others added 27 commits October 6, 2026 09:38
The memory pipeline had four structural defects: the scope column was a
label rather than a permission, so a project's memory could surface in —
and be rewritten by — an unrelated project; the routed `remember…` command
wrote through AddMemory, the one door with no duplicate check; every write
failure (extraction, settle, import row) vanished without a record; and the
legacy memory.md import renamed the file whether or not the rows landed,
ignoring the scanner's own error on the way.

Every memory now carries an owner (`user`, `machine`, `project:<key>`)
derived at write time, and every read that can put a memory in front of a
model — the router's shortlist, the dedup search, a forget match, a person's
query — filters on the owners its session can see and refuses an empty list
rather than widening. The project key is proved from the workspace's git
origin (normalized so an SSH clone and an HTTPS clone are one project), or
the canonical path when there is no remote. Rows the old schema left
unattributable are quarantined as `project:legacy` — never re-attributed to
the person at large — and a process-open pass re-homes the ones whose source
session provably ran in a workspace, reading each session's own meta.json.

All five mouths that say "keep this" — /remember, the remember tool, the
routed command, the post-turn extraction and the legacy import — walk one
store door ([store.Store.Write]) with same-owner dedup and journaled
outcomes. Failures are journaled (`memory_write_failed`), said once per
five-minute window as a dim line, and fall back to
`v3/memory-failures.log` when the store itself is what failed. The decider
no longer gates an explicit save: an outage journals and the words still
land. The import reads scanner.Err, renames only after every row landed,
and retries idempotently through the same door. Memory off no longer closes
the conversation index: `Config.MemoryIndex` is its own handle and the
search verb survives the memory row. The FTS5 index is proved by a probe at
open instead of assumed, keyset pagination arrives with MemoryPage, and the
sync seam (ExportMemoryEvents / ApplyMemoryEvents with a receiver's owner
policy, idempotent re-delivery, tombstone precedence) is tested but wired to
no transport — PR #1738's pairing carries that later.

docs/memory-architecture.md is corrected in the same change to state what
shipped and what is deferred on purpose (vectors, edges, team owners, the
transport), with the early draft's estimates marked as estimates.

submitted a change that the project's own build or tests do not pass. senior-dev's model said: Memory architecture implementation delivered and verified: owner isolation with quarantine migration, one write door across all five mouths, visible/journaled failures, lossless idempotent import, memory/search decoupling, FTS5 probe, keyset pagination, tested sync seam, corrected architecture doc, manual page, changelog entry — with Spark full-suite/race/real-model E2E and a recorded demo; PR-side gates (push, CI, labels, attach) are outside this run's reach and documented as such. senior-dev observed: 1 of the project's 2 build and test commands failed. submitted candidate failed verification (1 failing entrypoint); shipping it anyway as the run's own answer: 5600299 bytes across 1121 file(s), tree 90058f2

Assisted-by: CodeAF (glm-5.3-flash, deepseek-v4-flash-0731)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Hardened write-path owner enforcement at the store boundary with
scoped UpdateMemoryForOwners and SupersedeMemoryForOwners, making
the store the gate and not only the session's pre-check.

Fixed update guard fail-open: a store error on the visibility
check now refuses the write instead of silently proceeding.

Fixed supersede path: target visibility is now checked through
GetMemories (owner-filtered) rather than the unfiltered
MemoryRecord, and the scoped SupersedeMemoryForOwners door
enforces the owner inside the transaction.

Fixed ownerForReplay and memoryEventOwner: an unknown owner
(future team:..., unrecognised scope word) quarantines to
OwnerLegacyProject rather than widening to OwnerUser. A filter
that cannot be understood must never degrade to no filter.

Fixed rehome fold path: validates the target owner (refuses
invalid owners and rehome-to-quarantine no-ops) matching the
live door's contract.

Expanded dedup window: write neighbor lookups from 5 to 200
(FTS path) and scan limit from 1000 to 10000 (no-index path)
so aged exact duplicates are caught.

Added nine regression tests proving: foreign-ID update/supersede
refusal through scoped store doors, ownerForReplay/memoryEventOwner
quarantine for unknown owners, rehome fold refusal of invalid
target owner, aged-corpus dedup beyond old 5-row window, and
session-level post-turn extraction ownership pinning.
The background tidy read its batch with ListMemories(nil, …) — an unfiltered
walk across every owner namespace and the project:legacy quarantine — and
applied its plan through the raw by-id doors (UpdateMemoryFromSession,
SupersedeMemory). Under the owner model that made it a cross-owner writer:
a quarantined row, which the architecture says is never injected and only
moves by the person's explicit re-home, could have its text silently
rewritten, and one project's tidy pass could refine another project's rows
with no visibility proof and no person in the room.

The pass now carries an explicit authorized owners list (user + machine, plus
the Config's project key when one is set; never quarantine) derived in
NewMemoryTidy, reads its batch through ListMemories(p.owners, …), refuses a
nil/empty list with an error instead of widening, and routes every refine and
supersede through the store's ForOwners doors, which re-prove the target
row's owner inside the write transaction — so a plan that names a foreign id
fails closed even when the session layer was bypassed. The supersede
replacement carries the retired row's owner as well as its type and scope.
consolidateHighWater follows the same owner scope.

Regression tests prove the quarantine is absent from the listing the model
is shown and untouched even when a plan names it; that a pass scoped to one
project cannot refine or supersede another project's row or the person's
own; that a row whose owner moved between read and write is refused; that a
missing owners list fails closed with no call and no write; and that
legitimate same-owner refinement still lands. The manual's what-i-remember
page states the tidy's owner scope in the same change.

submitted a change that the project's own build or tests do not pass. senior-dev's model said: Round2 P1 fixed: the background consolidation pass is now owner-scoped (user+machine, never quarantine/nil, fails closed on an empty owners list), every refine/supersede writes through the store's ForOwners doors that re-prove ownership in-transaction, and five deterministic negative regressions plus the manual note are added. Runtime gates are explicitly UNVERIFIED because the one permitted Spark probe timed out. senior-dev observed: 1 of the project's 2 build and test commands failed. submitted candidate failed verification (1 failing entrypoint); shipping it anyway as the run's own answer: 23972 bytes across 3 file(s), tree b75b06575b21

Assisted-by: CodeAF (qwen3.6-plus, deepseek-v4-pro, deepseek-v4-flash-0731)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Keep mutation authorization on the observed row owner and preserve the retry watermark on partial failures. Source-only change; Spark verification remains blocked.

Assisted-by: CodeAF (gpt-6.1-sol-pro)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
…call

Capture failed and blocked tool outcomes at the executeTool boundary,
independent of model extraction, into the one canonical evidence journal
so a real failure outlives the turn that produced it. Surface a relevant
prior failure before the first request of a later turn, labelled with its
source circumstances and quoted as an observation.

- store: add EventContextualAttempt rows; failures need a receipt, refusals
  are blocked/unknown, duplicate receipts are idempotent, and a forget
  retires an attempt only through the provenance it shares with a claim.
- session: bound the source snapshot to HEAD plus a capped dirty/untracked
  overlay; propagate Git errors as unknown, hash real content bytes, detect
  listing and stat races, and never certify a truncated capture as current.
- document the new before-action recall on the memory page.

Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
… tri-state judgment, durable say

The standing ticker could judge only yes/no, fingerprint a file by name/size/mtime,
arm a watch's baseline on its first wake rather than at ratification, and deliver an
ActionSay with nothing durable behind it. This lands the core contracts in
internal/standing/** only.

- File-watch hints now receive the changed-file listing as evidence; the sentinel is
  no longer asked to judge a change it was handed an empty string for.
- Store.Arm captures a WhenFile baseline at ratification, under the item lock, reading
  and writing the on-disk document so a concurrent edit is not overwritten. A scan it
  could not read whole sets no baseline and refuses with errBaselineIncomplete. An
  unarmed item keeps the silent first-reading semantics exactly.
- fingerprint hashes bounded CONTENT, not mtime: a bare touch is not a change and an
  equal-size same-mtime edit is; rename and symlink target are seen; files/bytes/time
  are capped and truncation is reported, never certified as "unchanged".
- Judgment is three-valued (Verdict via the optional SentinelVerdict seam; the legacy
  Sentinel stays source-compatible). False and unknown are different facts: a refusal,
  timeout or sentinel error is unknown, writes nothing, consumes no opportunity, and
  leaves no negative in the item's history; a known cost is still charged.
- ActionSay writes a durable Pending identity to the item before the line is carried
  out and clears it on acknowledgement, so a restart re-delivers under the SAME
  identity instead of guessing. The optional Deliverer seam lets a Runner dedup on that
  identity; the Runner.Say fallback is at-least-once and says so. ActionTask makes no
  exactly-once promise: a task that fails part-way is left for the person.
- The pending list refuses rather than silently evicting authorized work at its cap.
- Inbox notes carry an identity and dedup by it across duplicate appends, restart and
  drain, via a bounded drained-identity record; identity-less notes pass through.
- A guarded write (Store.saveActive) refuses to overwrite a pause, stop, edit or
  pause-then-resume that landed during a probe, keyed on an Item.Revision generation
  rather than status alone, and never resurrects a deleted item.

Focused tests in deferred_test.go cover hints/evidence, arm-before-first-tick,
rename/equal-size-mtime/touch, true/false/unknown/refusal/timeout, nonzero predicate
evidence, durable and replayed say identity, pending backpressure, inbox dedup across
drain/restart, failed delivery, pause/retire/pause-resume mid-probe, missing-item and
stale-save safety, and max-per-day rails.

Integration still required from the session lane: call Store.Arm at ratification, set
SentinelVerdict for a three-valued judge, and implement Deliverer (note identity) for
exactly-once say delivery. This commit does not touch session/schema/prompt/manual.

Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
Fix the confirmed high/medium review defects on the outcome candidate, still
inside the one-journal design and the fixed 4800-rune budget.

- Snapshot integrity: a stat/read metadata mismatch now returns a dedicated
  race error and becomes unknown rather than certifying empty content; the
  clean path double-checks HEAD and the listing; the git listing is read
  through a byte-capped buffer instead of buffering it whole; an unknown
  snapshot never matches another, while an empty user revision stays a
  wildcard.
- Circumstances: pre- and post-action snapshots are captured and a tree that
  moved is labelled unknown, not post-state-earned-current; tool evidence uses
  its own receipt snapshot rather than the start-of-turn identity.
- Redaction: attempt Action and Goal are redacted before persistence and the
  action is re-emitted quoted, so secrets and injected newlines cannot reach
  the prompt or disk.
- Relevance: failed lookups, invalid-argument refusals and unknown-tool noise
  are not stored; matching requires two distinctive terms or an action token,
  with generic engineering words stopworded.
- Forget: an explicit forget now reaches a standalone attempt the extractor
  never made a candidate for, by its own source key, scoped so unrelated and
  fresh attempts survive and no owner is muted.
- Delegated evidence: a worker's true tool boundary feeds an observation-only
  collector tied to the root session/owner/task, and a promoted bash job's
  real failure is recorded when it settles. Bounded, idempotent, dropped on
  close.
- Retention: reads stay one bounded window; an events_index covers the
  suppression scan; physical pruning is refused by the journal's append-only
  law, so retention is logical and tombstones are never erased.

Adds adversarial tests for forget-no-candidate, generic-word quiet, action
injection, bounded listing, worker/closed-collector/settled-job evidence,
unknown-vs-empty revision, and before-first-request ordering through runTurn.

Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
…-flight, schema barrier, inbox recovery

Core blockers B1, B4, B7, B8, B9 from the independent review, in
internal/standing/** only. Session/CLI wiring (B2/B3/B5/B6) is the integrator's
lane; the exact contracts it must land are stated in the summary.

B1 — durable pending is reconciled independently of a fresh look. Tick settles
every active item's Pending BY IDENTITY at the top of the pass, before the rails
and before look(), so a quiet file watch, an advanced NextDue or a spent daily
rail can no longer strand an authorized line. The same identity is reused, so a
retry mints no new intent and adds no second ledger charge. A non-active item's
unresolved line is surfaced as NeedsPerson rather than forgotten; after
PendingGiveUp failed attempts an active one is too. fire no longer files a
second intent while one is unresolved (one() skips), and the pending cap refuses
rather than evicting.

B4 — a task has a durable in-flight marker. TaskInflight (run dir, started,
attempts) is written under the guarded save BEFORE Runner.Run and cleared only on
a terminal outcome; a failed or crashed attempt keeps it, NeedsPerson is set, and
one() refuses to run the task again while it is up. No exactly-once promise about
what a task did — the marker makes replay STOP, it does not make a run idempotent.

B7 — schema barrier. Schema is 3; SchemaOf writes 3 for any item carrying a
Pending or a TaskInflight marker, so a build that predates those fields (codeaf,
devaf, stageaf share one home) skips it instead of dropping the intent or
replaying the task. Ordinary and isolated orders keep versions 1 and 2, so
legacy reads and migrations are preserved. The revision guard is unchanged.

B8 — inbox crash/race/recovery/bounds. Deliver/Drain/seen mutations take a
per-inbox process flock (lock file in the system temp dir, keyed by the absolute
path, so session folders stay clean). Drain recovers leftover *.draining files
deterministically alongside a live inbox; a file it could not read whole (open,
read, or a line past inboxMaxLine) is KEPT and reported rather than deleted.
Deliver refuses a note larger than inboxMaxLine and an inbox past inboxMaxBytes
with visible backpressure. The guarantee is durable persistence and a recoverable
read handoff, not user display acknowledgement; readInbox reports scanner.Errors.

B9 — accounting and consent boundaries. A ledger/log append failure after a
successful delivery is an accountingError: Tick counts and logs it without
relabelling the item's outcome via noteFailure. Consent is checked before delivery
begins; a pause during an in-flight delivery is the person's act winning and is
never recorded as "could not check", and stale saves never undo it.

Fingerprint — boundedGlob bounds match ENUMERATION (directory entries capped,
collection capped) instead of materialising every match before the cap; a file
that vanished between listing and stat, or changed under the read, marks the
reading UNKNOWN (truncated) rather than "unchanged".

Tests: hardening_test.go adds failure-injection coverage — settle retries a
file-watch delivery under the same identity with no new look; a pending on a
paused/unexpired item is surfaced; an expired item's waiting line is honest; a
task that edits then errors is never replayed across a restart; a successful task
clears its marker; schema 3 excludes a ≤2 reader for pending items; orphan
.draining recovery; unreadable and oversized inboxes are kept and reported;
concurrent Deliver/Drain loses nothing; a delivery whose log cannot be written is
not relabelled; a pause mid-delivery is not claimed as cancelled; boundedGlob caps
enumeration; an unreadable file marks the fingerprint truncated. Focused suite
and -race green; go vet clean; session/cmd still build.

Assisted-by: CodeAF (deepseek-v4.1-flash)
Co-Authored-By: CodeAF <267109073+agentfield-bot@users.noreply.github.com>
…e tick, pinned model

Session and CLI half of the deferred-delivery core:

- standingRunner implements standing.Deliverer; new Deliver stamps the pending
  identity into the inbox Note and the note is written and Sync'd BEFORE the line
  is offered to an open window. Say/deliver return real address/inbox errors. The
  live handoff is at least once and that limit is documented; the durable note is
  what survives a window closing or a crash.
- Ratification arms a file watch at the yes (bounded retry); an incomplete scan
  sets no baseline, leaves a visible NeedsBaselineLead flag and the first complete
  look clears it.
- v3StandingTicker wires SentinelVerdict, and the parser reads whole tokens, so
  'yesterday'/'nothing' are unknown rather than yes/no.
- The ratifying session's one-model policy is frozen on the item origin, fenced
  at the schema barrier, and honoured by the reloaded background pass: the
  sentinel and every child text seat resolve from the item's pin with role pins,
  tiers, fallback chain and nearest-model guess withheld; the client cache is
  keyed by model AND policy.
- A paid undecided check persists its cost on the item without consuming the
  look (Store.NoteSpend), so the lifetime figure and the daily rail stay honest.
- Tests: real inbox write failure, identity dedup/restart, live-then-offline
  durability, arm incomplete visible, strict tri-state scoring, pinned model on
  the wire through a local fake provider (including a refusal that must not call
  another model), schema fence, folder-addressed lock across TEMP and symlink,
  bounded glob with a metacharacter segment, paid-unknown budget fence, real tick
  constructor.
Forward a worker's successful trusted full reads to the root's dependency
observation, not only its failures, under an immutable origin frozen at
worker creation (root session, user turn, goal, task, owner, run id). Fail
closed without that origin rather than reattributing a late arrival to the
live turn.

Make the collector's idempotency a bounded evictable hint backed by the
canonical journal (seen only after a successful append; failed writes stay
retryable and are exposed via errorState and the memory-failure lane), bound
per-origin emissions and origin bookkeeping (LRU, pending half-read pairs
protected), and dedup redundant repeats within one frozen origin while
preserving new-turn observations.

Route a promoted job's real ending through the root-owned collector for task
workers as well as sessions, freeze the launch turn/goal, include the run
identity beside the job id in the source key, refuse a missing frozen turn,
distinguish cancellation from execution failure, and admit through the same
in-flight gate as Close. Only the collector's owner may seal it.

forget now answers partial/error truthfully per owner and exact source
instead of swallowing read/write failures; a fresh independent attempt
remains learnable.
…ip diagnostics

Repair the concrete independent/root review failures at a6c2caf10:
- WhenFile: a truncated digest is never adopted as the baseline and is never
  compared with a complete one; a stored partial digest reads as no baseline.
  First complete look establishes the baseline and clears NeedsBaselineLead;
  later complete real changes fire.
- Drain: read everything, persist the dedup record, then remove. A seen/write/
  scanner/removal/read failure retains the payload and the dedup state; notes
  are returned for the caller to use even when an error is joined. IDs dedup
  across staged files in one drain. Deliver and Drain no longer read an
  unreadable or corrupt dedup record as empty.
- standingRunConfig: the OneModel pin suppresses a contradictory item model;
  unpinned items keep their explicit model.
- judge: a billed unknown whose ledger/NoteSpend write fails keeps the original
  WhenProbe NextDue instead of consuming it via noteFailure.
- Remove the unused live HandedOver API; state the at-least-once fold limit.
- ListChecked reports skipped newer/corrupt/unreadable documents; Tick and the
  wake log surface them (bounded) instead of passing by absence.
- standingArmBaseline arms once at ratification; the ordinary bounded Tick
  establishes an incomplete baseline on its first complete scan.
… producer effects, read-only standing binding

- Approved rules and confirmed decisions get their own bounded authority
  projection, so 128+ newer incidental observations cannot starve a live
  rare constraint; latest correction/demotion, expiry, conditions and
  suppression still win.
- Impact attention is driven by observed state, not a keyword list: a real
  producer change with an unchanged re-read consumer assumption is offered
  whatever the words were, and an unchanged file stays quiet.
- A delegated paired re-read of changed producer + unchanged consumer no
  longer resets the recorded baseline; it moves only when the consumer's
  own assumption changes, so the post-action check still sees the
  consequence.
- Dismissal is precise: named target, unambiguous singleton, or explicit
  batch; a bare ambiguous dismiss drops nothing and 'do not dismiss' is a
  negation. Held-offer state is bounded and evictable so new material can
  recur; dismiss identity persists as a canonical suppression.
- Derived receipts (history lookup, task summary, memory_evidence) can no
  longer become independent observation; genuine raw tool receipts and
  literal user intent are preserved. Suppressed history cannot re-enter as
  fresh proof.
- Authority and global scope are judged against the claim's own supporting
  span, not the whole utterance. Persisted/rendered evidence is passed
  through the secret redactor while literal validation reads the original
  words.
- A standing run authorized to act unattended is lent the owner-scoped
  canonical brain READ-ONLY and bound before its first action.
- Home filing Save failure is surfaced truthfully instead of claiming an
  unpersisted origin.
- Harness: optional --model/--one-model, explicit CODEAF_CALL_LOG_BODIES=1,
  dedicated pinned profile, fleet-first isolation and a credential-free
  metadata whitelist.
- Rename standing.Verdict to SentinelReading (one Verdict type law) and
  route the sentinel client through Config.clientConfig (gated-config law).
- Manual, docs and change entry 1777 updated; events_no_delete law intact.
Machine scope was kept for any supported user quote and its evidence was
pinned to the origin project, silently restricting an explicit machine-wide
rule. Machine scope now needs a self-contained machine-wide span, carries no
project condition (so it applies across projects on the same authorized
machine), and is never inferred from a tool observation. A project span beside
a global sentence and 'everywhere in this project' both stay project-local.
santoshkumarradha and others added 17 commits October 6, 2026 09:38
…se scope and redaction gaps

Scale: the binding projection asked for each candidate's latest evidence,
eligible guard and memory row in its own query. One indexed latest-by-memory
read (new events_contextual_memory expression index), one batched owner-scoped
eligibility read (canonical rows, derivation ancestry, memory suppression and
source retirement preloaded, evaluated by the SAME recursion the single-record
door uses) and one GetMemories call now answer the whole candidate set. The
per-memory latest read is not windowed: a live rule behind newer rows is still
answered, and the approved-rule floor stays on its partial index.

Complexity: every function in the owned contextual store and session memory
surface is at or under 15 decisions (batch helpers, filter/window splits,
shape/metadata validation splits, migration phase splits, binding fallback
split).

F2: scope widening is conservative. A bare "everywhere" stays project-local
unless the same clause names an explicit cross-project grant; project
qualifiers, exceptions and restrictions stay local. Regression cover for the
listed app/service/unless/project-local cases plus our repo/monorepo and
genuine explicit global positives, direct and through ground.

F3: validMemoryBody is the one mouth every memory write shares, and it now
redacts title, body and tags before the caps. Update and supersede store doors
are covered, with a direct-store regression.

Docs: PERF.md and docs/contextual-memory.md state the indexed/batched bounds and
the conservative scope rule truthfully; the journal still grows.
…e state from disk

F1: recordStandingFold built the fold with userText() and left authored false,
so the journal line was a plain PERSON message with an empty Deliveries list.
The at-arrival receipt the fix claims therefore did not exist: hasRecorded was
false for the fold's identity, a true crash before the ack folded the notice
twice, and the replayed page drew "while you were away" in the person's lane.
The shipped crash test only passed because its first drain had already ACKED
and left the id in the drained-identity record.

The fold now carries userMessage.authored through the existing note door, so
the journal keeps sessionEntry.Note + Deliveries together, replay draws it as
the session's own "aside", and no turn or model call is added. The person's
folder learns nothing: the authored branch never reaches stampUserLocked, so
Meta.LastUserAt and the placeholder title do not move.

F4 root: Store.Save unioned the caller's Pending into the disk document and
took TaskInflight only when the caller's was nil. A stale surface row could
therefore RESURRECT a delivery intent the ticker had already settled and pin an
older task in-flight marker over a newer one. Every production Save caller is a
surface act on an existing row (pause/stop/edit/exception/origin move) and none
authors those fields, so Save now copies both from the disk document verbatim;
the ticker's guarded saveActive stays the one runtime door. Tests that seeded
runtime state through Save now seed it through that door.

Regression tests (both fail on the base source): the fold's Note mark and
delivery id in a real journal plus a reconstructed crash-before-ack that folds
once, and stale-row Saves that must not resurrect a settled intent or an older
marker.
A: a shared interpreter is not shared work. The goal-any-file fallback now
requires a SHARED ACTUAL input operand when the failure demonstrated one; only a
failure that named no input file falls back to a goal file the success uses. The
real vendor.csv pair (failed utility read -> stdlib UTF-16 read of the same file)
still pairs; python parser_test.py -> python calc.py week.csv does not.

B: recognized wrappers and their options are normalized once through a spec
table (command -p/--/-v/-V, busybox -c, env -i/-u/...). command -v is the lookup
it is. An unknown wrapper option is refused rather than read as the program.

C: any top-level pipeline may mask its producer absent reliable pipefail
evidence, so a pipe into grep/awk masks like the formerly listed consumers; the
passive-consumer blacklist is gone. Quoted and heredoc pipes are not pipelines.

D: a submodule's snapshot identity is content-aware and bounded: its own HEAD
plus a hash of its changed and untracked file CONTENTS (--untracked-files=all),
recursing into nested submodules. A dirty/unreadable submodule propagates whole-
snapshot UNKNOWN instead of a stable ':unknown' string; the top-level status
reads with --ignore-submodules=none.

E: the dismissal path-token boundary includes the DOT (item.data.py no longer
names data.py) and batch dismissal needs an explicit notice-referent phrase
('dismiss all notices'), so 'that is all' drops nothing.

F: the persisted delegated SourceKey and AlternativeOf carry origin.Run, so two
workers of one node cannot collide; old opaque keys stay readable.

Also pay the last store gate debt: validateContextualDependency (22) is split
into five named field checks, each far under the ceiling. Regressions drive the
real recordOutcome/captureSourceSnapshot/dismissal paths.
…facts

An outcome fallback tied a failed action that named NO input file to any
success that used a file the frozen goal named. With a mixed goal such as
'validate parser.py and total week.csv' and a failure 'python -c import
pandas', a later calculation over week.csv could be paired as the answer to
that failure even though it says nothing about the parser.py work.

A no-input failure cannot say which goal artifact it was for, so the
fallback now requires the success to USE EVERY file the goal names: the
single-artifact goal is the ordinary case and the exact live compound pair
(week.csv AND ledger.py) still ties, while a success over one of several
artifacts fails CLOSED. A shared meaningful action token remains the
stronger tie and is decided before the fallback. A quoted file name with
spaces is no longer double-counted as a second artifact.

Also folds the four shell 'not the way work got done' refusals in
alternativeEligible into shellUnfitAsDemonstratedWork so the contextual
complexity gate measures 10 rather than 16, without weakening any rule.
…y ties

TestA4GuardTaxonomyPins pinned a bare pipeline into a consumer as
unmasked, contradicting shellMasksExit, which (correctly) refuses any
pipeline lacking its own pipefail. Pin the unsafe bare pipeline as MASKED
and leave the pipefail-cleared pipeline to the current policy gates.

The same run exposed a pre-existing failure: TestA4InterpreterTaxonomyIsExact
expected python2/python3/pypy to ground a goal-file operand, but the program
tie compared raw base words. Normalize a recognized interpreter word to its
family so the tie matches the closed interpreter taxonomy.
Batch retrieval is a regression, not a feature. 99328ddd6 replaced the
per-row evidence read in contextualEligibleFor with a verdict map, and
annotated every kept memory with verdict.evidence \u2014 including the
verdict of a memory that has NO journal row at all (verdict.has false).
That zero-valued struct wrote an empty "Memory id: ...\nEvidence: /"
block onto an ordinary legacy memory, so a plain "prefers tabs" record
rendered as if it carried contextual provenance and lost its legacy body
and age.

Absence is not a zero evidence: guard the annotation on verdict.has,
exactly as the two other annotate seams (bindingAuthorityFor and
bindingLexicalFallback) already do. An un-evidenced memory keeps its own
body and age; only a memory whose OWN latest row passed the guards is
annotated.

Regression: TestContextualMixedBatchAnnotatesOnlyMemoriesWithEvidence
projects one evidenced and one un-evidenced memory of the SAME exact
owner in one eligibility read, asserting the evidenced one keeps its
provenance and the un-evidenced one keeps its legacy body, age and an
unannotated rendered block. It fails on the pre-fix code with the exact
production symptom ("prefers tabs over spaces in Go\nMemory id: plain\nEvidence: /").

Contextual provenance, expiry, suppression and the 4800-rune ceiling are
untouched.
The fresh31ab run's goal named vendor.csv, but the bounded two-slot
<prior_outcomes> note rendered the newest relevant failure (a mixed probe
that read vendor.csv as utf-16 and then died on week.csv) plus the newest
remaining row, crowding out the demonstrated `ledger.py vendor.csv`
failure the goal was actually about.

priorOutcomeLines now ranks the relevant failures by evidence before
recency: the failure whose action actually USES a file the goal names,
weighted by how few non-goal, non-program inputs it also uses, so a newer
command that merely read the goal file on the way to failing elsewhere
cannot outrank a method that is squarely about the artifact. The second
slot prefers a DISTINCT method or input; when every remaining row repeats
the first row's method, a row carrying an observed success still renders,
and otherwise the slot stays empty rather than duplicating evidence.
Method identity falls back to the action's own text for actions with no
shell structure, so distinct failed tools are never merged as duplicates.
All reads, owner/suppression eligibility, the 2-slot bound and the shared
4800-rune ceiling are unchanged; no new cap and no new shell parser.

frameworkMethodPolicy's second exception is now attributed to the
person's own goal ("the person's own goal is itself to debug, test or
reproduce that method") rather than the ambiguous "the work is explicitly
to debug or test that method", closing the reading where a model's own
verification subtask qualified.

Focused tests drive the record/query/render path on the captured
representative pre-failure rows and on several newer paired variants, and
pin the preserved vs fixed selection, the alternative retention,
ownership/suppression, changed snapshots and distinct non-shell tools.
priorFailureMethodKey summarised a failure as its program families plus the
files it used, so two genuinely different invocations collapsed to one
identity and the second slot dropped a real distinct failure:
`ledger.py vendor.csv` and `ledger.py --strict vendor.csv` keyed the same, as
did two different inline programs with no operands, so only the newer of each
pair rendered.

The key is now the exact stored action -- tool name and full argument text,
trimmed -- rather than a semantic program-family+file-set summary. An
identical command recorded for two failures still dedups (the source key is
not part of the identity), while a different flag, inline program or wrapper
text stays a distinct row. Retaining an identical-looking wrapper as distinct
is the safe side of the trade: it never drops a real command, and the paired
observed alternative still rides its own failure line. The 2-slot bound, the
shared 4800-rune ceiling, owner/suppression eligibility and the ranking
heuristic are unchanged; only the lossy key and its now-unused workspace
parameter are removed.

Regressions drive the real priorOutcomeContext read seam and assert: same file
with different flags are two rows, different inline programs are two rows
(both fail on 5c7d59b), an identical command under different source keys
renders once, and the paired alternative survives the dedup.
…lure fits

The captured ledger note selected the exact failed row for the second
prior-outcome slot and then lost it at composition: the shared 4800-rune
ceiling was spent as 755 runes of framework policy plus a 2081-rune
approved-rule block, leaving 1964 for outcomes, while the two-record
<prior_outcomes> block measured 2136 and the trailing record was trimmed.

The ~167-rune caveat was appended to every observed alternative on top of
an equivalent ~235-rune block preamble. The caveat is now stated once in
the shared preamble (235 -> 221 runes), every semantic retained: quoted
history, untrusted, not instructions or authority, not causal or current
test proof, the user's current goal and words outrank them, a snapshot is
not the environment, and a working method holds only while conditions hold.
No cap, priority or authority order changed; whole-record trimming stays.

Regression at the whole compose/request boundary: a 2081-rune approved
block plus the captured two-record shapes now keeps both the working
alternative and the specific failure inside 4800, while the reconstructed
old block is shown to drop the failure.
…proposer's guess

A v4 watch reduced application readiness to curl -o /dev/null -w '%{http_code}'
whose output is 200 in both states, and the engine handed the model-authored
hint to the sentinel as WHAT A YES LOOKS LIKE with no provenance, so it fired
while the body said not ready. The sentinel now judges the person's own words;
a hint is a note of the proposer's guess, and a weaker fact than the ask (a
host answering where they asked whether an application is ready) is unknown.
The stand tool now says a probe must read what the person named, and the hint
field says reachability is not their condition. Prefix budget unchanged: the
description paid for itself out of a rails sentence and an op list the schema
already states.
d179d5aa1 made the person's sentence the sentinel's criterion and was still
wrong 10/10 times on the recorded shape: a status-only probe against a
readiness ask. The judging model was handed the output alone and inferred the
contract from it -- 'answered 200, so it is ready'. Nothing in the judgment said
the probe had thrown the body away, so no instruction could name what the look
could not see.

The judgment now carries the probe's own provenance -- the command it ran, or
the belt tool and arguments -- beside the evidence, under WHAT THE CHECK RAN,
read from the approved item and never re-derived; and the prompt states the rule
once: what a probe printed is all it saw, a status or reachability reading
proves transport and not the application contract the person named, and a
weaker witness is unknown unless the output itself carries what they named (a
file's contents, a JSON field, a real run's conclusion) or their own words made
that reading the criterion, as 'when it answers HTTP 200' does. The closing
question asked 'yes or no' where the prompt promised three answers; it now says
yes, no or unknown. The stand tool's probe sentence gains the read-only
qualifier the review asked for.

Measured at the production seam on deepseek/deepseek-v4.1-flash, pinned
OneModel: bare200_vs_readiness unknown 3/3 (was yes 6/6) and bare200_nohint
unknown 2/2 (was yes 2/2); explicit person-named HTTP 200 yes 2/2; JSON
false/true with the contradictory hint no/yes; already-ready first look yes;
missing body, failed command and clipped body unknown. The live matrix is an
opt-in e2e regression (internal/e2e/standing_sentinel_grounding_e2e_test.go),
skipped without a key, never a pass. The 231-line answer-encoding double is
gone: yes/no/unknown plumbing, clipping and probe dedup are covered where they
live, and the in-package witness keeps only what is mechanical.

Prefix budget unchanged and under: fixed 57150 B (cap 57218), lean 49572 B
(cap 49590); hint schema description 193 B (ceiling 200).
The regression fixture modelled the captured trailing failure WITHOUT its own
observed alternative, so the trailing record was roughly half its real size and
the test claimed the captured two-failure/two-alternative case fit after the
shared-caveat dedup when it does not. With the real mandatory approved-rule load
the two-record block exceeds the room left inside the 4800-rune ceiling, and the
whole-record trim correctly keeps the leading pair complete and omits the
trailing pair whole.

The fixture is now two failures that each carry their own working alternative on
anonymized relative paths, and the tests drive the real seams: with a mandatory
rule load sized from the contract the leading pair survives complete and the
trailing pair is absent whole with no orphan alternative, while both pairs fit
when the rule load is lower. A guard asserts each alternative actually indexes,
so an incomplete-pair fixture cannot pass silently. The brittle pre-dedup
reconstruction scaffolding is removed; the shared-caveat dedup stays covered by a
clearly synthetic two-pair test. No production code, cap, priority or authority
order changed.
…me-composition reviews

The entry carried the earlier wave but not the latest review fixes. It now
records, without duplicating existing lines: a condition watch handed the
sentinel the proposer's model-authored hint as the yes-shape with no
provenance, so a status-only probe against a readiness ask fired a false yes;
the judgment now carries the look's own command or belt tool and arguments,
read from the approved item and never re-derived, and a weaker witness than the
person named is unknown unless the output itself carries it or their own words
made that reading the criterion. The closing question now asks yes, no or
unknown, and the stand tool asks for a read-only look. The sentinel-grounding
matrix is now an opt-in, key-gated real-model e2e at the production seam that
skips rather than passes. The whole-request outcome-composition fixture no
longer claims both failure/alternative pairs always fit: inside the one shared
4,800-rune ceiling the leading pair is kept complete and a later pair is
omitted whole with no orphan alternative. No cap, default, priority, authority
or model changed. Final native terminal acceptance has not been run, and this
entry claims no native pass.

Docs only; no source touched.
The bounded <prior_outcomes> note showed at most two complete
failure/alternative pairs, but the captured rows repeated a bulky
`cd <workspace> &&` launcher on every action and the same source snapshot
on every row. Two genuine complete pairs measured 2446 runes against the
2156 left after the mandatory approved rules and the 755-rune policy, so
the whole-record trim dropped the trailing pair -- the literal
`ledger.py vendor.csv` failure and its working utf-16 alternative -- and
the model saw only the mixed probe whose 240-rune clip hides the extra
input.

The renderer now states a fact once where it is ACTUALLY shared: a source
snapshot and seen day that the raw structured rows genuinely share, and an
action launcher that every VISIBLE 240-rune preview shares. Two rows with
different commits that both render "different source snapshot" are never
unified; a shared span that stops inside a word, quoted argument or
heredoc is refused; the prefix is computed only over already-clipped
previews, so no clipped operand is inferred and no byte beyond the preview
is revealed; and the shortest honest candidate wins, with the old per-row
statement otherwise. Rows are still one physical line per complete pair,
success leading failure, every field escaped and clipped at 240, with the
cause advisory and reconsideration retained and no causal claim.

Measured read-only through the production functions against a narrow copy
of the captured owner database: the two genuine pairs are now 2094 runes
block-only (was 2446), and rules 1888 + outcomes 2094 + policy 755 = 4737
inside the one shared 4800 ceiling. No cap, preview window, slot count,
approved-rule or policy priority changed.

Tests drive the real composition seam with both captured-shape pairs and
assert the contract: both complete pairs fit under a representative
authority load, a long load still trims whole-record, distinct structured
sources/days keep per-row brackets, partial or quoted prefixes are
refused, and an injected marker in the shared prefix stays escaped.
The shared-prefix compaction asserted a preview prefix ended at a
top-level command boundary from a trailing `&& ` or `; `, balanced
quotes and no newline alone. That bounded test did not track $( ),
backticks or ( ) nesting, so a `; ` or `&& ` INSIDE an unquoted
substitution or grouping construct could be factored as if it were
top-level. The visible bytes still reconstructed, but the function and
the framing line promise a whole command.

Refuse these constructs conservatively instead of parsing shell: any
unquoted substitution or grouping opener outside single quotes makes
the prefix ineligible, so unknown syntax keeps the full per-row
actions. Quoted semicolons and literal non-ASCII launchers still
compact. No cap, preview window, slot count or authority priority
changed.
@santoshkumarradha
santoshkumarradha force-pushed the contextual-memory-20261005 branch from ea287a8 to d708504 Compare October 6, 2026 17:40
@santoshkumarradha
santoshkumarradha marked this pull request as ready for review October 6, 2026 18:45
santoshkumarradha pushed a commit that referenced this pull request Oct 6, 2026
… limits

Name the ambient side's real limits in the words a person asks with: an
app-closed watch has no helper to install, a once watch waits as an inbox note
rather than a notification and is delivered only once, a missing or disconnected
key queues no note and leaves the item active and due, and a check missed while
the computer was off or asleep is asked back once on return. Qualify the old
error-fix sentence so the prior-attempt advisory does not contradict it, and
replace 'a memory kept in one project is there in the next' with the closed
shelves: ownership is you, this workspace or this machine, while a manager, a new
member and a task in the workspace still read its scoped context. Add retrieval
probes for these questions and sync the #1777 entry.
@CLAassistant

CLAassistant commented Oct 6, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

… limits

Name the ambient side's real limits in the words a person asks with: an
app-closed watch has no helper to install, a once watch waits as an inbox note
rather than a notification and is delivered only once, a missing or disconnected
key queues no note and leaves the item active and due, and a check missed while
the computer was off or asleep is asked back once on return. Qualify the old
error-fix sentence so the prior-attempt advisory does not contradict it, and
replace 'a memory kept in one project is there in the next' with the closed
shelves: ownership is you, this workspace or this machine, while a manager, a new
member and a task in the workspace still read its scoped context. Add retrieval
probes for these questions and sync the #1777 entry.
…wers

The rerun of the native manual questions exposed four answers that were
materially wrong or unreachable. The judged-watch page now says a check needs
the credential for the model that watch judges on, not any connected provider
(reconnect OpenRouter when it serves that model, otherwise configure the model
and connect its matching service). The task and member pages now separate the
local, provider-free read of approved rules and confirmed decisions that binds
before the first request (bounded, read-only for workers) from the router's
advisory relevant-memory shortlist. The forgetting page now answers the old-
conversation search question in the words people use and grounds the guarantee
in suppression rather than in the search, and the loose 'every search and every
message' claims are narrowed to memory lists and searches. The catch-up text
now says the operating system is asked for one check and that it is not
guaranteed while the machine is off or asleep. Probes for all five questions
assert the section reached and its distinguishing content.
@santoshkumarradha
santoshkumarradha force-pushed the contextual-memory-20261005 branch from 0317833 to edd3a40 Compare October 6, 2026 19:27
q3's rerun answer got its main conclusion right \u2014 a remembered failure is
not a ban \u2014 but was materially wrong about the architecture and
overpromised: it said the line only reaches the model if it survives the
router's ranking and the small relevance model, and that with changed
inputs it WILL try the tool again.

The prior OBSERVED-outcome advisory is read locally from the saved
memories before the first request, ranked by its own relevance to the turn
and bounded to the relevant rows. It does not depend on the router's
ranking or its small relevance model. It bans no tool: the agent may
retry or adapt when the inputs, tree or circumstances change, and the
line never promises a retry.

what-i-remember now carries a heading in the asker's words so the actual
question and its paraphrases reach that section, opens with "No." so the
heading cannot be read as a yes-to-ban, and the generic router paragraph
names the one advisory that is not on its path. A retrieval probe asserts
the real question and a paraphrase reach the section and its
distinguishing content, not a shared topic.

The q2/q5 binding fixes are untouched, and there is no production change.
…t chat

A once watch's no-window paragraph and two neighbouring sections said a
firing with nothing open folds into the next conversation opened in the
project. The source walks four roads (internal/session/standing_run.go):
the origin conversation if live, any other live chat of the project, the
origin conversation's own inbox when nothing is open, and the project
inbox only for an item born in home's ask here exchange. An ordinary
chat's closed-inbox note is not handed to a new project chat. Name the
address for each origin, keep live-window routing as any same-project
chat, and carry the same correction into the PR 1777 change entry.
@santoshkumarradha
santoshkumarradha merged commit 3588adf into dev Oct 6, 2026
9 checks passed
@santoshkumarradha
santoshkumarradha deleted the contextual-memory-20261005 branch October 6, 2026 21:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants