Skip to content

perf(runtime): trim per-turn reminders (budget thresholds, minute time, session id once) - #459

Merged
SaladDay merged 1 commit into
mainfrom
jack/p1-8-turn-reminders
Oct 9, 2026
Merged

SaladDay merged 1 commit into
mainfrom
jack/p1-8-turn-reminders

Conversation

@hetaoBackend

@hetaoBackend hetaoBackend commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Problem (P1-8, 0.6.5)

Two reminders spend tokens on every request or turn:

  1. --timeout budget reminder (execution-budget-reminder.ts). This is request-only and appended to the end of every LLM request, so its ~40 tokens are never cached: Execution time remaining at request preparation: N seconds. Since the previous model request was prepared: M seconds ....
  2. Per-turn <agent-context>. This block is built in packages/agent-modules/system-reminder/src/blocks.ts buildSlimAgentContextBlock. The hosted reminder service in packages/local-runtime (hosted-agent-capabilities.ts → LocalSystemReminderService → LocalDataCollector) reaches it through v2 LocalTurnInputPreparer → reminders.buildSystem. The block is placed at the start of each turn's durable user message. On every turn ≥2 it repeats YOUR SESSION ID and date: ${env.date}. env.date was new Date(now).toString() (host-memory.ts), which has second precision.

Change

(a) Budget reminder: emit only at thresholds

  • Emit on the first request, then once each time the remaining budget first drops to ≤50%, ≤25% or ≤10% of the total. Inside the final 2 minutes, emit on every request.
  • The total budget is measured from hook creation (turn preparation), because requests carry no start time.
  • Coarser text: whole minutes while more than 2 min remain; inside the final 2 min, seconds rounded down to 10 s (under 10 seconds at the end).
  • Removed the "since the previous model request" sentence.
  • If admission (canAppendExecutionBudgetReminder) rejects a reminder, that threshold is retried on the next request.
  • Unchanged: the caller's cancellation timer stays authoritative, the reminder stays request-only, and the hook still runs before the appending background reminder.

(b) <agent-context>: minute-precision time, session ID only when needed

  • date: is now Date#toString() without the seconds, e.g. Fri Oct 09 2026 13:23 GMT+0800 (China Standard Time).
  • v2 preparation checks the provider-facing history for a user message that already contains this session's ID. The scan only covers messages after the latest compactionSummary (or legacy archonCompaction marker). If one exists, it passes sessionIdInContext: true → compat adapter → hosted buildSystem → LocalReminderSessionInfo → AgentEnv. The slim block then omits the YOUR SESSION ID line.
  • As a result the ID appears on the first turn, which always uses the full block. It reappears on the first turn after a compaction summarized it away, and it stays omitted while a kept message still shows it. The model can always find it in context.
  • The cloud scene and other callers never set the flag, so their behaviour is unchanged. The full first-turn block always includes the ID.

Design notes

Goal. Cut tokens that are re-sent on every request or turn without changing what the model can do, and without touching the stable cached prefix (system prompt, tool list, earlier history).

Why the two reminders are treated differently

  • The budget reminder is request-only and goes at the end of each request, so it is never part of the cached prefix. Every emission is pure uncached input. Fewer emissions is a direct saving, and changing its text or frequency cannot invalidate any cache.
  • <agent-context> is durable. Once written into a turn's user message it never changes, so older turns stay byte-identical and the prefix cache keeps hitting. We only change what is written into new turns. Nothing already in history is rewritten.

Budget reminder: why thresholds

  • The model only needs to know about the budget when it can act on it. A reminder at 50%, 25% and 10%, plus every request in the final 2 minutes, gives early notice and a dense warning when it matters.
  • Text is coarser (whole minutes, then 10 s steps) so consecutive reminders do not differ on every request for no reason.
  • A threshold rejected by admission is retried on the next request, so a reminder is not silently lost.
  • The caller's cancellation timer stays authoritative. The reminder is advisory only.

Session ID: why "only when not already in context"

  • The ID is needed by tools and skills only if the model can find it somewhere in context. We keep it in the first turn's full block and re-add it only when compaction has removed every message that contained it.
  • Detection is content-based over the provider-facing history after the latest compaction marker. The alternative, an explicit compaction-epoch fact in v2, would need a new contract between compaction and the preparer. We chose the content check because it needs no new interface and fails safe: if unsure, the ID is included.
  • Cloud and other callers never set the flag, so their behavior is unchanged.

Why minute precision for date:

  • Seconds change on every turn but carry no value for the model, and they make otherwise identical blocks differ.

Alternatives considered

  • Dropping the budget reminder entirely: rejected, the model would no longer see the deadline coming.
  • Emitting only once per threshold, with no first-request reminder: kept as an open question below.
  • Removing the session ID from all turns: rejected, tools may rely on it being findable after compaction.

Risks

  • A tool or skill that expects the ID in every turn's reminder (open question 4).
  • Content-based detection could miss an ID that is only present in a tool result. That fails safe by re-adding the ID.

Tests

Run with Node 22.23.3; the repo engines require ≥22.19.

  • node scripts/run-vitest-suite.mjs capability: 215 files / 5397 tests passed, 18 skipped. This includes:
    • executor.test.ts: rewritten budget tests (threshold steps for 10 min and 2 min budgets, admission retry, skipping straight to the deepest threshold, updated durable-reminder text) and new sessionIdInContext cases (empty history, string and text-part content, compaction summary, legacy marker, kept message after compaction).
    • compat/v1/agent-host.test.ts: checks the flag is forwarded.
    • new packages/local-runtime/test/unit/turn-reminder-context.test.ts: minute formatter, slim and full block, cloud unchanged. Registered in test/vitest-suites.json and release/public-source.json.
  • pnpm typecheck passed. pnpm lint passed. node scripts/source-inventory.mjs passed.
  • Not run: check:standalone needs dist/metafile.json from a build.

Token measurement

Method: a script runs the actual old (origin/main) and new hook code and the real buildSlimAgentContextBlock, then counts tokens with gpt-tokenizer 3.4.0 (already a repo dependency; o200k-family BPE). That is a proxy, not the MiniMax tokenizer, so treat the numbers as relative.

Scenario Before After
(a) --timeout 30m, 40 LLM requests evenly spaced (last one 15 s before the deadline) 40 reminders, 1609 tokens (all uncached, end of request) 7 reminders, 173 tokens
(b) 20-turn session, slim block on turns 2–20, no compaction 1235 tokens (65/turn) 969 tokens (51/turn)

In (b), the session ID line is about 11 tokens and dropping the seconds saves about 2. Because (b) is durable in history, most of the saving lands on the prompt-cache write and on context growth, not on every request.

Open questions (for Tao)

  1. Final 2 minutes: I chose to emit on every request there (coarse, 10 s steps), plus one announcement on the first request. That first-request announcement was not in the spec. Should it be strictly once per threshold, with no first-request reminder?
  2. Budget start: the total is measured from hook creation, i.e. turn preparation, not CLI start. For a continuation turn, the "total" is whatever remained when that turn started. OK?
  3. Session ID detection is content-based: it looks for the session ID string in user messages since the last compaction. The alternative is an explicit compaction-epoch fact in v2, which would need a new contract. Is this acceptable?
  4. Does any tool or skill rely on YOUR SESSION ID being present in every turn's reminder, not just somewhere in context?

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

…e, session id once)

P1-8: two reminders were spending tokens on every request/turn.

--timeout execution budget (request-only, appended per LLM request):
- Emit only on the first request, when remaining budget first drops to
  <=50%, <=25%, <=10% of the total, and on every request inside the final
  two minutes. Total budget is measured from hook creation (turn prep).
- Coarse text: whole minutes above 2 min, 10-second steps inside it.
- Drop the "since the previous model request" sentence.
- Admission rejection does not consume a threshold. Caller's cancellation
  timer stays authoritative; hook ordering constraint unchanged.

Per-turn <agent-context> system reminder (local runtime):
- date: Date#toString() without seconds (minute precision).
- The slim (turn >= 2) block omits YOUR SESSION ID when the provider-facing
  history since the latest compaction already contains a user message with
  this session's ID. First turn and the first turn after a compaction that
  dropped it still carry the ID. Full first-turn block and cloud scene are
  unchanged (flag absent => old behaviour).
@hetaoBackend
hetaoBackend marked this pull request as ready for review October 9, 2026 06:17
@hetaoBackend
hetaoBackend requested a review from SaladDay October 9, 2026 06:20

@yujiachen-y yujiachen-y left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants