Skip to content

perf(server): stop loading the whole database on hot orchestration paths - #297

Merged
leoisadev1 merged 6 commits into
mainfrom
perf/server-hot-paths
Sep 27, 2026
Merged

leoisadev1 merged 6 commits into
mainfrom
perf/server-hot-paths

Conversation

@leoisadev1

@leoisadev1 leoisadev1 commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

Several server hot paths did work proportional to the entire environment instead of the thread at hand. Starting a Codex or Kimi session loaded every project, bot, thread, message, and activity through a full projection snapshot just to find one thread. Every streamed assistant token re-read secret files from disk to check one settings flag. Each persisted event wrote all 14 projector cursors and reloaded the thread row, and checkpoint capture loaded a thread's full message history about four times per turn. Since node:sqlite is synchronous, these all block the event loop and grow with total chat history.

Session starts now use targeted by-id lookups with new narrow queries for delegations, groups, and bots. Materialized settings are cached and invalidated on writes, so secrets are read once instead of per token. Projectors declare the event types they handle and unrelated ones are skipped, with all cursors advancing in one batched write. Checkpoint capture uses the lightweight checkpoint-context query, ingestion caches the per-thread runtime context across content deltas, and the shell snapshot keeps open plus the 20 most recent finished delegations per thread with result text capped at 2,000 characters. Two behavior boundaries: the session-start lookup excludes archived and deleted threads, which the old snapshot included, and a thread whose project row is missing no longer gets a checkpoint.

Verified with 401 passing tests across the touched orchestration and provider suites, including new tests for the settings cache, projector skipping, and checkpoint linkage, plus 241 server, decider, and routine tests and a clean workspace typecheck. Fixing a pre-existing race in the ProviderCommandReactor test drain was required: the faster pipeline exposed that drain could return before subscribed events reached the worker.

Created with Claude Fable 5.1 in Claude Code.

🤖 Generated with Claude Code


Devin Review

Session starts, checkpoint capture, and per-event projection loaded full
projection snapshots or thread histories. They now use targeted by-id
queries, cache materialized settings so secrets are not re-read per token,
skip projectors that ignore an event, and cap shell delegation text.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@vercel

vercel Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
akeru-bot-landing Ready Ready Preview Sep 26, 2026 5:52pm UTC

Request Review

@github-actions github-actions Bot added type:provider Agent provider contribution. area:connectors Plugin and MCP connector runtime. size:XL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. labels Sep 26, 2026
devin-ai-integration[bot]

This comment was marked as resolved.

The shell snapshot now selects open delegations plus the newest terminal
ones per parent thread in SQL instead of hydrating every record, and
streamed delegation upserts carry the same capped result text. The
provider command reactor reads its drain baseline before subscribing so
buffered events are not marked as seen. The checkpoint test no longer
adds manual Effect runners, and the bot docs describe the card limit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
devin-ai-integration[bot]

This comment was marked as resolved.

@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Refactors database queries on hot orchestration paths.

No outstanding blocking findings; the PR is safe to merge based on the reviewed findings.

Summary

This PR replaces full-state reads on server orchestration paths with targeted queries, caches settings, reduces projection work, and limits delegation data sent to shell clients. The latest change makes reactor startup fail rather than lose provider work when it cannot replay events committed during startup.

Reviews (5) · Last reviewed commit: "fix(server): read the reactor startup ga..."

greptile-apps[bot]

This comment was marked as resolved.

@greptile-apps

This comment has been minimized.

The provider command reactor now retries its subscription until no
commit lands between the two sequence reads, so drain neither skips
buffered events nor waits for an event the subscription missed.
Concurrent settings reads after a change share one secret
materialization, and the shared shell reducer keeps the same number of
finished delegations per thread as the server snapshot.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
devin-ai-integration[bot]

This comment was marked as resolved.

greptile-apps[bot]

This comment was marked as resolved.

…order

The startup retry could drop provider intents that committed between its
sequence reads. Keep the first subscription and replay the gap from the
event store, skipping live events the replay already covered. The client
delegation cap now breaks updatedAt ties like the shell snapshot query.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
devin-ai-integration[bot]

This comment was marked as resolved.

greptile-apps[bot]

This comment was marked as resolved.

… failure

The startup gap replay stopped at the event store's default 1,000-event
limit, and a failed replay marked the whole gap as handled, so buffered
provider intents were filtered out. Replay the full bounded range, retry
from the last delivered event, and let the live stream deliver anything
the replay did not.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
devin-ai-integration[bot]

This comment was marked as resolved.

greptile-apps[bot]

This comment was marked as resolved.

A replay that failed after retries left drain unable to tell which gap
events the live buffer still held. Read the bounded gap eagerly during
start, retry transient failures, and fail startup if the store cannot
return it, so no intent is dropped and drain never waits on a lost event.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@leoisadev1
leoisadev1 merged commit 01161ce into main Sep 27, 2026
12 checks passed
@leoisadev1
leoisadev1 deleted the perf/server-hot-paths branch September 27, 2026 20:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:connectors Plugin and MCP connector runtime. size:XL type:provider Agent provider contribution. vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant