Every history pass still does O(sources) reads on a corpus of more than 1,000 Codex rollouts; this is what remains after #2786.
What happens
- Each pass re-reads the
session_meta head of every unchanged rollout. Meta is needed for the source identity, and the cursor lookup and the unchanged-generation lookup both depend on that identity.
- The meta cache (
CodexMetaFillClaim::publish) holds only shared_jsonl_preparation_capacity() entries, which is the CPU width, so it thrashes on any real corpus.
UNCHANGED_GENERATION_CACHE_CAP (4096) would thrash the same way past 4096 transcripts.
Measurement
Isolated daemon on an 8× store (about 60.6k messages, 800 rollouts), after #2786:
- raw reads are 3.11 MB per message, against 2.05 MB at the base store (about 7.6k messages);
- while settled, idle reads are 253 KB/s.
Fix direction
Skip sources whose (path, file identity, length, mtime) are unchanged at discovery, before deriving meta or identity, rather than raising either cache cap.
Every history pass still does O(sources) reads on a corpus of more than 1,000 Codex rollouts; this is what remains after #2786.
What happens
session_metahead of every unchanged rollout. Meta is needed for the source identity, and the cursor lookup and the unchanged-generation lookup both depend on that identity.CodexMetaFillClaim::publish) holds onlyshared_jsonl_preparation_capacity()entries, which is the CPU width, so it thrashes on any real corpus.UNCHANGED_GENERATION_CACHE_CAP(4096) would thrash the same way past 4096 transcripts.Measurement
Isolated daemon on an 8× store (about 60.6k messages, 800 rollouts), after #2786:
Fix direction
Skip sources whose (path, file identity, length, mtime) are unchanged at discovery, before deriving meta or identity, rather than raising either cache cap.