Skip to content

Commit 4aa36af

Browse files
author
rodrigedilson-ia
committed
fix(compiler): cap concept/entity brief lists in the plan prompt
`_read_concept_briefs` and `_read_entity_briefs` read every page in the KB and the result is rebuilt into the concepts-plan prompt for each compiled document. Their size is therefore O(KB), with no ceiling — so a large KB eventually pushes that call past the model's context window. The failure mode is unkind: it lands as a hard error partway through a recompile, not as degraded output. We hit it on a KB of ~260 documents, where seven consecutive `recompile` runs exited non-zero, all of them at the plan step, with prompts measured between 200k and 356k tokens against a 200k limit. Because it scales with the KB, every retry was guaranteed to fail the same way. This adds a per-list character budget (`briefs_budget_chars`, default 120_000 — roughly 30k tokens each, so a KB has to grow well past a few hundred pages before anything is trimmed). Set it to 0 to restore the previous uncapped behavior. Two details worth calling out: - **Ranking is by source count, not alphabetical.** That count is already the cross-document recurrence signal the plan call uses for create-vs-update, so it is also the right thing to keep when the list has to be cut. Ties break alphabetically, because prompt caching depends on byte-identical prefixes and an unstable order would silently cost cache hits. - **The trim is announced in the list.** A shortened list that reads as complete is worse than a long one: the planner uses these briefs to decide create-vs-update, so a silently dropped concept comes back as a duplicate page for something the KB already has. The trailing marker states how many were omitted and tells the model to prefer `update` when unsure. In the degenerate case where the budget cannot fit even one line, the output says the list was omitted rather than returning an empty one — an empty list reads as an empty KB, which would push the planner to recreate everything. Small KBs are byte-identical to before (covered by a test), so prompt caches are not invalidated for existing users.
1 parent ff54396 commit 4aa36af

3 files changed

Lines changed: 261 additions & 12 deletions

File tree

‎openkb/agent/compiler.py‎

Lines changed: 132 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -724,14 +724,117 @@ def _resolve_description(fm: dict) -> str:
724724
return ""
725725

726726

727-
def _read_concept_briefs(wiki_dir: Path) -> str:
727+
DEFAULT_BRIEFS_BUDGET_CHARS = 120_000
728+
"""Per-list character cap for the concept/entity briefs in the plan prompt.
729+
730+
Roughly 30k tokens each, so a KB has to grow well past a few hundred pages
731+
before anything is trimmed — the cap is a ceiling against unbounded growth, not
732+
a working limit. Overridable with the ``briefs_budget_chars`` config key; set it
733+
to 0 to restore the previous uncapped behavior.
734+
735+
Why a cap at all: ``_read_concept_briefs`` and ``_read_entity_briefs`` read
736+
*every* page in the KB and the result is rebuilt into the plan prompt for each
737+
compiled document. The size is therefore O(KB), and a large KB eventually
738+
exceeds the model's context window — which surfaces as a hard failure partway
739+
through a recompile, not as degraded output.
740+
"""
741+
742+
743+
def _resolve_briefs_budget(config: dict) -> int:
744+
"""Read ``briefs_budget_chars`` from config, falling back to the default.
745+
746+
A non-integer or negative value falls back rather than raising: a typo in a
747+
config file should not abort a compile, and a negative budget has no sane
748+
reading. ``0`` is meaningful and passes through — it disables trimming.
749+
"""
750+
raw = config.get("briefs_budget_chars", DEFAULT_BRIEFS_BUDGET_CHARS)
751+
if isinstance(raw, bool) or not isinstance(raw, int) or raw < 0:
752+
logger.warning(
753+
"briefs_budget_chars=%r is not a non-negative integer — using %d",
754+
raw,
755+
DEFAULT_BRIEFS_BUDGET_CHARS,
756+
)
757+
return DEFAULT_BRIEFS_BUDGET_CHARS
758+
return raw
759+
760+
761+
def _n_sources(fm_dict: dict) -> int:
762+
"""How many documents a page cites, from its ``sources:`` frontmatter list.
763+
764+
This is the cross-document recurrence signal already used in the entity
765+
brief line. It doubles as the salience ranking when the brief list has to
766+
be trimmed to fit a budget: a concept seen in many documents is the one the
767+
planner most needs to know about, because it is the one most likely to be
768+
updated rather than created.
769+
"""
770+
sources = fm_dict.get("sources")
771+
return len(sources) if isinstance(sources, list) else 0
772+
773+
774+
def _fit_briefs(ranked: list[tuple[int, str]], budget_chars: int, noun: str) -> str:
775+
"""Join brief lines under a character budget, ranked by salience.
776+
777+
``ranked`` is ``[(n_sources, line), ...]``. Lines are emitted most-cited
778+
first, then alphabetically (the caller sorts), until the budget is spent.
779+
780+
**The truncation is announced in the returned text.** A shortened list that
781+
looks complete is worse than a long one: the planner reads these briefs to
782+
decide create-vs-update, so a silently dropped concept comes back as a
783+
duplicate page for something the KB already has. The trailing marker tells
784+
the model the list is partial, so "not listed" stops meaning "not present".
785+
786+
A budget of 0 or less disables trimming, which keeps the previous behavior
787+
available for callers that want the full list.
788+
"""
789+
if budget_chars <= 0:
790+
return "\n".join(line for _, line in ranked) or "(none yet)"
791+
792+
kept: list[str] = []
793+
used = 0
794+
for _, line in ranked:
795+
if used + len(line) + 1 > budget_chars:
796+
break
797+
kept.append(line)
798+
used += len(line) + 1
799+
800+
dropped = len(ranked) - len(kept)
801+
if not dropped:
802+
return "\n".join(kept) or "(none yet)"
803+
if not kept:
804+
# Budget smaller than a single line. Say so rather than return an empty
805+
# list that reads like an empty KB.
806+
return f"(list omitted: {len(ranked)} {noun} exceed the brief budget)"
807+
808+
logger.info(
809+
"brief list trimmed to fit budget: kept %d of %d %s (%d chars)",
810+
len(kept),
811+
len(ranked),
812+
noun,
813+
budget_chars,
814+
)
815+
kept.append(
816+
f"- (… {dropped} more {noun} not listed — this list is truncated to the "
817+
f"{len(kept) - 1} most-cited; absence here does NOT mean the page is "
818+
f"missing, so prefer 'update' over 'create' when unsure)"
819+
)
820+
return "\n".join(kept)
821+
822+
823+
def _read_concept_briefs(wiki_dir: Path, budget_chars: int = 0) -> str:
728824
"""Read existing concept pages and return compact one-line summaries.
729825
730826
For each concept, reads the ``description:`` field (falling back to legacy
731827
``brief:``) from YAML frontmatter if present; otherwise falls back to
732828
truncating the first 150 chars of the body (newlines collapsed to spaces).
733829
Formats each as ``- {slug}: {description}``.
734830
831+
With ``budget_chars`` > 0 the list is capped at that many characters, most-
832+
cited concepts first (see :func:`_fit_briefs`). The cap exists because this
833+
list grows with the KB and lands in the plan prompt on every compiled
834+
document: an unbounded list eventually exceeds the model's context window,
835+
and the failure arrives as a hard error mid-compile rather than as degraded
836+
output.
837+
735838
Returns "(none yet)" if the concepts directory is missing or empty.
736839
"""
737840
concepts_dir = wiki_dir / "concepts"
@@ -742,7 +845,7 @@ def _read_concept_briefs(wiki_dir: Path) -> str:
742845
if not md_files:
743846
return "(none yet)"
744847

745-
lines: list[str] = []
848+
ranked: list[tuple[int, str]] = []
746849
for path in md_files:
747850
text = path.read_text(encoding="utf-8")
748851
fm_dict = frontmatter.parse(text)
@@ -752,17 +855,22 @@ def _read_concept_briefs(wiki_dir: Path) -> str:
752855
body = parts[1] if parts is not None else text
753856
brief = body.strip().replace("\n", " ")[:150]
754857
if brief:
755-
lines.append(f"- {path.stem}: {brief}")
858+
ranked.append((_n_sources(fm_dict), f"- {path.stem}: {brief}"))
756859

757-
return "\n".join(lines) or "(none yet)"
860+
# Most-cited first; alphabetical within a tier so the prompt stays stable
861+
# across runs (prompt caching depends on byte-identical prefixes).
862+
ranked.sort(key=lambda item: (-item[0], item[1]))
863+
return _fit_briefs(ranked, budget_chars, "concepts")
758864

759865

760-
def _read_entity_briefs(wiki_dir: Path) -> str:
866+
def _read_entity_briefs(wiki_dir: Path, budget_chars: int = 0) -> str:
761867
"""Read existing entity pages as compact lines for the plan call.
762868
763869
Formats each as ``- {slug} ({type}, {n} sources) — {brief}``. The source
764870
count is the cross-document recurrence signal the LLM uses to decide
765-
create-vs-update and salience. Returns "(none yet)" when empty.
871+
create-vs-update and salience — and, with ``budget_chars`` > 0, the ranking
872+
used to decide what stays when the list is capped. Returns "(none yet)"
873+
when empty.
766874
"""
767875
entities_dir = wiki_dir / "entities"
768876
if not entities_dir.exists():
@@ -772,21 +880,22 @@ def _read_entity_briefs(wiki_dir: Path) -> str:
772880
if not md_files:
773881
return "(none yet)"
774882

775-
lines: list[str] = []
883+
ranked: list[tuple[int, str]] = []
776884
for path in md_files:
777885
text = path.read_text(encoding="utf-8")
778886
fm_dict = frontmatter.parse(text)
779887
brief = _resolve_description(fm_dict)
780888
etype = str(fm_dict.get("type") or "").strip().lower() or "other"
781-
n_sources = len(fm_dict["sources"]) if isinstance(fm_dict.get("sources"), list) else 0
889+
n_sources = _n_sources(fm_dict)
782890
if not brief:
783891
parts = frontmatter.split(text)
784892
body = parts[1] if parts is not None else text
785893
brief = body.strip().replace("\n", " ")[:150]
786894
suffix = f" — {brief}" if brief else ""
787-
lines.append(f"- {path.stem} ({etype}, {n_sources} sources){suffix}")
895+
ranked.append((n_sources, f"- {path.stem} ({etype}, {n_sources} sources){suffix}"))
788896

789-
return "\n".join(lines) or "(none yet)"
897+
ranked.sort(key=lambda item: (-item[0], item[1]))
898+
return _fit_briefs(ranked, budget_chars, "entities")
790899

791900

792901
def _iter_h2_headings(lines: list[str]) -> list[tuple[int, str]]:
@@ -1603,6 +1712,7 @@ async def _compile_concepts(
16031712
doc_type: str = "short",
16041713
rewrite_summary: bool = False,
16051714
entity_types: list[str] | None = None,
1715+
briefs_budget_chars: int | None = None,
16061716
bundle=None,
16071717
) -> None:
16081718
"""Shared Steps 2-4: concepts plan → generate/update → index.
@@ -1624,8 +1734,14 @@ async def _compile_concepts(
16241734
valid_types = frozenset(entity_types)
16251735

16261736
# --- Step 2: Get concepts plan (A cached) ---
1627-
concept_briefs = _read_concept_briefs(wiki_dir)
1628-
entity_briefs = _read_entity_briefs(wiki_dir)
1737+
# Both lists grow with the KB and are rebuilt into the plan prompt for every
1738+
# compiled document, so their combined size is what eventually pushes this
1739+
# call past the model's context window. The budget caps each one; see
1740+
# ``_fit_briefs`` for why the trim is announced rather than silent.
1741+
if briefs_budget_chars is None:
1742+
briefs_budget_chars = DEFAULT_BRIEFS_BUDGET_CHARS
1743+
concept_briefs = _read_concept_briefs(wiki_dir, briefs_budget_chars)
1744+
entity_briefs = _read_entity_briefs(wiki_dir, briefs_budget_chars)
16291745

16301746
# Second cache breakpoint: end of the assistant summary message. Covers
16311747
# (system + doc + summary) for the plan call and every concept call.
@@ -2220,6 +2336,7 @@ async def compile_short_doc(
22202336
config = resolve_effective_config(kb_dir)[0]
22212337
language: str = config.get("language", "en")
22222338
entity_types = resolve_entity_types(config)
2339+
briefs_budget = _resolve_briefs_budget(config)
22232340

22242341
wiki_dir = kb_dir / "wiki"
22252342
schema_md = get_agents_md(wiki_dir)
@@ -2280,6 +2397,7 @@ async def compile_short_doc(
22802397
doc_type="short",
22812398
rewrite_summary=True,
22822399
entity_types=entity_types,
2400+
briefs_budget_chars=briefs_budget,
22832401
bundle=bundle,
22842402
)
22852403
finally:
@@ -2308,6 +2426,7 @@ async def compile_long_doc(
23082426
config = resolve_effective_config(kb_dir)[0]
23092427
language: str = config.get("language", "en")
23102428
entity_types = resolve_entity_types(config)
2429+
briefs_budget = _resolve_briefs_budget(config)
23112430

23122431
wiki_dir = kb_dir / "wiki"
23132432
schema_md = get_agents_md(wiki_dir)
@@ -2364,6 +2483,7 @@ async def compile_long_doc(
23642483
doc_brief=doc_description,
23652484
doc_type="pageindex",
23662485
entity_types=entity_types,
2486+
briefs_budget_chars=briefs_budget,
23672487
bundle=bundle,
23682488
)
23692489
finally:

‎openkb/config.py‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -36,6 +36,11 @@
3636
# global/KB list overrides it wholesale; resolve_entity_types cleans the
3737
# effective value on read.
3838
"entity_types": list(DEFAULT_ENTITY_TYPES),
39+
# Per-list character cap for the concept/entity briefs in the compile plan
40+
# prompt. Those lists are O(KB) and are rebuilt for every compiled document,
41+
# so without a ceiling a large KB eventually exceeds the model's context
42+
# window mid-recompile. 0 disables trimming.
43+
"briefs_budget_chars": 120_000,
3944
}
4045

4146
GLOBAL_CONFIG_DIR = Path.home() / ".config" / "openkb"

‎tests/test_compiler.py‎

Lines changed: 124 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,7 @@
99

1010
from openkb.agent.compiler import (
1111
_ENTITY_TYPE_LIST,
12+
DEFAULT_BRIEFS_BUDGET_CHARS,
1213
_add_related_link,
1314
_backlink_concepts,
1415
_backlink_entities,
@@ -23,6 +24,7 @@
2324
_read_entity_briefs,
2425
_read_wiki_context,
2526
_remove_source_from_frontmatter,
27+
_resolve_briefs_budget,
2628
_sanitize_concept_name,
2729
_update_index,
2830
_write_concept,
@@ -2830,3 +2832,125 @@ def test_concept_update_malformed_frontmatter_rebuilds(self, tmp_path):
28302832
assert 'type: "Concept"' in text
28312833
# Must have a properly closed frontmatter block (two '---' occurrences).
28322834
assert text.count("---") >= 2
2835+
2836+
2837+
class TestBriefsBudget:
2838+
"""The concept/entity brief lists are O(KB) and are rebuilt into the plan
2839+
prompt for every compiled document. Without a ceiling a large KB eventually
2840+
exceeds the model's context window, and the failure arrives as a hard error
2841+
partway through a recompile rather than as degraded output.
2842+
"""
2843+
2844+
@staticmethod
2845+
def _concept(wiki, slug, description, n_sources=1):
2846+
d = wiki / "concepts"
2847+
d.mkdir(parents=True, exist_ok=True)
2848+
sources = ", ".join(f'"summaries/doc{i}.md"' for i in range(n_sources))
2849+
(d / f"{slug}.md").write_text(
2850+
f'---\ntype: "Concept"\nsources: [{sources}]\n'
2851+
f'description: "{description}"\n---\n\n# {slug}\n',
2852+
encoding="utf-8",
2853+
)
2854+
2855+
def test_under_budget_is_unchanged(self, tmp_path):
2856+
"""Small KBs must see byte-identical output — the cap is a ceiling, not
2857+
a working limit, and the plan prompt is prompt-cached on exact prefixes.
2858+
"""
2859+
wiki = tmp_path / "wiki"
2860+
self._concept(wiki, "alpha", "first")
2861+
self._concept(wiki, "beta", "second")
2862+
2863+
assert _read_concept_briefs(wiki, 120_000) == _read_concept_briefs(wiki, 0)
2864+
2865+
def test_budget_keeps_most_cited_first(self, tmp_path):
2866+
"""Source count is the recurrence signal the planner uses for
2867+
create-vs-update, so it is also the right thing to keep when trimming.
2868+
"""
2869+
wiki = tmp_path / "wiki"
2870+
self._concept(wiki, "rare", "x" * 60, n_sources=1)
2871+
self._concept(wiki, "common", "y" * 60, n_sources=9)
2872+
2873+
out = _read_concept_briefs(wiki, 90)
2874+
2875+
assert "common" in out
2876+
assert "- rare:" not in out
2877+
2878+
def test_truncation_is_announced(self, tmp_path):
2879+
"""A shortened list that looks complete is worse than a long one: the
2880+
planner reads these briefs to decide create-vs-update, so a silently
2881+
dropped concept comes back as a duplicate page for something the KB
2882+
already has.
2883+
"""
2884+
wiki = tmp_path / "wiki"
2885+
for i in range(12):
2886+
self._concept(wiki, f"concept-{i:02d}", "z" * 80, n_sources=i)
2887+
2888+
out = _read_concept_briefs(wiki, 300)
2889+
2890+
assert "not listed" in out
2891+
assert "prefer 'update' over 'create'" in out
2892+
2893+
def test_budget_zero_disables_trimming(self, tmp_path):
2894+
wiki = tmp_path / "wiki"
2895+
for i in range(30):
2896+
self._concept(wiki, f"c{i:02d}", "w" * 100)
2897+
2898+
out = _read_concept_briefs(wiki, 0)
2899+
2900+
assert len(out.splitlines()) == 30
2901+
assert "not listed" not in out
2902+
2903+
def test_budget_smaller_than_one_line_says_so(self, tmp_path):
2904+
"""An empty string would read as an empty KB, which is a different
2905+
thing entirely and would push the planner to recreate everything."""
2906+
wiki = tmp_path / "wiki"
2907+
self._concept(wiki, "alpha", "x" * 200)
2908+
2909+
out = _read_concept_briefs(wiki, 5)
2910+
2911+
assert "list omitted" in out
2912+
assert out != "(none yet)"
2913+
2914+
def test_ordering_is_stable_within_a_tier(self, tmp_path):
2915+
"""Prompt caching depends on byte-identical prefixes, so equal-salience
2916+
entries must not reorder between runs."""
2917+
wiki = tmp_path / "wiki"
2918+
for slug in ("gamma", "alpha", "beta"):
2919+
self._concept(wiki, slug, "same", n_sources=2)
2920+
2921+
primeira = _read_concept_briefs(wiki, 120_000)
2922+
2923+
assert primeira == _read_concept_briefs(wiki, 120_000)
2924+
assert primeira.index("alpha") < primeira.index("beta") < primeira.index("gamma")
2925+
2926+
def test_entity_briefs_respect_the_same_budget(self, tmp_path):
2927+
wiki = tmp_path / "wiki"
2928+
d = wiki / "entities"
2929+
d.mkdir(parents=True)
2930+
for i in range(10):
2931+
(d / f"e{i:02d}.md").write_text(
2932+
f'---\ntype: "organization"\nsources: ["summaries/d{i}.md"]\n'
2933+
f'description: "{"q" * 90}"\n---\n\n# e{i:02d}\n',
2934+
encoding="utf-8",
2935+
)
2936+
2937+
out = _read_entity_briefs(wiki, 250)
2938+
2939+
assert "not listed" in out
2940+
assert len(out.splitlines()) < 10
2941+
2942+
2943+
class TestResolveBriefsBudget:
2944+
"""A typo in a config file must not abort a compile."""
2945+
2946+
def test_default_when_absent(self):
2947+
assert _resolve_briefs_budget({}) == DEFAULT_BRIEFS_BUDGET_CHARS
2948+
2949+
def test_explicit_zero_passes_through(self):
2950+
"""0 is meaningful — it disables trimming — so it must not be treated
2951+
as 'unset'."""
2952+
assert _resolve_briefs_budget({"briefs_budget_chars": 0}) == 0
2953+
2954+
@pytest.mark.parametrize("bad", ["lots", -1, None, 1.5, True])
2955+
def test_invalid_falls_back(self, bad):
2956+
assert _resolve_briefs_budget({"briefs_budget_chars": bad}) == DEFAULT_BRIEFS_BUDGET_CHARS

0 commit comments

Comments
 (0)