Skip to content

Repository files navigation

waymark

A knowledge base that lives next to your source and is queried from the command line.

A waymark is a marker left on a route so the next traveller does not have to work it out again. That is the whole idea: the expensive thing in a long-lived codebase is not writing code, it is re-deriving what somebody already learned — which hypothesis was tested and died, why a function looks wrong but isn't, what the wire actually carries as opposed to what the config claims.

waymark indexes your code, and lets you attach durable notes, claims with evidence, and cross-references to it. Then it checks itself, because a knowledge base nobody trusts is worse than none.

$ waymark/.tools/query_code_index.py claim --status dead
  status: dead
  evidence: measured
  entry: compaction
  text: The lost writes are caused by WAL_SYNC_BATCH not flushing before the segment swap,
        so setting WAL_SYNC_ALWAYS should fix it.
  killed_by: reproduced with WAL_SYNC_ALWAYS at 8x the write cost: identical growth and
             identical lost writes. The policy was never involved.

That entry is the point of the tool. Somebody spent days on that theory. Without it written down, the next person spends them again — and the search that finds it is a search for the symptom.

A knowledge base only helps if the way you work produces knowledge worth keeping, and does not destroy the evidence first. Docs/BEST-PRACTICE.md collects the working habits that make that true — distilled from a real project, each one written with the failure that produced it.

Quick start

git clone <this repo> && cd waymark
python3 .tools/index_code.py            # build the index
python3 .tools/query_code_index.py summary
python3 .tools/query_code_index.py selftest

Everything below runs against sample/, a small append-only key-value store included in this repository, and its knowledge base in Docs/source_index_annotations.json. The examples are real: copy and paste them.

This repository uses waymark on itself, which is also a working demonstration of the multi-root split described further down. kb.config.json names two knowledge bases:

"annotations": ["Docs/source_index_annotations.json", "Docs/kb-engine"]

The first is the sample project's, so the examples stay clean. The second is Docs/kb-engine/, waymark's knowledge about itself — the bug classes, the compatibility floors, the traps that cost someone a day. They merge into one index and one search. Try:

python3 .tools/query_code_index.py --full annotation IsADirectoryError
python3 .tools/query_code_index.py --full annotation walrus

If you are about to change the engine, read that KB first. It exists because the alternative is re-deriving the same four bugs.

Where does a knowledge base actually live? It is a folder of markdown files, one per entry, and where that folder sits — beside your code, on a branch of its own, or in a separate repository — is your choice. If that is the question you came with, skip ahead to The KB on disk and Managing it with git; the sample here uses the older single-JSON form, which the engine still reads, so it is not a good picture of the layout you want for a real project.

Finding things

python3 .tools/query_code_index.py symbol compact              # where is it, what is its signature
python3 .tools/query_code_index.py symbol wal_append --branches all
python3 .tools/query_code_index.py constant WAL_SEGMENT_SIZE
python3 .tools/query_code_index.py refs WAL_SEGMENT_SIZE       # who uses it
python3 .tools/query_code_index.py comment tombstone           # rules left at the code site
python3 .tools/query_code_index.py arch WAL                    # KB_ARCH: comments
python3 .tools/query_code_index.py api 'KV+PUT'                # your project's command dialect
python3 .tools/query_code_index.py file store.cpp

Output is compact by default and shows the first useful lines of long fields. Add --full to expand an entry before relying on its detail, or --json for machine-readable output.

What a refactor left behind

python3 .tools/query_code_index.py dangling-refs

Names this codebase used to define, no longer defines, and still references — with the enclosing function of every site, because "who still calls this" is the question you actually have.

A compiler catches that class. A scripting language does not, and the failure is patient: the real one this was written for sat in an error handler, so it threw only when a request failed — and the message that handler existed to print was the one thing that would have pointed at it. It shipped in two release images.

This is why waymark keeps looking for a name after its definition is gone. References are recorded only for names the index knows, so removing a definition would otherwise erase every call site of it — the evidence disappearing at exactly the moment it becomes interesting, and 0 references reading as "nothing uses it".

It answers from what the index remembers, so it needs a build from before the definition went and one from after; on a fresh clone with no history it has nothing to compare and says so by staying quiet.

Whether the knowledge base is still true

python3 .tools/kb_stale.py            # after index_code.py
python3 .tools/kb_stale.py --no-git   # skip the freshness pass, much faster

selftest checks a KB's internal consistency — links resolve, the vocabulary is known, a headline does not contradict its own status. Every one of those passes on a knowledge base that is perfectly consistent and completely out of date. kb_stale checks the other direction: does the entry still describe the code?

check a hit means
file-exists the entry's file: no longer resolves
symbol-exists the index does not know its name: — either it went, or the indexer never saw it
symbol-live every definition of that name is commented out
citation-range a file.cpp:NNN in the body points past the end of that file
freshness the cited file has commits after the entry's ts: — REVIEW, never a failure

Non-zero exit on a problem, so it can gate a build the way selftest does.

Two things it will not do. It does not judge prose — "this function is slow" has no definite answer. And it does not rewrite anything: a stale entry might need its citation corrected or its finding deleted, and only a person can tell which.

Read the hits before believing them. Its first run against a real 3000-file tree reported 149 problems and almost all were the tool's own fault — bare names against qualified symbols, a line number inside a frontmatter path, entries about code outside the indexed roots. 149 → 1, and the one that survived was real. Those cases are the test suite now, because they are the same mistakes any query is subject to.

The knowledge base

python3 .tools/query_code_index.py annotation compaction       # by name or keyword
python3 .tools/query_code_index.py annotation writes-disappear # by SYMPTOM -- this is the point
python3 .tools/query_code_index.py --full annotation durability-policy
python3 .tools/query_code_index.py concept store.durability
python3 .tools/query_code_index.py open                        # everything still unresolved

Notes live in Docs/source_index_annotations.json, which is the source of truth; the index is a build product and is regenerated with index_code.py. Keyword entries on the SYMPTOM, not the cause — a year later you will search for "writes disappear", not for "compaction snapshot".

annotation and concept match ONE SUBSTRING, not a set of keywords. Both are a single LIKE %term% over the entry's id, name, symbols and body. So a query written as loose keywords returns No matches for an entry that is certainly there, and that output is indistinguishable from the entry not existing. Observed 10-09-2026: a concept named "Reading branch_audit output …" was reported missing by concept "branch audit output" minutes after it was written and indexed — the name has an underscore, branch_audit.

No matches is not evidence of absence. Query a distinctive phrase you know is in the text, or list the table, before concluding anything is missing. This matters most in exactly the situation the KB exists for: you are checking whether something was already recorded, and a false "nothing here" sends you off to re-derive it, or worse, to re-implement it.

Each entry carries a short brief for triage and a long notes for the evidence. Compact output prefers the brief; --full gives you everything.

Claims: what was tested, and what died

A note that says "compaction is fine now" ages badly. A claim carries its provenance:

python3 .tools/query_code_index.py claim compaction
python3 .tools/query_code_index.py claim --status dead --dead-first
python3 .tools/query_code_index.py claim --evidence measured
  • --status — live (believed), open (unsettled), dead (disproved, with killed_by), done (was right, and has been acted on).
  • --evidence — measured, inferred, reported, mixed, unknown.
  • --dead-first — start with what has already been ruled out.

dead and done are opposites, and the difference is the whole point. dead means the claim was refuted — kept on purpose so nobody re-derives it. done means it was right and the code moved on. Recording a correct root-cause analysis as dead tells the next reader it was disproved, in the one field they consult to avoid repeating dead ends.

--dead-first is the query to run when you pick up an investigation. A hypothesis that was tested and killed, with the reason recorded, is the most reusable thing in the file and the easiest to spend a week re-deriving.

evidence takes the detail after the enum — and it should

The enum is what gets filtered and sorted on, so it stays a small closed set. But the question a reader actually has is does this still hold on the build in front of me, and only the date, the build and the rig answer that. Put both in the one field:

evidence: measured -- 2026-08-31 on the A/B bench, fw 2.0.20260828
evidence: measured 2026-08-31 on the A/B bench          # the bare form parses too

The leading word lands in evidence; the rest is kept beside it and shown with it. A value with no leading enum is unknown and is reported — see selftest below. This matters more than it sounds: before it existed, writing the useful half cost the entry its provenance, because the whole sentence failed the vocabulary check and was downgraded to "nobody said". In one real knowledge base 25 entries were in that state, several written the same week by people who had just read the rules. A field that looks like free text gets written as free text unless something says otherwise.

Cross-references

Entries link to code and to each other with see_also, and the links are validated on every build:

python3 .tools/query_code_index.py links compaction
python3 .tools/query_code_index.py broken-links

One knowledge base, several branches

selftest and broken-links ask whether a KB is consistent with itself. kb_stale.py asks whether it is still true of the code — the direction knowledge actually rots in.

python3 .tools/kb_stale.py                              # judge against the checked-out branch
python3 .tools/kb_stale.py --branches all               # ...or any branch, naming which one
python3 .tools/kb_stale.py --branches 1.18,2.0 --strict # ...and fail if it is not this branch

When branches are long-lived and not merged into one another, a referent has three possible answers, not two: here, nowhere, or correct on another branch. The third gets its own heading with the branch named, because "missing here" and "wrong" are different findings — and a checker that conflates them reports false alarms on a healthy KB, which is how a check gets switched off.

file-exists and citation-range are answered by git and widen for free. symbol-exists needs that branch's index, so where none exists the answer is unresolved, never missing.

A link may target symbol:, constant:, annotation:, concept:, route:, comment: or file:. Two conveniences worth knowing:

  • Bare names resolve to qualified ones. symbol:compact finds Store::compact, because that is how you looked it up. A name that matches more than one symbol stays ambiguous and asks you to qualify it, rather than silently picking the first.
  • @path picks one file. When the same name is defined in several files, as a page function is in every page that carries it, symbol:pick@src/p.html names one of them (so does constant:). The path is repository-relative, exactly as symbol prints it.
  • A concept and the annotation elaborating it are one subject. They share a concept_id, so annotation:compaction resolves to the entry holding the content instead of complaining that the name is ambiguous with its own concept.
  • A file: target that exists is resolved even outside the scanned roots. roots chooses what is scanned for symbols; it is not an inventory of the repository. A link to a real Docs/… file used to report missing forever.

Links into a store this index cannot read

A KB routinely points at knowledge kept somewhere else — a private per-machine note store, another team's base. A namespaced target (memory:no-force-push) resolves to external: recorded, shown, never counted as rot, because a permanent false alarm in front of the real ones is how a link report stops being read.

The trap is the bare name. no-force-push looks like an entry in this KB, finds nothing, and reports missing — sending the reader after something that was never supposed to be here. Name the stores and waymark says so instead:

"external_stores": [{"prefix": "memory", "path": "~/notes/project"}]

With that, the link still reports missing — the prefix really is absent — but carries not in this KB -- write it as memory:no-force-push, which is a one-word fix. It is deliberately a hint and not a resolver: silently marking it ok would delete the only signal that someone has to go and write the prefix. Configure none and nothing changes.

When the target lives on another branch

One KB, many branches, one index per branch. A symbol: link is validated against whatever is checked out, so a target that exists on only some branches reports missing on all the others — permanently, and correctly by its own logic. That is a false alarm standing in front of the real ones, which is how a link report stops being read.

Say so in the link:

"see_also": [{"type": "symbol", "target": "SingleWire::set_pins", "status": "branch_scoped"}]

It still resolves normally where the target is present, so it never hides a link that works here; it only changes what absence means. An ordinary link to nothing is still a defect.

And the claim is checked, not merely believed. .tools/code_index.*.sqlite is a set of sibling views of the same tree, so when the target is absent here the other branches' indexes are consulted:

what the siblings say status
one of them has it branch-scoped — confirmed, reported for review, not counted as rot
none of them has it missing — the claim is refuted, usually a typo in the name
there are no siblings yet branch-scoped — nothing to check against, so trust

A file: target is asked of git instead (git cat-file -e <branch>:<path> on each of commit_branches, or every local branch when that is not set). A file outside the scanned roots, a Docs/ document above all, is in no index, so the siblings used to refute every such link. Git needs no index, so the answer is the same on a fresh clone.

That last row matters: a fresh clone has one branch indexed, and refusing the claim there would break the feature exactly where it is needed. Unverifiable is not the same as refuted. This branch's own index is excluded from the check — a stale copy of it could otherwise confirm a claim using the very data being replaced.

Both spellings are accepted on input, branch_scoped and branch-scoped, because the validator reports the hyphenated one and copying a reported value back into the frontmatter is the obvious thing to do. Check a target by hand with:

python3 .tools/query_code_index.py symbol <name> --branches all

Two other authored statuses exist for the cases this does not cover: renamed and expired. Reach for those when the target is genuinely gone everywhere — never delete a stale link silently, since the reference is itself a record that the thing once existed.

Rebuilding

Rebuild after every KB edit, and after every git pull of a shared KB. The index is built per checkout and is not shared: a note somebody else committed does not exist for you until you rebuild. Nothing warns you — a query just answers "No matches", which reads as nobody wrote that down rather than your index is behind.

python3 .tools/index_code.py            # re-scans only what changed
python3 .tools/index_code.py --force    # everything, from scratch

A rebuild re-scans the files whose size or mtime moved, and reuses everything else. On a 678-file tree (11.6k symbols, 73k refs) that is 13.3 s down to 2.4 s for an ordinary edit. The build says which it did:

{ "scan": "incremental", "rescanned": 1, "dropped": 0, "refs_rescan": "changed" }

The rebuild is atomic: it builds into a private file and publishes it with one rename, so a concurrent reader sees the old index or the new one and never a half-built one, and an interrupted build leaves the previous index in place. Two builds racing each other each produce a complete index and the last rename wins — wasted work, never corruption, and no lock required.

One case costs more, and it is not a bug. References are found by matching tokens against the complete symbol table, so adding or renaming a symbol means files that did not change may now contain references to it. When the symbol name set moves, references are rebuilt for every file ("refs_rescan": "all"); when it does not, only the changed files are touched. Skipping that would leave references missing, and a missing reference is indistinguishable from one that was never written — the failure this whole tool exists to avoid.

Size and mtime is the same pair git's index uses, and it inherits the same caveat: a checkout can restore an old timestamp. --force is the answer to that, rather than making every build slow.

The KB on disk

A knowledge base is a folder of markdown files. That is the entire storage format — no database, no server, no lock file. You can read it with cat, edit it in any editor, and diff it in a code review. Everything else on this page is built from that one fact.

If you are starting from nothing:

mkdir -p kb/features kb/concepts            # 1. make the folder
$EDITOR kb.config.json                      # 2. "annotations": "kb"
python3 .tools/index_code.py                # 3. build the index

1. One entry is one file

The filename is the entry's name. A small frontmatter block, then ordinary prose:

--- kb/routes/kv-put.md ---
---
name: KV+PUT
file: sample/cli/kvctl.py
keywords:
  - route
  - write-path
  - put
---

## notes

Client -> server write path: KvClient.put -> KV+PUT -> Store::put -> wal_append.
The acknowledgement is sent only after the WAL accepts the record.

Long explanations go in the body, below the frontmatter — there they diff like prose instead of arriving as one line full of \n escapes.

The frontmatter dialect is deliberately tiny and is not YAML. Three forms, nothing else:

key: plain text to the end of the line
key: [{"json": "when you need structure"}]
key:
  - one list item
  - another

Anything the reader cannot parse is an error naming the file and line — never a silently dropped field. That is what makes the format safe to hand-edit.

2. The files sit in named folders

Each folder is a collection, and the names are a fixed vocabulary — not cosmetic. Each one lands in a different table and is reached by a different query:

kb/                                      ← whatever `annotations` in kb.config.json points at
├── kb.json                              ← lists the collections below
├── features/                            ← how something behaves; ported/verified status
│   ├── io32-slot3-guarded.md
│   └── wifi-apply-freeze.md
├── concepts/                            ← incidents, open bugs, root causes, invariants
│   └── aes-parallel-uart-crash.md
├── routes/                              ← endpoints and channels
├── symbols/                             ← notes bound to one symbol (add_note.py writes here)
└── symbol_comments/                     ← hard-won rules left at an exact code site
folder you read it back with
features/ query_code_index.py annotation <keyword>
concepts/ query_code_index.py concept <id>
routes/ query_code_index.py route <name>
symbols/ query_code_index.py notes <Symbol>
symbol_comments/ query_code_index.py comment <term>

Two more things may appear in that folder later: a per-version overlay JSON, which must stay inside the KB directory because it is found relative to it, and nothing else. If kb.json is missing, every non-dotted subdirectory is treated as a collection — which is why .git sitting beside the entry folders does not break anything.

3. What is NOT part of the KB

.tools/code_index.*.sqlite is generated — one per branch, rebuilt by index_code.py from the KB plus a scan of your source. Gitignore it. Never edit it, and never treat a query result as the record. Delete it and nothing is lost.

A rebuild against a missing KB produces an empty index without erroring, which is the failure that looks most like success. Check the entry count after a rebuild.

Managing it with git

The engine never looks at where the folder lives — that is your call. Three arrangements work:

arrangement good for cost
a plain folder, gitignored notes that must never leave the machine no history, no sharing
an orphan branch, as a worktree shipping the KB with the repo everyone already has none worth naming — start here
its own repository a KB spanning several projects, or a different audience one more thing to clone

Why not just commit it on your main branch

Because the KB is edited from every branch. A file that lives on main and is edited while you are on a feature branch either follows you (and shows up in your code diffs) or does not (and your notes vanish when you switch). An orphan branch sidesteps both: it is a branch in the same repository that shares no history with any other, so it never merges into your code and your code never merges into it.

A worktree is git checking out a second branch into a second folder at the same time. Together they give you a KB folder that stays put while you git checkout in the code tree.

git checkout --orphan kb          # a branch with no shared history
git rm -rf .                      # nothing but the KB lives here
mkdir features concepts
echo '{"schema": 2, "scope": "shared", "collections": ["concepts", "features"]}' > kb.json
git add kb.json && git commit -m "kb: initial import"
git checkout main                 # back to your code

git worktree add ~/kb/myproject kb    # the KB now lives here, permanently
$ git worktree list
/home/…/myproject     4f2a1c9 [main]      ← your code
/home/…/kb/myproject  a908d5c [kb]        ← your KB

$ git -C ~/kb/myproject rev-parse --git-common-dir
/home/…/myproject/.git                    ← the SAME repository, not a clone

Point kb.config.json at it and you are done:

"annotations": "~/kb/myproject"

Day to day

Record, rebuild, commit — KB edits in the KB worktree, code in the code tree. Two branches, never staged together:

python3 .tools/add_note.py <Symbol> "<finding>" --keywords "a, b"
python3 .tools/index_code.py
git -C ~/kb/myproject add features/<name>.md
git -C ~/kb/myproject commit -m "kb: ..."

What that buys you: git -C ~/kb/myproject log features/<name>.md answers when did this become true, and two people editing two different entries never touch the same file.

Four rules, each of which exists because breaking it cost someone a day:

  • Never commit the generated index.
  • Never git add -A in the KB — it sweeps in the index and whatever else is lying around. Name the paths, or use git add -u.
  • Never merge the KB branch with a code branch, in either direction.
  • Keep the version overlay inside the KB folder — it is found relative to the KB, so moving the KB without it orphans it silently.

Public and local halves

annotations also takes a list, so a shared KB can sit beside one that is in no repository at all:

"annotations": ["~/kb/myproject", "~/kb/myproject.local"]

Both are merged into one index and one search — only git tells them apart. Site-specific material (bench addresses, personal network details, customer particulars) goes in the local folder. Prefer that over marking an entry private: a file merely marked local still sits in the tracked worktree, and one git add publishes it. A folder in no repository cannot be pushed by accident.

Declare which is which in each root's kb.json ("scope": "shared" / "local"), and add_note.py will tell you where a finding went:

$ python3 .tools/add_note.py gadget "the FIFO drains before the ack"
noted gadget -> /home/you/kb/project/symbols/gadget.md   [shared]
  this KB is shared -- use --local for anything site-specific

$ python3 .tools/add_note.py gadget "on THIS bench, io3 is jumpered to io7" --local
noted gadget -> /home/you/kb/project.local/symbols/gadget.md   [local]

Naming the target matters because the default is the shared root, so every note is a publish-or-not decision — one that used to be taken silently. --local refuses rather than falling back when no second root is configured: a fallback there writes exactly the note that should not be published into exactly the place it should not go.

When an entry has both, split it rather than hiding it whole — the general finding goes public where other people can use it, the values stay local, and the two cross-link. Most bench notes are 90 % general; moving them wholesale would gut the shared KB of exactly the root-cause work it exists to carry.

Docs/STORAGE.md goes deeper, with the reasoning behind each rule.

Browsing it

python3 .tools/serve_code_index.py        # http://127.0.0.1:8765/

Every search the CLI has, the browser has — because it does not implement any of them. It asks /api/commands what exists and forwards each query to query_code_index.py --json. Add a subcommand and it appears in the picker; there is one implementation of each search rather than two that drift.

Both kinds of link are indexed: the declared see_also field, and [[name]] written inline in prose — which is how most of them are actually written. On the tree this engine came from there were 232 inline references against 29 see_also entries, none of them indexed, so nothing validated them and the graph showed an unconnected KB that is in fact densely cross-referenced. A declared link that does not resolve fails selftest; an inline one is reported but does not, because the convention allows a forward reference to something not written yet.

Press graph to draw the KB itself: entries as nodes, see_also and relations as edges, with broken links and violated relations in red. Drag to pin a node, click to open the entry. The layout is a few dozen lines in the page — no library, so it works offline like everything else here.

Relations: claims the build can test

see_also says two entries are related. A relation says something about the code, and every build goes and checks it:

{
  "name": "wal_open",
  "notes": "Opening the log is Store::open's job. A write path that reaches it opens a SECOND handle...",
  "relations": { "must_not_call_from": ["Store::put"] }
}
python3 .tools/query_code_index.py relations

must_not_call_from is tested against the call graph already in refs, transitively — a violation is reported with the chain that causes it (Store::put -> wal_append), and selftest fails on it, because an assertion the code has stopped honouring is a defect and not a note.

Know what this can and cannot tell you. A violation is a finding; silence is not evidence. refs.in_symbol is the nearest preceding definition, headers get no attribution, and a source-level graph cannot see calls the compiler invents — the bug that motivated this feature reached flash from an ISR through a GCC $constprop clone and had to be caught by disassembly. So the check catches paths that should not exist; it never certifies that none does.

For the same reason the engine is explicit about what it did not check. A target that resolves to no known symbol is unchecked, and a relation kind this engine does not implement is unknown-relation — both listed by selftest without failing it. The one thing worse than an unchecked invariant is an unchecked invariant that looks enforced.

Everything else — "writes EEPROM", "restarts WiFi" — belongs in claims, with the evidence that backs it. Prose about side effects cannot be tested, so it should carry provenance rather than sit next to a verified relation looking equally solid.

Checking the KB itself

python3 .tools/query_code_index.py selftest     # exits 1 on any problem

Each check corresponds to a way a knowledge base goes wrong quietly: an index that built from a missing annotation file and is silently EMPTY, a see_also pointing at something that no longer exists, a claim with no provenance so an inference reads like a measurement. It exits non-zero, so it belongs in a pre-commit hook or CI.

Including a status or evidence value waymark does not know. Those are silently downgraded — an unknown status becomes n/a, which drops the entry off the open list, and an unknown evidence becomes unknown, so a carefully sourced entry reads as having no provenance. The loader has always detected them, but it prints at the end of a build that actually re-reads the KB, and an unchanged KB short-circuits — so the warning only ever appeared under --force, which nobody runs without a reason. 65 had accumulated unseen in one real KB. They are recorded at load time and reported here instead, naming an offending value rather than only counting.

When the headline outlives the entry

The most-read part of an entry is its one-line brief, and that is the part most likely to go stale: the body gets corrected when the facts change, the summary above it does not. Waymark compares the verdict an entry OPENS with against its own status and says so in the reader:

$ query_code_index.py annotation wifi-apply-freeze
wifi-apply-freeze feature
  !! status=open, but the brief opens by calling it RESOLVED -- read the notes, not the headline

It is deliberately not a verdict about which side is right — a mixed state is legitimate, and a fix that has landed while its bench re-test is still outstanding is honestly both. It is a warning placed where somebody is about to act on the headline alone. selftest lists the same entries as REVIEW and does not fail for them, because a check that is permanently red is a check nobody reads. Only the opening of the brief is examined; a brief that narrates a history it has moved past ("this was believed FIXED in build 9 and is not") is left alone.

Skills: letting an agent run a procedure instead of re-deriving it

A procedure that lives only as prose gets re-derived. Opt an entry in, and the build projects it into a skill an agent can load:

skill: cross-branch-port
skill_when: "port this to the other branches, sync a fix across the release lines"

…plus a ## procedure section in the body, which is the only section copied. index_code.py regenerates .claude/skills/ on every run, so there is no restore step and no second copy to drift; gen_skills.py --check reports without writing and exits non-zero, so it can gate a build. skills_dir in kb.config.json moves the output.

A skill is a projection, never a source. It is also the narrowest place knowledge can sit: it means nothing to a human, nothing to another tool, and — if your repository publishes through a filtered mirror — it may not travel at all, while .tools/ and the KB do. So the finding stays in the entry, the doing stays in a script, and the skill is a thin adapter over both.

An entry with no ## procedure section is refused, not flattened. Most of a good entry is evidence, measurements and dead hypotheses — its whole value, and exactly what must not be loaded into a skill.

Tools beyond the query

Everything above is query_code_index.py. These are separate, and each answers one question:

gen_skills.py project opted-in procedures into agent skills (above)
kb_stale.py is the KB still true of the code — see Whether the knowledge base is still true
verify_port.py is this commit's substance on that branch? Scores the tokens the commit introduced, and marks +OLD where the branch still carries a line the commit eliminated -- the one signal an argument-only fix leaves. Reads branches as refs (--ref). Triage, not a verdict
xstatus.py a cross-branch register: why a fix is absent from a branch — never ported, deliberately rejected, not applicable — which git has nowhere to record. See Docs/CROSS-BRANCH-REGISTER.md
shared_paths.py are the directories you have DECLARED shared -- a toolset, an SDK, a test suite -- byte-identical on every branch? Compares git tree hashes and groups the branches that agree, so the report is "three agree, one differs" and names the files. BEST-PRACTICE section 8 states the rule; this is the instrument for it
css_factor.py which repeated CSS declarations can be folded into a selector list, and whether that is safe. Refuses any fold the cascade would change, and reports the refusals — usually the more interesting output

Handing work to the next context

Waymark keeps durable facts queryable; a handover records current working state. Use both when a session stops mid-investigation: put facts and dead hypotheses into the KB first, rebuild and run selftest, then write a short handover with repo state, hardware/runtime state, what is proven, and the next concrete action. See HANDOVER.md.

What it needs

Python 3.7 or newer, standard library only. No third-party packages, and nothing to install.

3.7 is not a guess: the engine is run on a work host with Python 3.7.8 and a SQLite built without the JSON1 extension, and both of those have already broken it once. Where SQLite lacks JSON1 the notes query registers a small pure-Python json_extract and carries on. Where the interpreter is 3.7, syntax is the thing that bites — a := does not degrade one query, it stops the module parsing and takes every query with it, selftest included. CI parses the tree under a real 3.7 for that reason, and the test suite refuses 3.8+ syntax on whatever machine you are working on, which is where a syntax error is cheap to fix.

Using it on your own project

Start here: SETUP.md — first-time setup, migrating an existing knowledge base to one file per entry, the layout that makes it shareable, and the data-safety cautions that go with moving a KB.

Copy .tools/ into your repository and write a kb.config.json at the root:

{
  "roots": ["src", "tests"],
  "annotations": "Docs/source_index_annotations.json",
  "version_file": "src/version.h",
  "api_regex": "(KV\\+[A-Za-z0-9_?=:+,.-]*|\\?[A-Za-z0-9_*][A-Za-z0-9_=&*.,+-]*)"
}

Every field is optional. With no config at all the engine indexes the repository it sits in, which is enough to try it. api_regex teaches it your project's command dialect — omit it and no api markers are indexed rather than a dialect being invented for you. version_file is read for a BUILD_VER define, which lets you keep version-specific overlays.

Languages recognised out of the box: C, C++, Python, JavaScript, HTML, shell. Parsing is regex over a lexical pass that classifies code, comments, strings and char literals — good enough to navigate by, and deliberately not a compiler.

Indexing a language that is not in that list

Four more optional fields decide what gets scanned and which grammar parses it:

{
  "roots": ["VEO"],
  "source_exts": [".cs"],
  "c_like_exts": [".cs"],
  "skip_dirs": [".git", ".tools", "bin", "obj", "packages"],
  "max_file_bytes": 2000000
}
  • source_exts — extensions to index. Replaces the built-in list. "cs" and ".CS" are accepted and mean .cs.
  • c_like_exts / js_like_exts — which extensions are parsed with the C-like or JS-like grammar, and lexed for their comment style. Setting source_exts alone indexes the files and their comments but finds no symbols, so a brace language wants both.
  • c_like_exts / js_like_exts / xml_like_exts replace their built-in sets too. Naming ".cs" alone means C and C++ are no longer parsed — list every extension the project wants, not just the new one.
  • xml_like_exts — markup parsed with the XML-like grammar: .xml, .xaml, .xsd, .xsl, .resx by default. It indexes the names a thing can be referred to by — x:Class, x:Name, x:Key, and id/name in plain XML — not element tags, which would bury the file in <Grid>. None of these are in the default source_exts, so nothing changes until a project asks for them.
  • skip_dirs — directories never descended into. Replaces the built-in set rather than adding to it, because a project that names them is describing its own tree.
  • max_file_bytes — size cap per file, default 2 MB.
  • index_ignored — true indexes files git ignores. By default they are skipped: a render, a scratch copy or a nested checkout never reaches a clone, and indexing one duplicates every definition in it. Outside a git repository nothing is skipped.

This buys navigation, not comprehension: a C# tree indexed as C-like yields methods and classes but misses properties and expression-bodied members. On a real 421-file C# project it found 5714 symbols where a purpose-built C# parser found 5959 — worth having, and not the same as language support.

Why these exist. Without them, pointing the engine at a language outside the built-in list produced an index with one file and no symbols and exited 0 — an empty index is indistinguishable from a repository with no code, so every later query answered "no matches", which reads as an empty topic rather than a broken setup.

Keeping your notes private

The engine is separate from what you write with it, and where your KB lives is your call — gitignored beside the code, an orphan branch in the same repository, or its own repo. All three work; Docs/STORAGE.md compares them. The KB in this repository is committed because it documents the sample.

For anything you intend to share, split the KB rather than redacting it. annotations takes a list, so a pushed root can sit beside a local one that is in no repository at all:

"annotations": ["~/kb/myproject", "~/kb/myproject.local"]

Site-specific material — bench addresses, home network details, customer particulars — goes in the local root. Prefer that over marking an entry private: a file merely marked local still sits in the tracked worktree, and one git add publishes it. A directory that belongs to no repository cannot be pushed by accident.

Backups are automatic: every run keeps a gzipped copy of each annotation file it read under .tools/kb-backups/, one per distinct content, sixty retained. The file is hand-written over months and is usually not in git, so it is the one thing here worth protecting from a bad edit.

Sample project

sample/ is a small append-only key-value store — a C write-ahead log, a C++ index on top, a Python client and a status page. It is not a toy for its own sake: its knowledge base carries a real-shaped investigation (writes appearing to vanish, two dead hypotheses, the actual root cause) so the queries above have something honest to return.

What has changed

CHANGELOG.md, newest first. There are no version tags yet, so entries are dated; each says what a user gets or stops getting, and the commit behind it carries the reasoning and the measurements.

Licence

MIT. See LICENSE.

About

A knowledge base that lives next to your source and is queried from the command line.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages