Skip to content

feat(plugin): a shared skills library across all six hosts - #26

Merged
ypflll merged 15 commits into
mainfrom
feat/shared-skills
Sep 17, 2026
Merged

ypflll merged 15 commits into
mainfrom
feat/shared-skills

Conversation

@Tian-yi-Sun

Copy link
Copy Markdown
Collaborator

Implements skillsearch-shared-skills-spec.md — all six designs — on top of f2a7403.

The problem

Each of the six hosts scans its own skills directory and only its own, so a skill you have in one agent is invisible to the other four. And a skill retrieved from a catalogue was extracted for one turn and thrown away — re-downloaded next turn, never seen by another agent, because the cache it landed in is deliberately outside every scan.

Raven has it worst: its default skills directory is a relative path, so two projects belonging to one user do not share even with each other.

What this does

One shared root, ~/.evermind-skillsearch/, the same expression on all three platforms. On Windows that is not the platform convention and shared.py says why consistency wins: it is the one path all five hosts must compute identically, and a platform branch is somewhere for a service-launched agent to quietly resolve elsewhere.

Hosts self-register. Hardcoding the five default directories is wrong three ways at once — versions change the default, users move it, and the Windows location is not knowable from here. Each host writes the absolute path it actually resolved and reads the table back, so what it sees is exactly "the other agents that also have this plugin".

Retrieved skills are kept, in the shared library, one directory per identity rather than per version. Each carries a .skillsearch-origin.json; removals append to uninstalled.log.

They still count once. Installing into a scanned directory recreates the double-counting the cache exclusion existed to avoid, and neither existing defence catches it: fusion keys on qualifiedId, and local/x and hub/x differ; the exact-body dedup compares a digest a bumped version defeats. So the marker carries an identity and fusion collapses on that.

Changes take effect next turn, not next restart. SkillSearch.invalidate existed for this and no adapter had ever called it. WorkBuddy's per-turn fingerprint is generalised to both engines, scoped to the shared directory alone — the other roots keep their host's existing behaviour, which for four of five still means restarting.

Two switches, deliberately opposite, because users conflate them: enabled in the registry is "the others cannot see me"; shareSkills / share_skills in a host's config is "I cannot see the others".

Also closes the gap the spec notes in passing: Raven and Hermes had only a singular skills_dir and no environment override, while the engine has supported several roots all along.

Verification

All nine acceptance items observed on real hosts, against the real EverMind SkillHub. Cases are tests/host-e2e/cases.md S1–S9; results are tests/host-e2e/reports/shared-skills.md and workbuddy-shared-skills.md.

# Acceptance How
S1 retrieved skill lands in the shared library live catalogue
S2 next turn finds it once, no restart live catalogue
S3 agent B retrieves what agent A installed Raven installs → OpenClaw 2.0 retrieves
S4 a skill in A's directory is retrievable in B OpenClaw's reply carried Wombat-Ledger-7, a token present only in Raven's SKILL.md
S5 hand-dropped skill found next turn two hosts, plus each host separately
S6 enabled: false holds across A restarting two hosts
S7 uninstall → gone, recorded, not retrievable two hosts + live catalogue
S8 a failed update leaves the old version working real install, real update, dead endpoint: 2546 bytes unchanged, no staging directory left
S9 corrupt registry costs sharing, not retrieval two hosts

Each of the five headless hosts is also asked separately whether it joins at all, because the registration is wired at six different call sites: openclaw 1.x, openclaw 2.0, DeepSeek Harness, Raven, Hermes — all register and all retrieve.

Suites: python 177, parity 54, hermes 26, raven 20 (+6 skipped without a checkout), openclaw 47, openclaw2 47, workbuddy 49. Gates and the release-version validator pass.

Bugs this found, and how

Six, five of them in code every package suite reported green.

Found by Defect
the live catalogue list_installed looked only at the top level, so a bundle that wraps the skill in a directory was invisible — dedup worked while listInstalled reported an empty library holding two skills and uninstall refused to remove anything
the same, one port later the identical bug, unfixed, in the TypeScript port
a differential between the ports the two computed different install directories for a non-ASCII slug, so one skill occupied two of them; every all-CJK slug collapsed to the same name and overwrote itself
the same differential a whitespace-only SKILLSEARCH_SKILLS_DIRS read as an override on one side and unset on the other, silently emptying the host's skills directory and with it the shared library
building the DSH host installRootFor(cfg.shareSkills) did not compile — the field is optional on the exported Config and the zod default only applies to config the harness parsed. No suite here compiles that file against that interface
re-reading the spec verbatim acceptance 7 asks for a record on uninstall. There was none

Three of the six are one shape: a fix applied to one port and not the other. Every package suite compares a port with itself, which is why tests/parity/ now exists — nine comparisons on identical inputs, two of them byte-level (Python writes a registry, TypeScript rewrites it, the file must come out identical).

Worth a reviewer's attention

  • shareSkills defaults to true. The spec does not say. It means every host with the plugin starts scanning the others' directories on upgrade. One constant if you want it conservative.
  • installRoot is a new parameter rather than a change to cache_dir semantics. Unset reproduces 0.3.0 exactly, which makes this reversible and comparable.
  • on-demand tool triggering is unreliable on WorkBuddy's default model. Measured: it answered from its own knowledge, and once named a skill that does not exist without calling the tool at all. The tool works — an explicit instruction triggers it. Whether the answer is a reworded description or recommending auto there is a product decision and is left as one.

Not done

  • The shared library is unverified on WorkBuddy's cross-host combinations. Its single-host behaviour is verified (report attached); nothing has confirmed a WorkBuddy↔other-agent pair, because that run had only WorkBuddy on the machine.
  • S5 and S9 on all five hosts. The spec words them as 五家; they ran on the hosts named in each table, and the per-host table covers the part that is per-host code.
  • Concurrent registration from several hosts at once. The write is atomic and tested as such, but no test runs two hosts registering simultaneously.

Untouched, per the spec's own out-of-scope list: PathGuard placeholder divergence, sandboxing, cross-machine sync, native host skill loading, and anything in the router, fusion or gate.

🤖 Generated with Claude Code

Tian-yi-Sun and others added 13 commits September 9, 2026 02:28
Spec sections 1, 2 and the multi-directory gap it lists alongside them. This
is the foundation the install/dedup/invalidation work sits on; it lands on its
own because it is independently useful and independently reviewable.

Today each of the five hosts scans its own directory and only its own, so a
skill a user has in one agent is invisible to the other four. Raven has it
worst: its default is a *relative* path, so two projects belonging to one user
do not share even with each other.

`shared.py` / `shared.ts` add one root — `~/.evermind-skillsearch`, the same
expression on all three platforms — and a registry inside it. On Windows that
is not the platform convention, and the module says why consistency wins here:
this is the one path all five must compute identically, and a platform branch
is somewhere for a service-launched agent to quietly resolve elsewhere.

Hosts self-register rather than us hardcoding their five directories, because
hardcoding is wrong three ways at once: versions change the default, users
move it, and the Windows location is not knowable from here. Each writes the
absolute path it actually resolved and reads the table back, so what it sees
is exactly "the directories of the other hosts that also have this plugin".

Three disciplines the file's editability depends on, all tested:

- registering is idempotent — the steady state is one small read and no write,
  which matters because WorkBuddy's hook is a fresh process every turn, on the
  turn's hot path, inside an 8s budget;
- re-registering never overwrites the user's `enabled`, so switching a host
  off survives that host restarting;
- everything fails open. A missing, truncated or hand-corrupted registry costs
  this machine the sharing feature and nothing else — never a turn.

`enabled` in the registry and `shareSkills`/`share_skills` in a host's own
config are deliberately two switches, because users conflate them: the first
is "do the others read me", the second is "do I read the others".

Also closes the gap the spec notes in passing — Raven and Hermes had only a
singular `skills_dir` and no environment override, while the engine has
supported several roots all along. Both now take `skills_dirs` and
`SKILLSEARCH_SKILLS_DIRS`, split the same way the three TypeScript packages
already split it.

The two ports write this file and read each other's writes on one machine, so
`fixtures-registry.json` pins eleven parse cases and five unparseable
documents that both suites assert against. A disagreement there would not be
cosmetic; it would split a user's agents apart with nothing logged.

Two things found by building it:

Registration is a side effect of constructing an engine, which means the test
suites started writing to the developer's real home directory — and one
OpenClaw ranking assertion changed because retrieval had picked up whatever
was installed there. Every suite now redirects `SKILLSEARCH_HOME` to a scratch
directory, autouse on the Python side so a new test cannot forget.

The project-inventory gate walks the working tree rather than the index, so a
gitignored `.ruff_cache` still failed it. Dot-directories are now skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…estart

Spec section 6. Without it the feature's headline does not work: a user
installs a skill in agent A and agent B cannot see it until B restarts.

The cause was already documented and already unused. `LocalPool.__init__`
does one eager `rebuild_index()` and never rebuilds; `SkillSearch.invalidate`
exists for exactly this and `local_pool.py` says so in a comment — and no
adapter has ever called it. WorkBuddy was the accidental exception: its hook
is a fresh process per turn, so it had to restore from a disk cache, and it
fingerprints by path and mtime to decide whether that cache is still good.

`watch.py` / `watch.ts` generalise that fingerprint, and the engines check it
before each retrieval. Three things about the scope, all deliberate:

**Only the shared directory.** A host now scans five or more roots rather than
one, and walking all of them every turn would charge deployments that are not
using the shared library. The shared directory is ours and its size is
something we control; it is also the only root that changes behind the host's
back. Everyone else's directories keep whatever their host already did, which
for four of them still means a restart. That is a real limitation, and it is
the spec's call rather than an oversight.

**A fingerprint, not a revision file.** A counter the plugin bumps when *it*
installs something is one `stat`, but it cannot see a user dragging a
directory in by hand — which is one of the acceptance cases. A revision file
could only ever be a fast path in front of this, never a replacement.

**The walk is not pure overhead.** Measured on WorkBuddy: 34ms to fingerprint
46 skills against 53ms to read and parse them. The walk is flat in corpus size
and the parse it avoids is not, so it pays off harder as the library grows.

Two details worth the reader's time. The walk sorts each level, because a
filesystem may return `readdir` entries in any order and an unsorted walk
would invalidate the cache at random. And the TypeScript side reads
nanoseconds via `statSync(path, {bigint: true})` rather than `mtimeMs`:
millisecond resolution misses an edit made inside the same millisecond as the
previous walk, and that failure is silent and reads as "my change did
nothing".

Both suites assert the headline end to end — drop a skill into a watched
directory mid-run and it is retrieved on the next call, exactly once, with
nobody having called `invalidate()`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…anage them

Spec sections 3, 4 and 5. They land together because they only work together:
installing into a scanned directory is what creates the double counting that
identity dedup exists to prevent, and the marker carrying that identity is
also the ledger management reads.

**Where they land.** A skill retrieved from a catalogue was extracted into
`.skillsearch-cache/`, used for one turn and thrown away — re-downloaded next
turn, and never visible to another agent. That cache is deliberately outside
every scan, and for a good reason: underneath a scanned directory each
downloaded bundle would be picked up as a *local* skill and compete with
itself in the same ranking. Installs now go to the shared directory instead,
one directory per identity rather than per version, so an update replaces
rather than accumulating a second ranked copy.

Not into a host's own directory, which was the other option and is worse:
`~/.openclaw/skills` is loaded by OpenClaw's native skill system too, so
putting a third-party download there decides on the user's behalf that it
applies to the host as well; ownership becomes unguessable; and the same skill
ends up copied per host.

**Why it still counts once.** `provenance.py` / `provenance.ts` write a marker
inside each installed skill — source, slug, version, body digest, timestamp —
and the scanner carries the identity on the hit. Fusion collapses on it.
Neither existing defence could: fusion's key is `qualifiedId`, and
`local/pdf-tables` and `hub/pdf-tables` are different ids, while the exact-body
dedup compares a digest that a bumped version or a changed line ending defeats.
Identity is what the skill *is*, so the collapse holds whatever the bytes are.

**Managing them.** Installing puts files on someone's disk, so `listInstalled`
/ `uninstall` and a human-readable marker per skill are not optional. The
ledger is the directory rather than an index: a skill the user deleted by hand
is simply gone, not a stale row nobody can explain.

Updates are "build the new one, switch, then delete the old". `rename` alone
cannot do it — on POSIX it fails onto a non-empty directory — so the old copy
is moved aside and moved *back* if the switch does not complete. A failure at
any point leaves the previous version in place and working, which is
acceptance 8.

Two things found by building it:

A same-version reinstall is skipped. Without that, any turn whose retrieval
happens to hit an already-installed skill rewrites it on disk, which is both
wasteful and enough to trip the directory watch every single turn.

Creating the shared directory eagerly made every deployment "enabled", because
an empty shared directory is still a source — so a host configured with no
skills directory at all started scanning and registering. It does not now: a
deployment that configured no skills directory is stating it has no local
skills, and joining the shared library would contradict that. Three existing
tests caught this, which is what they were for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Everything else about the shared library is unit-tested in-process, which
proves the mechanism and cannot prove the claim. The claim is that two
different agents on one machine see each other's skills, and there is no way
to observe that without running two of them — so the acceptance table in the
previous commits was asserting a mechanism and reading as an observation.

`e2e_shared.py` runs one Python host and one TypeScript host against one
`SKILLSEARCH_HOME` and a real model. Two hosts rather than five because two is
what the two *ports* are: a disagreement between them is the failure a
single-host run cannot see, and a third host of either kind re-runs the same
code.

Observed, at this commit, against Qwen3.6-27B:

- OpenClaw 2.0 retrieved a skill that exists only in **Raven's** directory —
  its reply carries `Wombat-Ledger-7`, a token present in nothing but that
  file, which Raven registered and OpenClaw read from the registry;
- a skill dragged into the shared directory by hand was found on the next
  turn, with nothing restarted;
- setting `enabled: false` on Raven's line made OpenClaw answer that it has no
  such internal procedure — and Raven re-registering afterwards left the flag
  off, which is the rule that makes the file editable;
- a registry corrupted to `{ this is not json` left retrieval working.

That is acceptance 3, 4, 5, 6 and 9. Acceptance 1, 2, 7 and 8 are about
installing and need a live catalogue, so they stay unit-tested, and the
script's own docstring says which is which rather than implying otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found by running the install path against the real EverMind SkillHub instead
of a fixture, which is the whole reason to do that: a hand-built fixture
agrees with itself, and this shape only exists because a real catalogue sends
it.

Hub bundles wrap the skill in one directory, so `SKILL.md` — and the marker
beside it, which is where the scanner reads provenance from — sits one level
below the directory the install created. `list_installed` looked only at the
top level and found nothing.

The effect was a split brain rather than an outright failure, which is worse:
deduplication worked, because the scanner reads the marker from the `SKILL.md`'s
own directory, while every management call was blind. `listInstalled` reported
an empty library with two skills in it, and `uninstall` refused to remove
anything.

Resolution goes through the wrapper the same way the bundle root does, one
level and no further — a marker deeper than that was not written by this code,
and treating arbitrary depth as an install would let a skill that ships another
skill be uninstalled out from under its owner. `find_installed` returns the
directory the install *created*, so removing it does not leave an empty husk.

`e2e_install.py` is the run that found it: retrieve from the live catalogue,
then assert on acceptance 1, 2 and 7. Acceptance 8 stays unit-tested, and the
script says so — it is a property of the swap primitive, not of the catalogue,
and forcing a real download to fail halfway would be theatre.

Three defects in the harness itself, all of which had made a broken thing look
fine or a working thing look broken:

- it counted `### Skill: <slug>`, but the block prints the *frontmatter name*,
  and the two differ for every catalogue skill — a working dedup reported as
  0 occurrences;
- each turn ran under its own `asyncio.run`, so the engine's HTTP client died
  with `Event loop is closed` after the first, silently turning the dedup check
  into a local-only check that could not fail. Real hosts have one loop;
- it looked for the marker at the outer directory, which is exactly the bug
  above, and so would have passed against the pre-fix code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs, both of the same kind, and the second only found because the first
made me go looking: a fix applied to one port and not the other.

**The wrapper bug, in TypeScript.** `listInstalled` and `findInstalled` looked
only at the top level, so a catalogue bundle that wraps the skill in a
directory was invisible to them — the identical defect fixed in the Python
port one commit ago, still sitting here. Measured before fixing:
`listInstalled` returned `[]` and `uninstall` returned `false` against a
correctly installed skill. Both suites were green because both fixtures were
flat, which is the shape no real catalogue sends.

**Where a skill lands, disagreeing across the ports.** Worse, because it is
silent and permanent. `slug_dir` sanitises a catalogue slug into a directory
name, and the two ports did not compute the same one:

    中文技能        Python  hub__中文技能        TypeScript  hub______
    café-export     Python  hub__café-export     TypeScript  hub__caf_-export

Python's `str.isalnum` is Unicode-aware; the TypeScript character class is
ASCII. So one skill installed by Raven and by OpenClaw occupies two
directories in the shared library — two ranked copies of one skill, which is
precisely what identity dedup exists to prevent. And every all-CJK slug
collapsed to the same `hub______`, so distinct skills overwrote each other.

Both ports are now ASCII-only and identical, and both append the identity's
digest. Sanitising alone is lossy: two skills collide whenever their slugs
differ only in dropped characters, and one silently overwrites the other. With
the digest the map from identity to directory is injective, which is what "one
directory per identity" has to mean.

One more character of drift after that: a JS regex walks UTF-16 code units, so
an emoji is two of them and became two underscores against Python's one. Fixed
by iterating code points.

`fixtures-slugdir.json` pins thirteen cases — CJK, accented, astral, traversal,
over-length, empty — and both suites assert against it. The class of bug here
is that each suite only ever compared a port with itself.

Re-verified on real hosts after the change: the install run against the live
EverMind SkillHub still passes acceptance 1, 2 and 7, and the two-host run
still passes 3, 4, 5, 6 and 9.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Third divergence between the ports, found by finally doing the
function-by-function differential I had said was missing rather than
asserting the two agreed.

`SKILLSEARCH_SKILLS_DIRS="   "` read as an override on the TypeScript side and
as unset on the Python side:

    python      ['/cfg']
    typescript  []

Emptying the list takes the host's own skills directory with it, and since a
host with no directory of its own does not join the shared library, it
silently takes cross-agent sharing too. A variable holding spaces is
indistinguishable from an absent one to whoever set it, so Python's reading is
the right one; `pick()` in all three TypeScript packages now trims before
deciding. The empty-string case was already handled — this is the same
judgement, one character wider.

Pre-existing on that side, but this branch is what made it costly.

`tests/parity/` is the tool that found it, checked in with its README. Nine
comparisons on identical inputs, and the two that matter most are byte-level:
Python writes a registry and a marker, TypeScript reads both and rewrites the
registry, and the file must come out byte for byte identical. A shape-only
comparison passes while the two quietly write different JSON.

It is a hand-run tool, not a CI job, and the README says why: it needs both
toolchains in one place and the pipeline runs the two languages in separate
images. The durable half is the two shared fixtures, which do run in CI in
both languages.

Three bugs have now come out of this one blind spot — a fix applied to one
port, two ports computing different install directories, and this. The README
names all three, because the lesson is about the method: every suite here
compares a port with itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Re-read the spec's nine acceptance items against what had actually been run,
rather than against my summary of it, and four did not hold up.

**Acceptance 7 was missing a feature, not a test.** "卸载 → 目录消失、记录留痕、
检索不到" — uninstall deleted the directory and left nothing behind. A user
asking what this plugin ever put on their machine had no way to find out, and
a skill that vanished was indistinguishable from one never installed. Removals
now append to `<shared root>/uninstalled.log`: JSON Lines, never rewritten, a
sibling of `skills/` rather than a file inside it, because everything under
`skills/` is walked by the scanner every turn.

Its last clause is now asserted by retrieving rather than by looking at the
disk. "The ledger is empty" and "the engine no longer answers with it" are
different facts, and only the second is what the spec asks for.

**Acceptance 3 was never actually run.** "打开 agent B → 也能检索到它", where
"它" is the skill A installed in item 1. I had verified B seeing a skill in A's
own directory (item 4) and B seeing a hand-dropped one (item 5), and reported
those as item 3. The chain now runs end to end: Raven retrieves against the
live EverMind SkillHub, which installs into the shared directory; OpenClaw — a
different host, the other language port, its own process — then retrieves it
with nothing told to it. Uninstall is verified across hosts the same way.

Two harness defects surfaced by adding those, both of which made a working
feature look broken:

- the drop test reused the standard `pdf-tables` fixture, which competes with
  the catalogue skill for the same query while the hosts run `topK: 1`. It was
  measuring which of the two ranked higher. Given its own subject now;
- the chain asserted on one particular installed skill's name, but a retrieval
  installs whatever the catalogue ranked — usually several — so with one slot
  the assertion measured the ranker. Any installed skill satisfies it now.

And one thing the harness could not previously distinguish: retrieval fails
open, so a catalogue that was unreachable for a run looked exactly like an
install that did not happen. That is a flaky service versus a broken plugin,
and only the second is a failure — it now reports BLOCKED and says so.

`README.md` gains a shared-library section and, more importantly, a
correction: "What leaves your machine" still described downloads landing in a
throwaway cache outside every scanned directory. They are kept now, and where
they are kept is exactly the kind of thing that section exists to state
plainly.

Verified after the change: ten checks across two real hosts, all passing, and
the install run against the live catalogue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… library

I had said the four unmeasured hosts were "the same code run again on another
host". That was wrong, and this is the evidence: self-registration is wired at
six separate call sites, each with its own config key and its own place in that
host's engine builder, and three "fixed on one side only" bugs have already
come out of this branch.

`e2e_shared_hosts.py` puts one skill in the shared directory, leaves each
host's own directory empty, and asks the one question that skill answers. Same
setup for all of them, so a difference in the result is a difference in that
host's wiring.

All five headless hosts register and retrieve: openclaw 1.x, openclaw 2.0,
DeepSeek Harness, Raven, Hermes.

It found a real defect in the DSH path. `installRootFor(cfg.shareSkills)` did
not compile: `shareSkills` is optional on the exported `Config` interface,
while the zod default only applies to config the harness itself parsed. No
suite in this repository compiles that file against that interface, so it
passed here and failed the moment the harness built it. `?? true` at the call
site, and the comment says why the default is not enough.

Two things about the verdicts, both of which cost a rerun to learn:

Registration and retrieval are separate questions and are now reported apart.
In on-demand mode retrieval goes through the model choosing to call the tool,
and a model that answers from memory instead leaves the wiring untested rather
than broken — measured on both OpenClaw generations, one run each, and driving
1.x's engine directly returned the shared skill in full. A host that registers
but whose model never reached for the tool is INCONCLUSIVE; a host that does
not register is a failure, because then the others cannot see it.

And the label a script uses is not the id a plugin registers. The 1.x package
predates the 2.0 split and still registers as `openclaw`, so comparing against
`openclaw1` reported a correctly registered host as failed.

WorkBuddy has no headless path and is not run here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nnot be inconclusive

The two remaining gaps that were mine to close.

**Acceptance 8 now runs against the live catalogue.** I had argued it was a
property of the swap primitive and therefore adequately unit-tested. That was
half right and the wrong half: the unit test drives `swap_into_place` with a
staging directory that does not exist, which covers the primitive and nothing
in front of it. The path a user hits is decide-to-update, download, extract,
swap — and only the last step was covered.

So: install for real from EverMind SkillHub, then attempt an update to a
version the marker does not have, from an endpoint nothing is listening on.
Measured — the download fails, the directory is unchanged, the previous body
is byte-identical at 2546 characters, and no `.incoming-` or `.retiring-`
directory is left behind. That last one matters on its own: a surviving
staging directory is picked up by the scanner as a skill.

**The OpenClaw verdict no longer depends on the model's mood.** In on-demand
mode retrieval goes through the model choosing to call `skill_search`, and one
that answers from memory instead leaves the wiring untested rather than
broken — which is how a correctly wired host came out INCONCLUSIVE, on
different generations on different runs.

Each generation now also gets an engine probe: the plugin's own `buildEngine`,
given the config the host profile carries, asked to retrieve. No model in the
loop, so it answers the question the wiring actually owns. The model-facing
turn stays, reported beside it as the end-to-end observation.

Both are recorded, and the difference between them is informative rather than
noise: across three runs `engine_probe` was true every time while
`model_found` flipped on both generations.

Five hosts, all passing: openclaw 1.x, openclaw 2.0, DeepSeek Harness, Raven,
Hermes. WorkBuddy has no headless path and still needs a hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The scripts were checked in as they were written; the cases and the results
were not. So the acceptance for this feature lived in a chat log — which is
exactly the failure the earlier E2E documentation request was about, and I
reproduced it.

`cases.md` gains **S1–S9**, numbered the way the shared-skills spec numbers
them so a result can be read against it line by line. They are a separate
family from P1–P6 and the file says why: P1–P6 ask whether one host retrieves
correctly, S1–S9 ask whether hosts can see each other, which needs two of them
running and leaves its evidence in files rather than in one transcript.

It also records the three-fixture rule and the mistake behind it. The hosts run
with `topK: 1`, so two fixtures on one subject makes a case measure which
ranked higher instead of what it was written for — which is how the
hand-dropped case failed once while the feature worked.

`reports/shared-skills.md` is what actually happened at ca32489: all nine
acceptance items, five hosts one at a time, two hosts sharing, and the install
run against the real EverMind SkillHub. With the numbers, not adjectives — S8's
2546 unchanged bytes, the exact skill ids installed, which token proved which
claim.

It also lists the six defects this found, five of them in code every package
suite reported green, and says plainly what was **not** run: the whole library
on WorkBuddy, S5 and S9 on all five hosts rather than the ones named, and
concurrent registration.

`README.md` gains the three scripts and the three things needed to read their
output: registration and retrieval are separate verdicts, a silent catalogue is
BLOCKED rather than FAIL, and the install script talks to the live service on
purpose. INCONCLUSIVE joins the verdict table, since it is new here and means
something specific — the plugin did its part and the model did not exercise it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The feature was documented for whoever reads `tests/host-e2e/`. It was not
documented anywhere the installing agent or the user looks, and it puts a new
directory in their home.

**WorkBuddy gets its acceptance as steps.** It is the one host with no headless
path, so S1–S9 there are a checklist rather than a script — same corpus, same
questions, driven through the UI. Including the one that only matters on this
host: note the registry's mtime, run several turns, check it is unchanged. The
`UserPromptSubmit` hook is a fresh process every turn, so "register at startup"
means "every turn" here, on the turn's hot path inside an 8-second budget, and
the short circuit is what makes that free.

**Both install playbooks gain the library**, because an agent following them
should tell the user what is appearing in their home directory and where their
skills are now visible from. Both switches are named together, since they are
opposites and get conflated: `enabled` in the registry is "others cannot see
me", `shareSkills` in a host's config is "I cannot see others". And deleting a
registry line is explicitly called out as not lasting.

**Uninstall is now correct in four places.** `~/.evermind-skillsearch/` is
shared, so removing it while another agent still has the plugin takes that
agent's library too. The narrow cleanup is this host's line in `registry.json`.
The previous wording would have had an agent delete the lot.

**Both READMEs gain the section**, and the Chinese one gains a correction the
English got earlier: "what leaves your machine" still said downloads land in a
throwaway cache outside every scanned directory. They are kept now, and where
they are kept is exactly what that section exists to state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The WorkBuddy run found a release defect I had left: the branch carries 0.4.0
code while every manifest still says 0.3.0. A marketplace host compares the
version and does not update, so the tester had to uninstall, delete the cache
and reinstall to get the new dist at all. This is the same shape the 0.3.0
review caught, one release later.

Bumped in all fourteen places: six packages, four host manifests, the
marketplace entry, two runtime constants, the root `__version__` and Hermes's
`plugin.yaml`. `verify_release_versions.py` is what proves the set is
complete — it caught the last two after I thought I was done.

`reports/workbuddy-shared-skills.md` is that run, checked in. It is the only
real-host evidence for the one host with no headless path, and it was living in
a share rather than in the repository. Its header says how to read it against
the scripted reports, and a closing table says what happened to each of its
findings.

Two of them changed the cases file.

**P2 failed a new way and the old note did not cover it.** On WorkBuddy's
default model the reply named a skill that does not exist —
`scanned-pdf-invoice-ocr` — without calling `skill_search` at all. "I don't
know" is legible as a miss; a fabricated skill name reads as though retrieval
worked. The case now says so, and says why the assertion is on the fixture's
facts and the tool call rather than on the reply sounding right. Whether the
answer is a reworded tool description or recommending `auto` on that host is a
product decision and is left as one.

**The manual checklist did not say which steps need a second agent**, so a
machine with only WorkBuddy left S2, S5, S6, S7 and S9 blank as "not executed"
when all five were verifiable there. There is now a table, and the steps say
which half of S6 needs company.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ypflll
ypflll self-requested a review September 10, 2026 03:31
ypflll
ypflll previously approved these changes Sep 10, 2026
The release job's version check reads `docs/releases/v<version>.md`, and this
branch bumps the root package to 0.4.0, so the file has to exist before
`build release artifacts` can get past its first step. The requirement landed
in #24, after this branch was written, so it only appeared once main was
merged in.

`check-notes` also returns the previous tag out of the Full Changelog line and
the job then asserts that tag exists and is an ancestor of HEAD, so the line
names v0.3.0 rather than the plugin-scoped tag v0.3.0's own notes used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ypflll
ypflll merged commit 16b5bdf into main Sep 17, 2026
26 checks passed
@ypflll
ypflll deleted the feat/shared-skills branch September 17, 2026 06:09
ypflll pushed a commit that referenced this pull request Sep 17, 2026
Latest Updates stopped at 0.3.0's delivery modes, so the shared skills
library that shipped in #26 does not appear on the front page at all; the
plugin README documents it, the root one never mentions it.

While here: the full corpus had no row of its own in Public artifacts, only
an unticked roadmap box and a line saying it is not published yet. It gets a
row and a "coming soon".

Both READMEs, kept in step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ypflll pushed a commit that referenced this pull request Sep 17, 2026
Latest Updates stopped at 0.3.0's delivery modes, so the shared skills
library that shipped in #26 does not appear on the front page at all; the
plugin README documents it, the root one never mentions it.

While here: the full corpus had no row of its own in Public artifacts, only
an unticked roadmap box and a line saying it is not published yet. It gets a
row and a "coming soon".

Both READMEs, kept in step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ypflll pushed a commit that referenced this pull request Sep 17, 2026
Latest Updates stopped at 0.3.0's delivery modes, so the shared skills
library that shipped in #26 does not appear on the front page at all; the
plugin README documents it, the root one never mentions it.

While here: the full corpus had no row of its own in Public artifacts, only
an unticked roadmap box and a line saying it is not published yet. It gets a
row and a "coming soon".

Both READMEs, kept in step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ypflll added a commit that referenced this pull request Sep 17, 2026
Latest Updates stopped at 0.3.0's delivery modes, so the shared skills
library that shipped in #26 does not appear on the front page at all; the
plugin README documents it, the root one never mentions it.

While here: the full corpus had no row of its own in Public artifacts, only
an unticked roadmap box and a line saying it is not published yet. It gets a
row and a "coming soon".

Both READMEs, kept in step.

Co-authored-by: yao pengfei <yaopengfei@shanda.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants