Skip to content

Add jfr-mcp --attach: one shared daemon for many stdio clients, with per-client sessions - #119

Merged
jbachorik merged 5 commits into
mainfrom
feat/mcp-attach-bridge
Oct 3, 2026
Merged

jbachorik merged 5 commits into
mainfrom
feat/mcp-attach-bridge

Conversation

@jbachorik

@jbachorik jbachorik commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

What & why

Every MCP client session used to start its own jfr-mcp JVM: 4 s or more before the first useful answer, and nothing shared between clients. jfr-mcp --stdio --attach is a stdio server that is really a thin client of one shared per-user SSE daemon, started on demand. A warm attach answers initialize in about 0.2-0.3 s.

Making a shared daemon the normal setup exposed three gaps in the SSE mode that this PR also closes: a client that left was never noticed, one agent's closeAll closed everyone's sessions, and a stale daemon from an older jar kept serving after an upgrade.

Review commit by commit. Each commit is one step and its message carries its own evidence. The branch can be split into a stack if you would rather review it that way (the natural cut is 32c217b per-client sessions, then cc30c96/6576208 plumbing, then the bridge and skew policy), but the later commits build on the earlier ones.

Commit Change
cc30c96 jafar.state.dir, authenticated GET /mcp/health, idle exit for auto-started daemons, SSE keep-alive pings
6576208 McpEndToEndTest no longer reads and rewrites the developer's real ~/.jafar (its server child did not inherit the isolation property)
cf6a979 The bridge, the Main entry class, and the daemon's AOT cache
32c217b Per-client session ownership across all four registries
4f4aac3 Replace a stale idle auto-started daemon (POST /mcp/shutdown), and the docs

Design points worth a reviewer's attention:

  • Raw relay, replayed handshake. The bridge relays JSON-RPC lines and understands only what it needs to survive a daemon restart: it caches the client's initialize and notifications/initialized and replays them on the new session; requests in flight get a JSON-RPC error instead of hanging.
  • Cheap to start. io.jafar.mcp.Main is now the jar's Main-Class and dispatches --attach before JafarMcpServer is loaded, since loading that class alone pulls in Jetty, the SDK and Reactor (~850 classes, over a second). The bridge also avoids HttpClient and ObjectMapper (about 300 ms each) and uses HttpURLConnection and Jackson's streaming parser. A fresh-JVM test asserts none of those load.
  • Keep-alive is load-bearing. The SDK drops an SSE session only when a write to it fails, and nothing writes to an idle stream. I measured a closed client still counted after 150 s. Without the pings, activeSessions over-reports and the idle exit can never fire.
  • Ownership. SessionOwnership (shell-core) holds owner-per-session, per-client aliases and visibility, and all four registries delegate to it. closeAll closes the caller's own sessions plus shared ones restored after a restart (they appear in the caller's listings and nobody owns them), never another client's. Persistence is deliberately unchanged: stdio sessions have a scope too, so persisting only unowned ones would have broken restart recovery.
  • AOT cache for the daemon only. The first daemon the bridge starts trains an AOT cache (-XX:AOTCacheOutput, written when it exits on idle) and later ones use it: daemon start 1.0 s to ~0.35 s. It is named by jar version, mtime and JVM version; a marker holding the training daemon's pid stops a second JVM writing the same file while the first dumps it. Needs JDK 25+ (guarded), and JDK 27 is not required.
  • Rejected: HttpClient/ObjectMapper in the bridge (startup cost); a shared JFRSession per path (measured: nothing is cached per session, so it would buy nothing); starting the daemon through jbang (a launchd/systemd environment often has no jbang on PATH); AOT flags on the bridge's own JVM (about 60 ms gain, and JVM logging defaults to stdout, which would corrupt the MCP stream).

How verified

Command Result
./gradlew spotlessCheck exit 0
./gradlew :jfr-mcp:test (forced rerun, clean test store) 389 passed, 0 failed, 0 skipped
./gradlew :jfr-mcp:test -DenableIntegrationTests=true on SessionRegistryTest, SessionRegistryClientIsolationTest, SessionRegistryOwnershipTest 26 passed
./gradlew :shell-core:test 342 run, 7 fail, the same 7 by name as on main (5 need test-jfr.jfr, which is not downloaded; 2 LlmConfigTest read the local model config)
Built jar, real daemon two bridges started at once spawn one daemon; it survives its bridges' process groups being killed; a bridge recovers after kill -9 of the daemon; stdout carries only JSON-RPC
Built jar, ownership two clients both open alias=cpu; B cannot read or close A's id; A's closeAll reports 1 and leaves B alone; B's sessions are released after it disconnects
Built jars, version skew an old idle auto-started daemon is replaced (new pid and version, old one exits after writing its cache); a daemon with clients connected is left alone and the stderr warning reads correctly
Built jar, AOT training daemon writes a 44 MB cache on idle exit; the next cold attach including daemon start: 2.83 s to 0.74 s
Real Claude Code 2.1.288 (claude -p, strict scratch MCP config) through the bridge server connected; jfr_open, jfr_summary, jfr_close all succeed; two concurrent Claude sessions both use alias='real' and succeed against one daemon

Several tests were shown to fail with their fix removed (reconnect replay, spawn lock, --attach entry class, per-client closeAll, the connected-clients guard).

Coverage notes

Tests cover the bridge over real HTTP against a fake SSE daemon (relay, auth, reconnect and handshake replay, in-flight failure, a replay the daemon never answers), the locator (discovery, token permissions, spawn race, stale ports, skew policy), the launcher's command building and AOT plan, the ownership rules and their use in each registry, the close tools' per-caller counts, and the jafar://sessions resource. Tests deliberately do not cover starting a real detached daemon from inside the suite; that is exercised by the end-to-end runs above.

Not verified / known limitations

  • Windows. The launcher refuses to start a daemon there with an explicit message; attaching to one you started yourself should work but has not been run.
  • The released flow. jbang jfr-mcp@btraceio --stdio --attach cannot work until a release contains the flag (the published jar ignores it and runs a private server, so nothing breaks). All end-to-end runs used the built jar directly.
  • A training daemon killed mid-dump could leave a partial AOT cache; the JVM falls back to a normal start but would not retrain until the file is deleted.
  • Disconnect release was timed with a 5 s keep-alive; at the default 30 s a departed client is noticed within about a minute.
  • The version compare uses the manifest version, so rebuilt SNAPSHOTs of the same version are not seen as different.
  • The heap registry's lock-free open is covered by unit tests, not under real contention.

Opened as a draft because of the items above; it was marked ready for review once CI was green (Tests (JDK 21) and Tests (JDK 8) pass, 2,294 tests, 0 failures, 16 skipped). They remain known limitations, and none of them blocked merging.

Breaking changes / migration

For SSE/daemon users, behaviour changes in ways that are the point of this PR but are visible:

  • Sessions are per client: other clients' sessions are no longer listed, reachable by id, or closable (Session not found), aliases are unique per client, and closeAll closes the caller's sessions plus shared restored ones. A client that disconnects has its sessions closed.
  • Keep-alive pings are on by default for every SSE daemon, supervised ones included (-Dmcp.sse.keepalive.seconds=0 turns them off).
  • The jar's Main-Class is io.jafar.mcp.Main. JafarMcpServer.main still works and also honours --attach, but callers that invoke it directly pay for loading the server first. The published jbang alias resolves the jar's manifest, so it is unaffected.
  • SessionRegistry, HeapSessionRegistry and SamplingSessionRegistry closeAll() now return a count (int, was void): source-compatible, not binary-compatible for precompiled external callers.
  • SessionRegistryClientIsolationTest now cleans up with shutdown(), since its cleanup relied on closeAll() reaching other clients' sessions.

No migration is needed for stdio users. --attach is opt-in.

Links

Companion PR: btraceio/agent-plugins#8 moves the jafar-perf plugin there and points both jfr-mcp plugins at --attach; it depends on this one. Docs: doc/mcp/Daemon.md, "One shared daemon for many stdio clients".

Noticed, not changed: DefaultQueryEvaluatorConsumeTest never closes the session it opens, leaving an entry in the shared test persistence store that a global closeAll() used to sweep up; McpJfrTransportTest.jfrDiagnoseReturnsReport takes ~14 s of its 15 s timeout, on main too.

🤖 Generated with Claude Code

jbachorik and others added 5 commits October 3, 2026 10:54
Groundwork for a stdio bridge that auto-starts and attaches to one shared
SSE daemon per user.

- jafar.state.dir redirects the port file, auth token and persisted
  sessions (default stays ~/.jafar), so tests and side-by-side runs can
  no longer touch the real daemon's files.
- GET /mcp/health, behind the bearer filter, reports version, pid,
  startedAt, activeSessions and whether the daemon was auto-started, so
  a client can decide to use, replace or leave a daemon it found.
- An auto-started daemon (-Dmcp.daemon.autostarted=true) exits after
  mcp.daemon.idle.timeout.minutes (default 30) with zero connected
  clients. Supervised daemons are unaffected.
- SSE keep-alive pings (mcp.sse.keepalive.seconds, default 30, 0 turns
  them off). The SDK drops a session only when a write to it fails and
  nothing writes to an idle stream, so a closed client stayed counted
  for 150s+ and the idle exit could never fire. This applies to every
  SSE daemon, supervised ones included.

Verified: spotlessCheck clean; :jfr-mcp:test 309 passed, 0 failed (with
test-dd.jfr present). On the built jar: health answers 401/401/200, a
connected client holds the daemon up past the timeout, and the daemon
exits 0 about 110s after the client leaves, leaving no state files.
The keep-alive fix was red (session still counted after 150s) before
and green (dropped by ~60s) after. StateDirTest's wiring check fails
with the SseAuthToken change reverted.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The test starts the server as a separate JVM, so the isolation Gradle
sets for the test JVM (jafar.mcp.sessions.file) never reached it. Every
run read and rewrote the developer's real ~/.jafar/mcp-sessions.json,
and the registry's restore step deletes entries whose recording file is
missing. Pass -Djafar.state.dir=<fresh temp dir> to the child, for both
the java and the jbang launch paths.

Verified: before the change, running only McpEndToEndTest changed the
real file's mtime (bisected over four test classes; it was the only one
that did). After it, a forced rerun leaves the mtime untouched, with the
test actually executing (tests=2, skipped=0, failures=0).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Every MCP client session used to start its own jfr-mcp JVM, paying 4s+
for startup and JIT warm-up and sharing nothing. The bridge is a thin
stdio server that relays to one long-lived per-user SSE daemon, starting
it if nobody has. Measured: a private stdio server needed ~4s before its
first useful answer; a warm attach answers initialize in 0.2-0.3s.

- StdioSseBridge: raw JSON-RPC relay, one stdin line to one POST, one SSE
  event to one stdout line. It caches the client's initialize and
  notifications/initialized and replays them on a fresh session when the
  daemon forgets this one (the SDK answers 404 after a restart), and
  fails requests that were in flight with a JSON-RPC error instead of
  leaving the client waiting.
- DaemonLocator: finds the daemon through the port file (SsePortRegistry,
  so one definition), refuses a token file others can read, and
  serialises starting it with a file lock so several clients opening at
  once launch exactly one daemon.
- DaemonLauncher: runs `java -jar` directly (never jbang, which a
  launchd/systemd environment may not have on PATH) from a shell with job
  control and nohup so the daemon outlives the bridge's process group.
  POSIX only; on Windows it fails with an explicit message.
- AOT cache for that daemon: with no cache yet the first daemon is the
  training run (-XX:AOTCacheOutput, written when it exits on idle), and
  later ones start from it (-XX:AOTCache), cutting daemon startup from
  ~1.0s to ~0.35s. The cache is named by jar version+mtime and JVM
  version, other versions' files are deleted, and a marker holding the
  training daemon's pid stops a second JVM writing the same file during
  the seconds the first spends dumping it.
- io.jafar.mcp.Main is now the jar's Main-Class and dispatches --attach
  before JafarMcpServer is loaded: loading that class alone pulls in
  Jetty, the MCP SDK and Reactor (~850 classes, over a second). The
  bridge also avoids HttpClient (TLS init, ~300ms) and ObjectMapper
  (~300ms) in favour of HttpURLConnection and Jackson's streaming parser.
  HttpURLConnection.disconnect() blocks on an idle chunked SSE stream, so
  Connection.close() does it off-thread.
- An older jar ignores --attach and runs a private stdio server, so the
  flag degrades to the old behaviour rather than failing.

Verified: spotlessCheck clean; :jfr-mcp:test 350 passed, 0 failed (41 in
the bridge package, incl. a fresh-JVM test that attaching never loads the
server, Jetty, the MCP SDK, databind, HttpClient or TLS). On the built
jar: two bridges started at once spawn one daemon; the daemon survives
its bridges' process groups being killed; a bridge recovers after the
daemon is kill -9'd; stdout carries only JSON-RPC frames; the training
daemon writes a 44MB cache on idle exit and the next one starts from it
(attach incl. daemon start 2.83s -> 0.74s). Reconnect-replay, spawn-lock
and --attach-entry tests were each shown to fail with their fix removed.
Not verified: Windows, a real Claude Code client, JDK 27 GA.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
One registry serves every client of the SSE daemon, so one agent could
close another's recordings with closeAll, read them by guessing a numeric
id, or be refused an alias another agent already held. With the bridge
making a shared daemon the normal setup, that stops being theoretical.

A session now belongs to the client that opened it (RequestScope), tracked
by the new shell-core SessionOwnership and used by all four registries
(JFR, heap, and the shared pprof/OTLP one):

- Other clients cannot see, reach or close it; aliases only need to be
  unique per client. Sessions restored after a restart belong to nobody
  and are visible to all. A caller with no scope (the server itself,
  tests) is unrestricted.
- closeAll closes what the caller can see: its own and the shared restored
  ones, never another client's, and returns the count. Shared ones are
  included because they appear in the caller's listings; a closeAll that
  left them behind would look broken.
- retainScopes also closes the sessions of clients that have gone, which
  nobody can reach any more.
- open() builds the session outside the registry lock, so one client
  parsing a large recording no longer stalls everyone's lookups; the alias
  is re-checked afterwards and a loser of that race closes its session.
- jfr_close and hdump_close report the caller's own counts, and the
  jafar://sessions resource runs as the asking client instead of listing
  every client's file paths.
- Persistence is unchanged: stdio sessions have a scope too, so persisting
  only unowned ones would have broken restart recovery there.

SessionRegistryClientIsolationTest now cleans up with shutdown(): its
cleanup relied on closeAll() reaching other clients' sessions.

Verified: spotlessCheck clean; :jfr-mcp:test 376 passed, 0 failed;
integration-tagged SessionRegistryTest, SessionRegistryClientIsolationTest
and SessionRegistryOwnershipTest 26 passed. :shell-core:test has 7
failures identical by name to the previous commit (5 need an absent
test-jfr.jfr, 2 LlmConfigTest read this machine's model config). On the
built jar with two bridge clients: both used alias=cpu, B could not read
or close A's id, A's closeAll reported 1 and left B alone, and B's two
sessions were released after it disconnected. The per-client closeAll test
was shown to fail with its fix removed.
Noted, not changed: DefaultQueryEvaluatorConsumeTest never closes its
session, leaving an entry in the shared test store that a global closeAll
used to sweep up; McpJfrTransportTest.jfrDiagnoseReturnsReport takes
~14s against a 15s timeout, on the previous commit too.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A bridge that finds a daemon of another version used to talk to it
silently, so an upgraded jar kept running old code until someone noticed.

The bridge now reads /mcp/health and compares versions. Only a daemon
that a bridge started, that has no client connected, and that is the
wrong version is replaced: the bridge POSTs the new /mcp/shutdown, waits
for the old daemon to go (its token file disappears before its port, so
the wait tolerates that), and starts a current one. Anything else is left
alone and reported once on stderr: a daemon with clients would take their
sessions with it, a supervised one is its supervisor's to restart, and a
daemon from before /mcp/health cannot be verified at all.

POST /mcp/shutdown sits behind the bearer filter and answers 202 only for
an auto-started daemon with zero clients; a daemon that cannot count its
clients refuses too, since unknown is not idle. A 409 also protects a
client that connected just after the bridge looked.

doc/mcp/Daemon.md gains the --attach section (what it starts, the
settings, per-client sessions, upgrades, /mcp/health and /mcp/shutdown)
and the isolation paragraph is corrected; jfr-mcp/README.md points to it.

Verified: spotlessCheck clean; :jfr-mcp:test 389 passed, 0 failed (83 in
the bridge and lifecycle packages). On real jars (one with a patched
manifest version): an old idle auto-started daemon was replaced by the
new version, the old one exited ~4s later after writing its AOT cache, and
a newer daemon with clients connected was left alone with the stderr
warning. The documented curl commands were run against a scratch daemon
and the internal doc links checked against the headings. Removing the
connected-clients guard makes its test fail. Not run: the jbang and
claude mcp add lines in the docs, since the released jar has no --attach.
Version comparison uses the manifest version, so rebuilt SNAPSHOTs of the
same version are not seen as different.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Combined JUnit Test Report

  • Total: 2294
  • Passed: 2278
  • Failures: 0
  • Errors: 0
  • Skipped: 16

HTML Test Reports

Run artifacts: https://github.com/btraceio/jafar/actions/runs/37145655941

@jbachorik
jbachorik merged commit ab4129b into main Oct 3, 2026
4 checks passed
@jbachorik
jbachorik deleted the feat/mcp-attach-bridge branch October 3, 2026 19:07
jbachorik added a commit to btraceio/jafar-perf-box that referenced this pull request Oct 3, 2026
The jafar-perf plugin moved into the btraceio/agent-plugins marketplace so the btraceio
agent plugins sit in one place instead of two with overlapping names (this repository's
perf-engineer agent and jfr-analyzer's, for one). The README now says so at the top, with
the commands to switch: for Claude Code, remove this marketplace (named btraceio), add
btraceio/agent-plugins and install jafar-perf from it; for pi, remove this package and
install btraceio/agent-plugins; or re-run Jafar's installer, which migrates both.

Pushed once the move was complete: btraceio/agent-plugins#8 and #10 and btraceio/jafar#119 and
#120 are merged, jfr-mcp 0.29.0 is released, and the jbang catalog resolves to it. The
repository is archived right after.

Verified: `claude plugin marketplace remove`, `pi remove` and `pi install` exist in the
installed CLIs and are what the notice says; the installer's migration of an install made from
this repository was run in a scratch HOME (Claude Code and pi) and left only the new copies.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant