Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .agents/plugins/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,18 @@
},
"category": "Developer Tools"
},
{
"name": "jafar-perf",
"source": {
"source": "local",
"path": "./plugins/jafar-perf"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Developer Tools"
},
{
"name": "perf-engineer",
"source": {
Expand Down
6 changes: 6 additions & 0 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,12 @@
"version": "0.1.0",
"source": "./plugins/jfr-analyzer"
},
{
"name": "jafar-perf",
"description": "JVM performance engineering with the Jafar MCP server: triage, CPU, latency, GC, memory-leak and regression playbooks, plus specialist subagents.",
"version": "0.1.0",
"source": "./plugins/jafar-perf"
},
{
"name": "perf-engineer",
"description": "Evidence-driven Java optimization investigations with JFR, BTrace, and JMH.",
Expand Down
91 changes: 91 additions & 0 deletions .github/workflows/tool-drift.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
name: Tool drift

# The skills and agents of the jfr-mcp plugins name the server's tools explicitly, and the server is
# developed in btraceio/jafar. The drift this job exists to catch therefore happens when *that*
# repository changes, not when this one does, so a push-triggered job alone would never run at the
# moment it matters. The weekly schedule is the point; the push and pull_request triggers only catch
# typos and keep the script itself working.
on:
schedule:
- cron: '17 6 * * 1' # Mondays, 06:17 UTC
push:
branches: [main]
paths:
- 'plugins/**'
- 'scripts/check-tool-references.js'
- '.github/workflows/tool-drift.yml'
pull_request:
paths:
- 'plugins/**'
- 'scripts/check-tool-references.js'
- '.github/workflows/tool-drift.yml'
workflow_dispatch:

permissions:
contents: read

jobs:
check:
name: Skills reference tools that exist
runs-on: ubuntu-latest
timeout-minutes: 15
permissions:
contents: read
issues: write # only used to report a scheduled failure
steps:
- name: Check agent-plugins
uses: actions/checkout@v5

- name: Set up Node.js
uses: actions/setup-node@v5
with:
node-version: 20

- name: Set up Java for JBang
uses: actions/setup-java@v4
with:
distribution: temurin
java-version: '21'

# Deliberately the same install path the jfr-mcp README tells users to run, so a break in
# that path shows up here rather than in someone's terminal.
- name: Install the published MCP server
run: |
curl -Ls https://raw.githubusercontent.com/btraceio/jafar/main/jfr-mcp/install.sh | bash
echo "$HOME/.jbang/bin" >> "$GITHUB_PATH"

- name: Check every referenced tool exists
run: node scripts/check-tool-references.js

# A scheduled run failing means upstream moved, and nobody is watching a cron job. An issue
# is how that reaches a human.
- name: Open an issue when the scheduled run finds drift
if: failure() && github.event_name == 'schedule'
uses: actions/github-script@v7
with:
script: |
const title = 'Skills reference MCP tools the published server no longer exposes';
const existing = await github.rest.issues.listForRepo({
owner: context.repo.owner, repo: context.repo.repo,
state: 'open', labels: ['tool-drift'],
});
const body = [
'The weekly tool-drift check failed: at least one skill or agent names an MCP tool',
'that `jfr-mcp@btraceio` does not expose. That usually means a tool was renamed or',
'removed in btraceio/jafar.',
'',
`Run: ${context.serverUrl}/${context.repo.owner}/${context.repo.repo}/actions/runs/${context.runId}`,
'',
'The log names each tool and the exact file and line that references it.',
].join('\n');
if (existing.data.length === 0) {
await github.rest.issues.create({
owner: context.repo.owner, repo: context.repo.repo,
title, body, labels: ['tool-drift'],
});
} else {
await github.rest.issues.createComment({
owner: context.repo.owner, repo: context.repo.repo,
issue_number: existing.data[0].number, body,
});
}
13 changes: 13 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,3 +13,16 @@ scripts/validate-marketplace.sh

For behavioral changes, add or update an eval case and include a captured response review. Use
conventional commit messages and do not change plugin versions until a release is being published.

Plugins that start the `jfr-mcp` server (`jafar-perf`, `jfr-analyzer`) name its tools in their skills
and agents, and the server lives in another repository. After changing those names, or when a new
`jfr-mcp` release lands, check that every referenced tool still exists:

```sh
node scripts/check-tool-references.js # published server, via jbang --fresh
node scripts/check-tool-references.js --jar path/to.jar # a locally built shadow jar
```

It starts the server, asks it for `tools/list`, and fails naming each missing tool with its file and
line. It is not part of `validate-marketplace.sh`, because it needs the published server.

1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ one another so the workflow instructions and supporting scripts are maintained o
| `btrace-development` | Repository conventions and build guidance for BTrace development. |
| [`btrace-observability`](plugins/btrace-observability/README.md) | A composable SRE skill suite for diagnosing Java applications with BTrace probes. |
| [`jfr-analyzer`](plugins/jfr-analyzer/README.md) | Systematic profile analysis with optional BTrace live-probe correlation. |
| [`jafar-perf`](plugins/jafar-perf/README.md) | Guided JVM performance analysis on the Jafar MCP server: playbooks for CPU, latency, GC and memory, and specialist subagents. |
| [`perf-engineer`](plugins/perf-engineer/README.md) | Evidence-driven optimization investigations with JFR, BTrace, and JMH. |

## Layout
Expand Down
1 change: 1 addition & 0 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@
"./plugins/btrace-development/skills",
"./plugins/btrace-observability/skills",
"./plugins/jfr-analyzer/skills",
"./plugins/jafar-perf/skills",
"./plugins/perf-engineer/skills"
]
}
Expand Down
22 changes: 22 additions & 0 deletions plugins/jafar-perf/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
{
"name": "jafar-perf",
"version": "0.1.0",
"description": "JVM performance engineering with the Jafar MCP server: triage, CPU, latency, GC, memory-leak and regression-comparison playbooks, plus specialist analysis subagents.",
"author": {
"name": "BTrace",
"url": "https://github.com/btraceio"
},
"repository": "https://github.com/btraceio/agent-plugins",
"license": "Apache-2.0",
"skills": "./skills/",
"mcpServers": "./.mcp.json",
"keywords": [
"jfr",
"jvm",
"performance",
"profiling",
"heap-dump",
"pprof",
"otlp"
]
}
38 changes: 38 additions & 0 deletions plugins/jafar-perf/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
{
"name": "jafar-perf",
"version": "0.1.0",
"description": "JVM performance engineering with the Jafar MCP server: triage, CPU, latency, GC, memory-leak and regression-comparison playbooks, plus specialist analysis subagents.",
"author": {
"name": "BTrace",
"url": "https://github.com/btraceio"
},
"repository": "https://github.com/btraceio/agent-plugins",
"license": "Apache-2.0",
"skills": "./skills/",
"mcpServers": "./.mcp.json",
"keywords": [
"jfr",
"jvm",
"performance",
"profiling",
"heap-dump",
"pprof",
"otlp"
],
"interface": {
"displayName": "Jafar Performance Engineer",
"shortDescription": "Guided JVM performance analysis with evidence for every claim.",
"longDescription": "Turns the Jafar MCP server into a guided JVM performance analyst: methodology skills for CPU, latency, GC, memory and heap investigations, plus specialist subagents that cite the tool call behind every claim.",
"developerName": "BTrace",
"category": "Developer Tools",
"capabilities": [
"Guidance",
"Analysis"
],
"defaultPrompt": [
"Analyse this JFR recording and tell me why p99 latency doubled.",
"Compare these two recordings and tell me what regressed.",
"Use perf-lead to review this heap dump and recording."
]
}
}
13 changes: 13 additions & 0 deletions plugins/jafar-perf/.mcp.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
{
"mcpServers": {
"jafar": {
"type": "stdio",
"command": "jbang",
"args": [
"jfr-mcp@btraceio",
"--stdio",
"--attach"
]
}
}
}
97 changes: 97 additions & 0 deletions plugins/jafar-perf/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
# jafar-perf — performance engineer in a box

A plugin (Claude Code and Codex) that turns the [Jafar MCP server](https://github.com/btraceio/jafar/blob/main/jfr-mcp/README.md) into a guided
JVM performance analyst.

The MCP server already exposes 37 analysis tools. What it does not carry is the *methodology*:
which question to ask next, which tool answers it, what counts as evidence, and how to report.
This plugin is that layer.

## Install

```
/plugin marketplace add btraceio/agent-plugins
/plugin install jafar-perf@btraceio-agent-plugins
```

The plugin bundles `.mcp.json`, so installing it also registers the `jafar` MCP server
(`jbang jfr-mcp@btraceio --stdio --attach`). [JBang](https://www.jbang.dev) must be on your PATH;
it fetches the server on first use. No separate `claude mcp add` is needed.

`--attach` makes every session share one daemon, started on demand, instead of each starting its
own JVM, so a new session answers in a fraction of a second and agents working in parallel get
their own sessions on the same server. It needs a `jfr-mcp` release that has the flag; an older
one ignores it and runs a private server, which still works. See the
[daemon documentation](https://github.com/btraceio/jafar/blob/main/doc/mcp/Daemon.md#one-shared-daemon-for-many-stdio-clients---attach).

## What is in it

### Skills

Invoked automatically when the work matches, or explicitly as `/jafar-perf:<name>`.

| Skill | Covers |
|---|---|
| `triage` | First step on any unfamiliar artifact: what it contains, what is anomalous, where to go next |
| `cpu` | Hot methods, call paths, convergence points, attributing samples to work |
| `latency` | Contention, parking, executor queue saturation, blocking I/O, per-endpoint attribution |
| `gc` | Pause distribution as a fraction of wall clock, heap behaviour, allocation hotspots |
| `memory-leak` | Retained sizes, dominators, GC root paths, leak detectors, heap-to-JFR correlation |
| `heap-diff` | Proving growth with two dumps instead of inferring it from one |
| `compare` | Before/after regression checks with a stated noise floor |
| `jfrpath` | Syntax reference for JfrPath, HdumpPath and SamplesPath |
| `report` | The output format and the evidence discipline every finding must meet |

### Agents

| Agent | Role |
|---|---|
| `perf-lead` | Triages, dispatches the specialists the evidence justifies, merges and ranks their findings |
| `perf-engineer` | General-purpose analyst for a single artifact, end to end |
| `cpu-analyst` | CPU-bound analysis |
| `concurrency-analyst` | Thread states, contention, queues |
| `memory-analyst` | GC and allocation |
| `heap-analyst` | Heap dumps and retention |
| `io-analyst` | File and socket I/O |

Specialists carry narrow tool allowlists, so each one works within its dimension rather than
wandering across the whole surface.

## Using it

Point it at an artifact and ask:

> Analyse `/tmp/recording.jfr` and tell me why p99 latency doubled after the last deploy.

For a broad investigation, ask for the lead agent, which fans out to specialists and merges
their findings:

> Use perf-lead to review `/tmp/recording.jfr`.

For a regression check, open both recordings and compare:

> Compare `/tmp/before.jfr` against `/tmp/after.jfr` and tell me what regressed.

## The standard these skills enforce

Every skill in this plugin pushes the same discipline, because it is what separates a
performance report from a guess:

- **Every claim names the tool call that produced it.** If you cannot cite the call and the
numbers, the claim does not go in the report.
- **Rates, not counts.** Absolute counts are meaningless without the recording's duration and
misleading across recordings of different lengths.
- **Sampling is not measurement.** Sampled data is labelled as sampled, and frames below the
noise floor are not findings.
- **Absence of evidence is reported as such.** "No allocation hotspots found" is wrong when
allocation profiling was never enabled; `jfr_diagnose` returns `capabilityGaps` for exactly
this reason, and they belong in the report.
- **No claimed improvement without a measured comparison.**

## Without Claude Code

The methodology is also available from the server itself, so other MCP clients get it too:
prompts (`triage`, `compare`, `leak-hunt`, `latency`) and resources (`jafar://sessions`,
`jafar://help/jfrpath`, `jafar://help/hdumppath`, `jafar://help/tools`). The skills here go
further — they carry the interpretation rules and the failure modes — but the prompts cover
the sequence.
24 changes: 24 additions & 0 deletions plugins/jafar-perf/agents/concurrency-analyst.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
---
name: concurrency-analyst
description: Specialist for thread and contention analysis of a JFR recording — thread states, monitor contention, parking, executor queue saturation, and per-endpoint latency attribution. Dispatch when triage shows threads blocked or waiting rather than running, or when the complaint is p99 latency rather than throughput.
tools: mcp__jafar__jfr_tsa, mcp__jafar__jfr_use, mcp__jafar__jfr_query, mcp__jafar__jfr_list_types, mcp__jafar__jfr_stackprofile, mcp__jafar__pprof_tsa, Read, Grep, Glob
skills: latency, report
model: sonnet
---

You analyse what threads are waiting for. Follow the `latency` skill; report in the
`report` format.

Run `jfr_tsa` with `correlateBlocking=true` first and read `stateDistribution` before
anything else — if the recording is RUNNABLE-dominated this is a CPU question and you
should say so rather than manufacturing a contention story.

Keep `jdk.JavaMonitorEnter` (blocked acquiring) separate from `jdk.JavaMonitorWait`
(waiting on a condition); they mean different things. Rank monitors by summed duration
relative to wall clock, never by event count. Treat executor queue saturation as
first-class: queued work cannot be recovered by faster methods.

Two honesty requirements: JFR monitor events have a duration threshold, so absence of
events is not absence of contention — check `jdk.ActiveSetting` if it matters. And
`decorateByTime` correlations are concurrency in time, not causation; report them as
"concurrent with".
20 changes: 20 additions & 0 deletions plugins/jafar-perf/agents/cpu-analyst.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
---
name: cpu-analyst
description: Specialist for CPU-bound analysis of a JFR recording or sampling profile — hot methods, call paths, convergence points, and per-thread or per-endpoint attribution of execution samples. Dispatch when triage shows high execution-sample counts or a RUNNABLE-dominated thread state distribution.
tools: mcp__jafar__jfr_hotmethods, mcp__jafar__jfr_flamegraph, mcp__jafar__jfr_callgraph, mcp__jafar__jfr_stackprofile, mcp__jafar__jfr_query, mcp__jafar__jfr_list_types, mcp__jafar__pprof_hotmethods, mcp__jafar__pprof_flamegraph, mcp__jafar__otlp_flamegraph, Read, Grep, Glob
skills: cpu, report
model: sonnet
---

You analyse where CPU time goes. Follow the `cpu` skill; report in the `report` format.

Start with `jfr_hotmethods` to learn whether the profile is concentrated or flat, then pick
the follow-up that shape calls for — bottom-up for a concentrated profile, top-down or
callgraph for a flat one. Confirm every hotspot against the three tests in the `cpu` skill
(above the noise floor, steady across time buckets, not one unrepresentative thread).

Stay in your lane: time spent parked, blocked or waiting on I/O is not CPU cost. If the
profile shows the cost is waiting, say so and hand it back rather than analysing it here.

Return findings with the tool call and numbers behind each, and the source location if you
can find it in the working tree.
Loading
Loading