Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 43 additions & 18 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,14 +2,27 @@

A structured, agent-queryable knowledge base of **AMD Instinct GPU kernel
optimization** for CDNA3 (gfx942 / MI300) and CDNA4
(gfx950 / MI350–MI355X), packaged as a **Codex CLI skill** (and compatible with Claude Code).
The repository root **is** the skill directory, so the full corpus remains
git-updatable instead of being copied into a separate wrapper.

> **Corpus dates:** the generated
> [`data/corpus-manifest.yaml`](data/corpus-manifest.yaml) owns the merged-PR
> cutoff; each doc/blog page carries its own retrieval date. The nod-ai AMDGPU optimization guide is synced
> through commit `efa471ae` on **2026-07-20**. Tool versions remain pinned in
(gfx950 / MI350–MI355X). It is packaged as a **Codex CLI skill**, remains
compatible with Claude Code, and uses the repository root as the skill
directory so one `git pull` updates both tooling and corpus.

ROCmKernelWiki couples architecture-scoped synthesis to merged-PR and
primary-source provenance, [real-silicon validation](VERIFICATION.md),
and a [maintainer-controlled evidence flywheel](#self-evolution-boundary).
Automation discovers, triages, and proposes evidence in a rolling PR that starts
in Draft; maintainers remain responsible for accepting facts and merging changes.

> **Corpus freshness:** the generated
> [`data/corpus-manifest.yaml`](data/corpus-manifest.yaml) records the last
> complete baseline PR-harvest cutoff from
> [`data/refresh-cutoff.yaml`](data/refresh-cutoff.yaml). Incremental refreshes
> may add selected evidence after that date;
> [`data/evolution-state.yaml`](data/evolution-state.yaml) records each source's
> latest discovery position, not a complete corpus through-date. Doc/blog
> retrieval dates and guide-sync boundaries advance independently. The captured
> nod-ai AMDGPU guide commit and retrieval date live in its
> [canonical source page](sources/blogs/blog-amdgpu-kernel-opt-guide.md). Tool
> versions remain pinned in
> [`data/tool-versions.yaml`](data/tool-versions.yaml). gfx950 facts and examples
> were verified on MI350X, with guide-specific device/LDS checks repeated on
> MI355X — see below.
Expand Down Expand Up @@ -208,33 +221,44 @@ Three layers (after MIT Han Lab's KernelWiki, in turn after Karpathy's LLM-wiki)

Supporting files: `data/` holds the schema and controlled vocabulary
(`schemas.yaml`, `tags.yaml`, `aliases.yaml`, `inclusion-policy.yaml`,
`tool-versions.yaml`, `refresh-cutoff.yaml`, `hardware-verified.yaml`);
`scope.yaml`, `sources.yaml`, `tool-versions.yaml`, `refresh-cutoff.yaml`,
`hardware-verified.yaml`);
`candidates/` holds per-repo PR ledgers; `references/` holds the primer, schema, and
worked examples.

## Maintenance Tooling

[`data/sources.yaml`](data/sources.yaml) is the canonical registry for the
refresh pipeline. Each run enforces file and line budgets so the rolling PR
remains reviewable. The daily worker creates that PR as Draft, then updates it
from a disposable clone. Approved hardware tasks are evaluated separately on the
trusted MI355 node against an exact candidate SHA.

| Script | Purpose |
|---|---|
| `scripts/evolve/discover.py` | Incrementally discover PR/tree evidence from `data/sources.yaml` |
| `scripts/harvest_prs.py` | Compatibility wrapper for merge-safe incremental discovery |
| `scripts/evolve/gaps.py` | Turn uncovered evidence clusters into synthesis proposals |
| `scripts/evolve/synthesize.py` | Validate a credential-free, path-bounded synthesis adapter |
| `scripts/evolve/refresh.py` | Run one budgeted discovery→triage→eval→validation refresh |
| `scripts/evolve/daily_worker.py` | Update the rolling `bot/evolution` Draft PR in a disposable clone |
| `scripts/evolve/corpus.py` | Generate or check the canonical corpus inventory and cutoffs |
| `scripts/evolve/daily_worker.py` | Create or update the rolling `bot/evolution` PR from a disposable clone |
| `scripts/evolve/mi355_worker.py` | Run exact-SHA approved evidence tasks on the trusted MI355 node |
| `scripts/backfill_diffs.py` | Fetch real upstream diffs for top-ranked kernel PRs |
| `scripts/enrich_facets.py` | Infer techniques/hardware_features/kernel_types from paths + diffs |
| `scripts/link_prs.py` | Build the bidirectional PR↔wiki bridge |
| `scripts/generate-indices.py` | Regenerate `queries/*.md` from frontmatter |
| `scripts/evaluate_skill.py` | Score held-out retrieval, citations, and architecture safety |
| `scripts/evaluate_answers.py` | Check reference-answer facts and source citations |
| `scripts/verify_provenance.py` | Re-hash artifact bundles and backfill immutable merge SHAs |
| `scripts/validate.py` | Validate pages, evidence, candidates, manifests, claims, and provenance |

CI (`.github/workflows/ci.yml`) gates every push on the validator, the query-tool
smoke tests, scored retrieval/answer evals, provenance, and index freshness.
CI (`.github/workflows/ci.yml`) gates every pull request and push to `main` on
the validator, query-tool smoke tests, scored retrieval/answer evals,
provenance, and index freshness.
`main` is protected by required checks, CODEOWNER review, one human approval,
linear history, and resolved conversations.
linear history, and resolved conversations for non-admin merges. Repository
administrators can bypass these protections.

```bash
python3 -m venv .venv
Expand All @@ -247,10 +271,11 @@ python3 -m venv .venv
### Self-evolution boundary

The system discovers, proposes, evaluates, and collects evidence; it does not
decide its own truth. Every automated change stays in a Draft PR and the bot has
no approval or merge capability. PR/blog text and stored diffs are rendered as
`UNTRUSTED-UPSTREAM-*` data. The public repository is not connected to a
persistent self-hosted runner: [`ops/mi355/`](ops/mi355/) documents the
decide its own truth by design. Automated changes are published to a rolling PR
that is created as Draft; they are not auto-approved or auto-merged. Maintainers
make the acceptance and merge decisions. PR/blog text and stored diffs are
rendered as `UNTRUSTED-UPSTREAM-*` data. The public repository is not connected
to a persistent self-hosted runner: [`ops/mi355/`](ops/mi355/) documents the
node-local, exact-SHA approval and sandbox contract.

### Quality Gates
Expand Down Expand Up @@ -291,7 +316,7 @@ If you use this knowledge base, please cite both:

```bibtex
@misc{rocmkernelwiki2026,
title = {ROCmKernelWiki: An AMD CDNA/RDNA GPU Kernel Optimization Knowledge Base},
title = {ROCmKernelWiki: An AMD CDNA GPU Kernel Optimization Knowledge Base},
author = {ROCmKernelWiki contributors},
year = {2026},
howpublished = {\url{https://github.com/jhinpan/ROCmKernelWiki}},
Expand Down
15 changes: 8 additions & 7 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,11 +5,12 @@ description: Search and apply ROCmKernelWiki when optimizing AMD Instinct kernel

# ROCmKernelWiki — AMD CDNA Kernel Optimization Wiki

> **Corpus dates:** `data/corpus-manifest.yaml` is the generated source of truth
> for the merged-PR cutoff; doc/blog pages carry individual retrieval dates.
> The nod-ai AMDGPU optimization guide is synchronized through commit
> `efa471ae` on **2026-07-20**. Re-run the relevant harvest or source-sync work
> before advancing either boundary.
> **Corpus freshness:** `data/refresh-cutoff.yaml` owns the last complete
> baseline PR-harvest cutoff, `data/evolution-state.yaml` records per-source
> incremental discovery positions, and `data/corpus-manifest.yaml` reports the
> generated inventory. Doc/blog pages carry their own retrieval dates. The
> nod-ai AMDGPU guide snapshot is recorded in
> `sources/blogs/blog-amdgpu-kernel-opt-guide.md`.

Query a structured, cross-referenced knowledge base of AMD GPU kernel
optimization for CDNA3 (gfx942 / MI300) and CDNA4
Expand Down Expand Up @@ -128,8 +129,8 @@ When answering from this KB:
- SHA-256-pinned upstream diffs plus gfx950-first example suites
- **Validator** `scripts/validate.py` — schema, vocabulary, link-integrity (0 errors)

`data/corpus-manifest.yaml` is the generated source of truth for inventory and
freshness; do not duplicate those values in this skill prompt.
`data/corpus-manifest.yaml` is the generated inventory and baseline-cutoff
projection; do not treat it as the owner of rolling or per-page freshness.

## Quality Guarantees

Expand Down
2 changes: 1 addition & 1 deletion data/corpus-manifest.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,4 +12,4 @@ counts:
cutoffs:
merged_prs: '2026-05-30'
harvested_at: '2026-05-30'
source_registry_sha256: a8f03b8b6d6690443821596cd173a3f32172eebdae475668ba489e8a61364caa
source_registry_sha256: 2e2e40c259dd6db885c61badcd42103b47abde1c3623608765295b928d384490
10 changes: 10 additions & 0 deletions data/evolution-state.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,16 @@ sources:
pr: 9351
merge_sha: d04f6751f3df4a716fdd5d2ead6c2918d285964b
captured_at: '2026-07-23'
rocm-triton:
merged_at: '2026-05-30T00:00:00Z'
pr: 0
merge_sha: ''
captured_at: '2026-07-23'
rocm-flash-attention:
merged_at: '2026-05-30T00:00:00Z'
pr: 0
merge_sha: ''
captured_at: '2026-07-23'
vllm-rocm:
merged_at: '2026-07-23T07:17:25Z'
pr: 49523
Expand Down
1 change: 1 addition & 0 deletions data/sources.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -189,6 +189,7 @@ sources:
license: Apache-2.0
source_category: community-note
include_paths:
- "docs/amdgpu_kernel_optimization_guide.md"
- "docs/AMDGPU-Kernel-Optimization/**"
- "docs/empirical-lds/**"

Expand Down
6 changes: 3 additions & 3 deletions ops/evolution/README.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# Daily evolution service

The service creates a disposable clone, resumes the single `bot/evolution`
branch when a Draft PR is already open, runs bounded discovery/triage/evals,
and opens or updates that Draft PR. It never pushes to `main` and cannot approve
or merge its own work.
branch when its PR is already open, runs bounded discovery/triage/evals, and
creates the PR as Draft or updates the existing PR. It never pushes to `main`
or invokes approval or merge; maintainers make those decisions.

Install the service and timer like the MI355 units, copy
`evolution.env.example` to `/etc/rocm-kernel-wiki/evolution.env`, and use a
Expand Down
3 changes: 2 additions & 1 deletion ops/evolution/evolution.env.example
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
# Short-lived GitHub App installation token: contents/pull-requests write,
# metadata read. The bot may update only bot/evolution and Draft PRs.
# metadata read. The service uses it only to update bot/evolution and create or
# edit that branch's rolling PR.
GH_TOKEN=

# Required only until the first refresh PR persists per-source watermarks.
Expand Down
2 changes: 1 addition & 1 deletion scripts/evolve/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
"""Bounded, review-gated evolution tooling for ROCmKernelWiki."""
"""Bounded, maintainer-controlled evolution tooling for ROCmKernelWiki."""

SCHEMA_VERSION = 1
10 changes: 10 additions & 0 deletions scripts/evolve/discover.py
Original file line number Diff line number Diff line change
Expand Up @@ -816,6 +816,16 @@ def run_discovery(
**newest,
"captured_at": capture_date,
}
elif watermark is None and since:
# A successful empty first scan still needs a durable baseline.
# Reusing the caller's lower bound is conservative: the next
# run may rescan the interval, but it cannot skip evidence.
state.setdefault("sources", {})[source_id] = {
"merged_at": since,
"pr": 0,
"merge_sha": "",
"captured_at": capture_date,
}
else:
changes = list(fixture.get(source_id) or []) if fixture_path else []
tree_head = None
Expand Down
6 changes: 3 additions & 3 deletions scripts/evolve/draft_pr.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
#!/usr/bin/env python3
"""Commit a bounded refresh and open or update its rolling Draft PR."""
"""Commit a bounded refresh and open or update its rolling PR."""

from __future__ import annotations

Expand Down Expand Up @@ -49,12 +49,12 @@ def build_pr_body(summary: dict[str, Any]) -> str:

## Review contract

- This PR is intentionally draft and cannot approve or merge itself.
- This PR is created as Draft; maintainers decide acceptance and merge.
- Upstream PR/blog text is untrusted data, not agent instructions.
- Hardware or performance confidence cannot be promoted without a linked,
immutable MI355 evidence bundle.
- Generated indices, provenance checks, retrieval evals, and schema validation
must pass before human review.
must pass before maintainer acceptance.
"""


Expand Down
59 changes: 59 additions & 0 deletions tests/test_evolution.py
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,27 @@ def test_registry_contains_current_first_party_sources():
assert by_repo["ROCm/rocm-blogs"]["kind"] == "github-tree"


def test_amdgpu_guide_registry_matches_canonical_source_path():
from evolve.discover import _tree_candidate
from evolve.registry import load_registry, source_by_id

registry = load_registry(ROOT / "data" / "sources.yaml")
source = source_by_id(registry, "amdgpu-optimization-guide")
candidate = _tree_candidate(
{
"filename": "docs/amdgpu_kernel_optimization_guide.md",
"status": "modified",
"sha": "a" * 40,
"commit": "b" * 40,
},
source,
"2026-07-23",
)

assert candidate["decision"] == "defer"
assert candidate["relevance_reason"] == "allowlisted source path changed"


def test_unknown_architecture_is_quarantined_not_defaulted_to_gfx942():
from evolve.discover import classify_pr
from _scope import is_active
Expand Down Expand Up @@ -301,6 +322,44 @@ def test_fixture_discovery_writes_run_ledger_and_state():
assert second_ledger["counts"] == {}


def test_empty_initial_pr_scan_persists_safe_since_watermark():
from evolve.discover import run_discovery

with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
(root / "data").mkdir()
(root / "data" / "sources.yaml").write_text(
yaml.safe_dump({"schema_version": 1, "sources": [_source()]}),
encoding="utf-8",
)
(root / "data" / "evolution-schemas.yaml").write_text(
(ROOT / "data" / "evolution-schemas.yaml").read_text(encoding="utf-8"),
encoding="utf-8",
)
fixture = root / "fixture.json"
fixture.write_text(json.dumps({"example": []}), encoding="utf-8")

run_discovery(
root=root,
source_ids=["example"],
fixture_path=fixture,
captured_at="2026-07-22",
run_id="20260722T000200Z",
since="2026-05-30T00:00:00Z",
dry_run=False,
)

state = yaml.safe_load(
(root / "data" / "evolution-state.yaml").read_text(encoding="utf-8")
)
assert state["sources"]["example"] == {
"merged_at": "2026-05-30T00:00:00Z",
"pr": 0,
"merge_sha": "",
"captured_at": "2026-07-22",
}


def test_corpus_manifest_is_generated_from_the_checkout():
from evolve.corpus import build_manifest

Expand Down
Loading