Skip to content
Open
24 changes: 9 additions & 15 deletions plugins/aidd-refine/skills/05-improve/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,29 +8,23 @@ argument-hint: conversation | export
```mermaid
flowchart LR
start([conversation ID or export]) --> read-conversation
read-conversation -->|complete| scopes{two or more scopes?}
read-conversation -->|complete| recommend --> target-edits
read-conversation -->|unavailable| unavailable([stop])
scopes -->|no| recommend-local[recommend locally] --> target-edits
scopes -->|yes| isolation{isolated artifact context?}
isolation -->|yes| recommend-parallel[recommend in parallel] --> target-edits
isolation -->|no| recommend-local
target-edits --> question([ask next intent]) --> stop([stop])
```

## Actions

Run the flow above. Read only the next action file.
Run all three actions in order without confirmation. Read only the next file in `actions/`.

| Action | Does |
| Order | Action |
| --- | --- |
| read-conversation | freeze complete evidence and measure visible cost |
| recommend | analyze relevant scopes and merge grounded findings |
| target-edits | render minimal edits and an executable prompt |
| 1 | `01-read-conversation.md` |
| 2 | `02-recommend.md` |
| 3 | `03-target-edits.md` |

## Transversal rules

- After resolving the source, run all three actions without pausing for confirmation.
- Analyze only a complete conversation and the exact skills or documents it names.
- Exclude every current or previous `improve` invocation and report from the analysis evidence.
- Never invent time, tokens, source coverage, or document status.
- Do not modify, stage, or persist any project file.
- Locate conversation data through [conversation sources](assets/conversation-sources.md); access only the selected conversation and its relevant sources.
- Use recorded evidence only; mark unavailable measurements as such.
- Keep project files read-only; write only a unique temporary report.
Original file line number Diff line number Diff line change
@@ -1,35 +1,6 @@
# 01 - Read conversation

Load complete conversation evidence once and profile its visible cost.

## Input

A current conversation, an exact conversation ID, or a complete transcript export.

## Output

An evidence boundary, a `## Timing` table with `Activity | Observed time | Share | Evidence`, a `## Usage` table with `Metric | Value | Evidence`, and a scope index.

## Process

1. **Resolve.** Choose the host-specific transcript source from [conversation sources](../assets/conversation-sources.md).
- Stop when no complete transcript can be resolved for the exact conversation.
2. **Bound.** Freeze the evidence at the message before the current invocation, or at the supplied export boundary.
3. **Read.** Load every in-scope message, tool call, tool result, and timestamp from that transcript.
4. **Index.** Record relevant turns and invoked or named skills and knowledge files for downstream analysis.
5. **Measure.** Calculate visible time and usage only from timestamps, elapsed records, or host usage data.
- Group explicit tool activity as `research and diagnosis`, `implementation`, or `validation`; keep mixed or unknown time `unattributed`.
- Include tokens, requests, and monetary cost only when the source exposes them.
- Never label unattributed wait as reasoning time.
6. **Render.** Mark a missing metric `unavailable` and order known activity times from longest to shortest.

## Test

| Case | Pass |
| --- | --- |
| A conversation is analyzed | its source resolves to one exact complete transcript |
| The same conversation is analyzed again | its evidence excludes every `improve` invocation and report |
| A timing metric is shown | its evidence identifies visible timestamp or elapsed records |
| A metric is unavailable | the report does not estimate or call it reasoning time |
| A usage metric is shown | its evidence identifies the host record that exposes it |
| A scope is indexed | it identifies exact turns or artifact paths, not a generated summary |
1. **Read.** Load the complete selected conversation before this invocation, or through its export boundary. Exclude previous Improve runs and reports. Stop if incomplete.
2. **Index.** Retain chronological prompts, tool calls/results, timestamps, and exposed usage, including failures, retries, agents, and linked subcalls. List invoked skills and tools.
3. **Sources.** Read their relevant maintained files, applicable `AGENTS.md`, and referenced memory. Reuse valid reads; never target install copies.
4. **Measure.** Report elapsed seconds, tokens, call counts, and exposed cost. Pair each user prompt with its recorded response end; leave unpaired durations unavailable. Distinguish elapsed gaps, tool durations, and background-process lifetime; never sum overlapping intervals.
49 changes: 3 additions & 46 deletions plugins/aidd-refine/skills/05-improve/actions/02-recommend.md
Original file line number Diff line number Diff line change
@@ -1,50 +1,7 @@
# 02 - Recommend

Analyze each relevant scope with the smallest grounded change.
Reflect internally:

## Input
> Based on the conversation, what shorter, equally reliable path could have achieved the same result, and what minimal reusable change would reduce time, tokens, or cost next run?

The evidence boundary, timing and usage tables, complete conversation, and scope index.

## Output

A `## Recommendations` table with `ID | Question | Type | Diagnostic | Evidence | Smallest change | Target | Saving`.

## Process

1. **Scope.** Select `behavior`, plus `skill` and `knowledge` only when the scope index names relevant artifacts.
2. **Dispatch.** Analyze locally when one scope exists; otherwise dispatch one read-only analyst per scope in parallel.
- Prefer a lightweight available model and low reasoning effort when the host supports per-agent overrides; otherwise inherit the run defaults.
- Give the behavior analyst the complete frozen transcript; give artifact analysts the same boundary, indexed turns, and exact artifact paths.
- Request isolated or minimal context for artifact analysts when supported; otherwise analyze every scope locally instead of duplicating the transcript.
- Do not dispatch an analyst with no relevant evidence.
- Require table rows only and forbid file writes.
3. **Question.** Make each analyst answer every prompt for its scope.
- How could the next run be faster or better?
- What information should be removed or clarified?
- Where should the change live?
- How could it save time or tokens?
- What work was counterproductive?
4. **Verify.** Read a named skill or knowledge file before assessing its information.
5. **Assess.** Label relevant information `obsolete`, `over-specific-or-time-bound`, `duplicate`, `inconsistent`, `counterproductive`, or `correct`.
- Use `correct` when no evidence supports another label, and never render it as a recommendation.
6. **Merge.** Deduplicate findings across scopes and verify only their cited evidence against the frozen source.
7. **Render.** Order by question then `behavior`, `skill`, `knowledge`, and describe each change with the fewest unambiguous words.
- Use `skill`, `behavior`, `knowledge`, or `tooling` as the target type.
- State `time`, `tokens`, `both`, or `unknown` as its saving.
- Render `no change` when a scope has no evidence-backed recommendation.

## Test

| Case | Pass |
| --- | --- |
| A question is shown | it is answered from conversation evidence |
| One relevant scope exists | no parallel analyst is dispatched |
| Several relevant scopes exist | their analysts run in parallel and return the same columns |
| An artifact analyst cannot receive isolated context | every scope is analyzed locally instead of duplicating the transcript |
| A type is shown | it is `skill`, `behavior`, `knowledge`, or `tooling` |
| Information is assessed | it has one allowed label backed by evidence |
| Information is correct | its scope says `no change` and no recommendation is rendered |
| A named target is assessed | the target file was read before the verdict |
| A saving is shown | it is categorical and never an invented amount |
| The same evidence is analyzed again | recommendations keep the same order and do not cite an earlier `improve` report |
Prefer removal or consolidation over adding instructions. Return only useful, verified edits: a clear title, relative source path, exact before/after, and a short execution prompt. No useful edit means no recommendation.
Original file line number Diff line number Diff line change
@@ -1,35 +1,5 @@
# 03 - Target edits

Map recommendations to minimal edits and render the report.
Fill [the report template](../assets/report-template.html) with the conversation, measurements, and edits. Replace its fictional example, escape inserted text, and copy `report.css` and `report.js` alongside it. Localize the UI, not source excerpts.

## Input

The timing, usage, and recommendations tables.

## Output

An HTML report in a temporary directory with local `report.css` and `report.js`, plus its path and the next-intent question.

## Process

1. **Target.** Map each file-targeted recommendation to a real project path.
- Omit behavior-only recommendations from this table.
2. **Summarize.** Build a `Fichier | + Ajout | − Retrait ou clarification` table.
- Consolidate repeated targets and render each diagnostic's smallest change.
- Render one `Aucun fichier recommandé | — | —` row when no recommendation needs a file edit.
3. **Render.** Fill [the report template](../assets/report-template.html) with only measured values and grounded findings, then copy its local CSS and JavaScript beside it.
- Remove every sample value and sample finding from the produced report.
- HTML-escape every injected value; allow only template-owned markup and local asset references.
- Write only to a unique temporary directory, never the project.
4. **Return.** Provide the report path and end with this exact question: `Quel changement d’intention général, même minime, appliquons-nous au prochain run pour rendre notre amélioration cumulative et mesurable ?`

## Test

| Case | Pass |
| --- | --- |
| A file is shown | it is the target of an earlier recommendation |
| No file needs editing | the empty-table row is shown |
| The report is rendered | its HTML, CSS, and JavaScript resolve locally with no network dependency |
| Conversation evidence is rendered | it is escaped as text and cannot create markup, scripts, or URLs |
| A value is unavailable | the report says `unavailable` instead of showing sample data |
| The run closes | the returned chat response ends with the exact next-intent question |
Return the HTML path. Ask which small general intent change to try next run.
Original file line number Diff line number Diff line change
Expand Up @@ -7,11 +7,3 @@ Use this static routing map to load one complete conversation. Do not fill or co
| Codex | host thread reader with the exact thread ID, then TUI `/export` Markdown | `CODEX_HOME/history.jsonl`, normally `~/.codex/history.jsonl` | TUI `/status` estimated thread credits or cost; tool elapsed records when exposed |
| Claude Code | `/export` text, or the documented script interface for an exact session ID | `~/.claude/projects/{project}/{session-id}.jsonl` | transcript timestamps and tool records when exposed |
| OpenCode | `opencode export {session-id} --sanitize` JSON | `~/.local/share/opencode/project/{project-slug}/storage/` | `opencode stats`; `~/.local/share/opencode/log/` |

## Rules

- Prefer the complete export or host reader over direct storage parsing.
- Use direct storage only after matching one exact session ID.
- Search shared history or logs only for the exact session ID and ignore unrelated records.
- Never enumerate unrelated sessions or inspect configuration or authentication files.
- Treat private reasoning duration as unavailable unless the host exposes it directly.
Loading
Loading