Skip to content

feat(framework): a unit of AIDD work carries its cost #631

Description

@blafourcade

A unit of AIDD work carries its cost: when a task closes, we know what it consumed, where it went, and how much to trust the figure.

Context and Value

Providers can say "this developer burned 4M tokens on Tuesday". None can say "story 428 cost €61, 60% of it in specification". The gap between those two sentences is exactly what the framework knows and they do not: the task, the step, the skill.

A measurement campaign on 13-14 August 2026 established that most of the machinery already exists upstream. Claude Code exports tokens, cost in USD and active time per session. What it cannot know is which task a session belonged to. That link is the whole contribution.

The same campaign moved the risk. The load-bearing assumption — that the session id a hook sees is the one the export carries — is proven. What remains dangerous is elsewhere: no probe worked on the first attempt, and never for a different reason. Every tool gates its hooks differently, and a hook installed without lifting that gate is silent and raises nothing.

A measurement layer that fails quietly is worse than none, because it produces figures that look right.

Boundaries

  • Includes: the run journal, the diagnostic that proves the pipe flows, the reading that turns it into a figure, a readable sink, and the per-tool facts everything rests on.
  • Includes: Claude Code as the first and only tool proven end to end.
  • Includes: the collector. Claude Code supports no file exporter for metrics or logs, so a receiving endpoint is a component of this epic, not an externality.
  • Excludes: the four remaining tools, whose mechanics are identical but whose export configuration differs.
  • Excludes: aggregation per person, per team or per epic, and the upload that would feed it.
  • Excludes: the commit trailer, and linking a delivery folder to its backlog artefact.

Success Evidence

On a real repository, one skill answers what a task cost, broken down by step and by model, and the same skill proves no session was silently lost. The failure signal is an installation that reports itself healthy while producing nothing — the mode this epic exists to make impossible.

Design: aidd_docs/specs/2026_08/2026_08_13-work-tracking-linkage.md
Plan: aidd_docs/plans/2026_08_14-telemetry-v1/

Activity

  1. added theissue type on Aug 14, 2026
  2. moved this from Ideation to Todo in AIDD Roadmapon Aug 14, 2026
  3. changed the title [-]feat(framework): un travail AIDD porte son coût[/-] [+]feat(framework): a unit of AIDD work carries its cost[/+] on Aug 14, 2026
  4. blafourcade commented on Aug 20, 2026

    @blafourcade
    ContributorAuthor

    Sequenced 2026-08-20

    The backlog under this epic had grown to twenty-odd tickets with no stated order, which is how #630/#679 and #676/#682 became duplicate pairs. Here is the order, with the reason each item sits where it does. Anything not listed is deliberately after all of it.

    Now — the chain has no consumer

    # Why here
    #663 In review. Records which step was running, on four tools. Without it #629 has no per-step block, which is the half of the report that no provider can produce.
    #684 Decide before building the reader. Whether the local collector stays at all, or the governor exposes the endpoint and holds the price table. It changes where redaction happens and whether a cost is billed or computed — both of which the report has to state on every figure.
    #629 The reader. Everything is recorded and nobody reads. Blocked on the two above only for its per-step block and its cost provenance; the rest is buildable today.

    Next — the data is wrong for two tools out of four

    # Why here
    #681 The journal never writes on Copilot. One tool in four is silent, and a silent tool looks exactly like an idle one.
    #680 Cursor's turn-end never fires headless, so no step ever closes there.
    #617 The diagnostic. Until it exists, an inert installation and an idle developer produce the same empty report — and #681 and #680 are precisely the failures it would have caught first.

    Then — coverage and honesty

    # Why here
    #653 Halved by #663, which already put four tools in the journal. Needs re-scoping to what remains: the export side for Codex and OpenCode.
    #654 The price table. Promoted if #684 says the cost is computed rather than billed — then it stops being a nice-to-have and becomes load-bearing.
    #676 OpenCode, through its plugin API. The only tool with no hook mechanism at all.
    #630 The commit trailer. Closes the gap where a session spent entirely outside the task flow leaves a cost with no content.
    #658 The FAQ promises no telemetry while we ship it. Blocks a public release, not a technical v1 — but it blocks it absolutely.

    Later — identity, transport, aggregation

    #660 and #661 (anonymous versus named, and resolving one person across tools), #655 and #662 (redaction and transport on the upload path, both folded into #684's decision), #656 (per person, team, epic), #659 (keep the join honest in CI), #657 (version-control status for the journal; its retention and schema-versioning halves are already delivered), #683 (one file per tool in the hook — the cost of the next tool, paid early).

    What a v1 actually is

    #663 committed, #684 decided, #629 built. Tokens, models and time are already stored and sliceable by period — proven end to end on 2026-08-20 by running the receiver against real captures. What is missing is not data, it is a reader. Everything else on this list widens coverage or hardens the figures; none of it unlocks the value.

  5. blafourcade commented on Aug 20, 2026

    @blafourcade
    ContributorAuthor

    Re-sequenced after #684

    The spike changed where the data comes from, so the order changed with it. Replaces the sequence posted earlier today.

    Now

    # Why here
    #663 Done. The journal records which step was running, on four tools. Committed, not merged.
    #684 Done. Local reading first, receiver optional.
    #685 Read what a session cost from the files the tool already wrote. Replaces the receiver on the critical path and removes the process-must-be-running failure mode.
    #654 The price table. Load-bearing now, not optional: none of the local files carries a dollar amount.
    #629 The reader. Blocked on the two above for its figures; its output format must carry a computed-versus-billed marker from the first version.

    Next

    # Why here
    #681 The journal never writes on Copilot. A silent tool looks exactly like an idle one.
    #680 Cursor's turn-end never fires headless, so no step ever closes there.
    #617 The diagnostic. #681 and #680 are precisely what it would have caught first.

    Then

    #630 the commit trailer · #676 OpenCode through its plugin API · #653 to re-scope, now essentially Copilot's export alone · #658 the FAQ, which blocks a public release absolutely though not a technical v1 · #683 one file per tool in the hook.

    Later

    #660 #661 identity · #655 #662 redaction and transport on the upload path · #656 aggregation · #659 the join in CI · #657 version-control status for the journal.

    Two facts to stop rediscovering

    Cursor cannot be measured. No token count in any file it writes, and an export gated behind an Enterprise team setting nobody here can enable. Uncovered by both routes. It belongs in the documentation as a known limit, not in the backlog as pending work.

    Copilot cannot give a per-step breakdown. Its local file carries outputTokens per turn and nothing else; input, cache and reasoning arrive once, at shutdown, for the whole session. Only its OTLP export would close that, and only if the user sets the variable themselves.

    What a v1 is, restated

    #685 + #654 + #629. Tokens, models and time are already stored and sliceable by period. What is missing is not data, it is a reader - and now, the table that turns tokens into money.

  6. blafourcade commented on Aug 20, 2026

    @blafourcade
    ContributorAuthor

    Re-sequenced: the rates move to the SaaS

    #654 is closed and moved to the governor. Once the rates live there, this repository has no pricing code to write — its job is upstream of pricing: emit metrics complete enough to price, and let the service holding the rates do it.

    That deletes a phase from the critical path and adds a smaller, harder one.

    Now

    # Why here
    #687 The metrics contract. A stored record does not name the tool that produced it, and does not carry the skill. Everything downstream — pricing, per-person, per-epic — reads this shape; building on one that cannot name its own tool means every consumer re-derives it, differently.
    #629 The reader. Buildable today for tokens, models and agents; needs #687 for anything per-tool or per-step.

    Next

    #681 the journal never writes on Copilot · #680 Cursor's turn-end never fires headless · #617 the diagnostic that would have caught both first.

    Then

    #662 transport out of band and #655 redaction on that path — now the pair that actually reaches the governor, so they move up · #686 a synthetic transcript message is not a billed request · #630 the commit trailer · #676 OpenCode through its plugin API · #653 to re-scope, essentially Copilot's export alone · #658 the FAQ, which blocks a public release absolutely · #683 one file per tool in the hook.

    Later

    #660 #661 identity · #656 aggregation · #659 the join in CI · #657 version-control status for the journal.

    What a v1 is, restated again

    #687 + #629. Proven end to end on 2026-08-20 against this repository's own session: aidd telemetry read produced 5134 records with tokens, models and per-agent attribution — aidd-dev:executor at 2543 calls against 1611 for the main loop, 98% of consumption in cache reads. The data is there. What is missing is a contract a consumer can read, and something that reads it.

  7. blafourcade commented on Aug 21, 2026

    @blafourcade
    ContributorAuthor

    A plan to a clean v1

    aidd_docs/tasks/2026_08/2026_08_21_clean-v1/

    Where this starts

    Delivered and unmerged: #663, #684, #685, #687, #629, #689, #690, #691, #692. The chain works end to end on Claude Code, proven on live headless sessions rather than on fixtures — a real session measured itself, and the figures were recomputed by hand from the transcript.

    What that does not mean, and the plan exists for the gap:

    Reads as Actually
    tested 2614 CLI tests, 177 hook tests, two live probes — and nobody has run it for a week
    works on Claude Code proven live; Codex proven on captured files; Copilot and Cursor record nothing
    shipped on a branch, behind a FAQ that promises the opposite

    Four milestones, in order

    0 — What exists reaches someone. Nothing new is built. Merge the branch, finish #658, write the delivery page. Nine tickets on one branch is the largest risk here and it grows every hour; every later milestone touches the same files. Half a day, and deliverable on its own.

    1 — Every declared tool records. #681 Copilot silent · #680 Cursor turn-end · #693 the worktree · #676 OpenCode through its plugin API. A tool that records nothing shows a user a zero, which is the one thing this layer exists never to show — worse than a missing breakdown, and cheaper to fix. Two to three days.

    2 — It cannot lie quietly. #617 the diagnostic · scale · a live multi-step flow · #686. Written before milestone 1, #617 would spend its time reporting the two failures we already know about; after it, everything it reports is news. Two days, plus a real session.

    3 — The figures leave the machine. #662 transport · #655 redaction in flight · #660 anonymity · #661 identity · #656 aggregation. Last because nothing above needs it, and because it is the only milestone that can leak — sending figures nobody has proven trustworthy, to a place they cannot be recalled from. Three to four days.

    Where a v1 actually is

    Milestone 0 is a v1 for one tool. Honest, installable, and narrow — Claude Code measured, everything else named as uncovered.

    Milestone 1 is the v1 worth announcing. Four hosts recording, coverage stated per tool, no silent zero.

    Milestone 2 is what makes it trustworthy without the reader having built it. Milestone 3 is the product beyond one laptop.

    What is deliberately not in any milestone

    Pricing. The rates live in the governor, closed as #654. This repository's job ends at emitting figures complete enough to price, and every milestone respects that.

  8. blafourcade commented on Aug 21, 2026

    @blafourcade
    ContributorAuthor

    Milestone 0 landed

    Five commits on claude/aidd-telemetry-layer-e403uf, thirteen in total, nothing outstanding in the tree.

    f4971c32 refactor(framework): what the plugin ships can be read
    8d76560c docs(framework): what measurement does, what it cannot, and where it goes next
    821b2dcf fix(framework): a file written through the shell still reaches its task
    4253a74c feat(framework): the plugin measures on its own, with no CLI installed
    3e5bfe3d feat(cli): what a period cost, and what each figure is worth
    

    Closed: #629, #689, #690, #691, #692, #658.

    Grouped by what they deliver rather than one per ticket — #629, #689 and #690 touch the same lines of the same files, and splitting them would have produced commits that do not compile. The messages say so rather than simulating a precision that was not there.

    What the plugin ships is now readable. Both scripts were minified; neither had to be. The switch is hand-written plain CommonJS like the hooks — it is the file someone reads before allowing anything to be recorded, and sixty commented lines answer that better than an artefact. The reporter stays generated but unminified, keeping real function names and a // src/... marker above every block. Forty percent more bytes on a file that is copied rather than downloaded, in exchange for one that can be audited.

    Not pushed. That step is the user's.

    The rest of the road

    Milestone 1 — every declared tool records. #681 · #680 · #693 · #695 · #676

    #695 is new and split out of #693: session_start does not name the worktree, so two sessions from two worktrees are indistinguishable. It is worth doing whichever way #693 decides, and it is one field.

    Milestone 2 — it cannot lie quietly. #694 · #686

    #694 is new and holds what #617 was for, restated against what now exists: four induced failures each named as itself, a period that has met a hundred sessions, a timed budget for the turn-end walk, and a live multi-step flow. Everything shipped so far has met three sessions at most — "it scales" is currently a hope with no number behind it.

    Milestone 3 — the figures leave the machine. #662 · #655 · #660 · #661 · #656

    The plan, with what each milestone is done when: aidd_docs/tasks/2026_08/2026_08_21_clean-v1/

  9. blafourcade commented on Aug 21, 2026

    @blafourcade
    ContributorAuthor

    Every boundary, against coverage as it actually stands

    The epic's Success Evidence: "On a real repository, one skill answers what a task cost, broken down by step and by model, and the same skill proves no session was silently lost." Both halves now exist and have been run against the same real session.

    Boundary as written Where it stands
    The run journal Met. Writes on Claude Code and on Codex, both proven by a live session rather than by a test.
    The diagnostic that proves the pipe flows Met. skills/02-check, five claims, each failure induced and proved to be the only claim that turns red.
    The reading that turns it into a figure Met. A real three-skill chain reconciles to its total, integer-exact on all five token fields.
    A readable sink Met, and now measured: a hundred sessions over a year of day files answer in under 80ms.
    The per-tool facts Met, and they moved: see below.
    Claude Code end to end Met, and exceeded — Codex journals end to end too.
    The collector Not exercised in this work. It exists and is not part of what was proven here.
    The four remaining tools — excluded Partly overtaken, and not in the direction the exclusion assumed. Stated below rather than quietly claimed.
    Aggregation per person, team, epic — excluded Unchanged.
    The commit trailer and the delivery-folder link — excluded Unchanged.

    The exclusion has moved, and saying so matters more than closing cleanly

    The epic excluded four tools on the reasoning that their mechanics are identical and only their export configuration differs. Measurement since says otherwise, tool by tool:

    None of that blocks this epic. It does mean closing it must not be read as claiming those four work.

    What this cost to prove, since that is the point

    $2.10 for the flow itself. $1.27 for a first attempt that reported success at every step and produced nothing — #703, the same silent-health failure this epic exists to remove, one layer below it.

    That the layer correctly reported nothing, and that a person could tell the difference, is the whole result.

    Open, and named

    #699 Codex hook trust · #700 the model setup writes · #701 Copilot steps · #702 the tool declarations exist twice · #703 headless plugin loading. None is a prerequisite for what this epic claimed; all five were found by proving it.

  10. blafourcade commented on Sep 2, 2026

    @blafourcade
    ContributorAuthor

    Delivered by #706, squash-merged into next as 627408f.

  11. blafourcade commented on Sep 11, 2026

    @blafourcade
    ContributorAuthor

    Closed for the outcome, with one wording correction: cost reporting and health diagnosis ship as two skills, not the “same skill” described in Success Evidence.

    01-cost reports what work consumed. 02-check establishes whether the measurement chain is recording. The separation is intentional: one command should not pretend to answer both questions.

    The remaining scale evidence is tracked in #694.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    High

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions