Repository navigation
feat(framework): a skill that reports cost per task and per step #629
Description
Activity
- added a parent issue
on Aug 14, 2026 - changed the title
[-]feat(cli): read one number per task and per step[/-][+]feat(framework): a skill that reports cost per task and per step[/+]on Aug 14, 2026 Simplified by #663 moving into v1, 2026-08-16
The framework now emits its own step boundaries in the same milestone. This issue's join changes shape.
Primary path: read the framework's step boundaries. They are exact, they need no provider flag, and they do not suffer the stickiness of
skill.name.Secondary path, free and worth keeping: where
skill_activatedexists — Claude Code only — compare the two. Agreement is a correctness check that costs nothing; disagreement is a signal worth surfacing rather than hiding.The
event.sequencecarry-forward rule stays documented as the fallback for a session where boundaries are missing, and as the mechanism that will be needed on tools that emit neither.Two acceptance criteria relax: the warning about interleaved skills is no longer required to be approximate, since boundaries are observed rather than inferred; and the by-step block no longer depends on
OTEL_LOG_TOOL_DETAILSbeing set.Two things the reader must not get wrong, measured 2026-08-20
Ran the real chain end to end — receiver up, the three captured OTLP payloads posted, day file read back.
Metrics are delta, so metrics and logs measure the same thing twice
claude_code.cost.usage sum aggregationTemporality=1 monotonic=true claude_code.token.usage sum aggregationTemporality=1 monotonic=true pts=4 claude_code.active_time.total sum aggregationTemporality=1 monotonic=trueaggregationTemporality: 1is DELTA: each export carries the increment since the previous one, every 10 s. So the metric datapoints are the same tokens and the same dollars theapi_requestlog records already carry, sliced by time instead of by request.On the captured session that is visible directly: the request lines sum to $0.1605, the metric lines to $0.0151. Not a contradiction — one is every billed request, the other is one 10-second window. Adding them double counts.
The rule for this report: cost and the four token counters come from
kind: "request"lines and nowhere else.active_time_scomes fromkind: "session"lines and nowhere else, because no log record carries it. Never take the same quantity from both.kind: "session"is one line per datapoint, never mergedSix lines for one session in the capture — cost, then each of the four token counters separately, then active time. They are not merged on the way in, deliberately: joining them would assume an ordering no tool documents. A reader that expects one session line per session will silently read a fifth of the truth.
What the aggregate looks like today
Produced with a throwaway script over the stored JSONL, which is exactly the gap this ticket closes:
modèle appels entrée sortie coût USD claude-sonnet-5 3 114485 165 0.173744 query_source appels coût USD sdk 2 0.121843 agent:builtin:general-purpose 1 0.051901 active_time_s 9.714query_sourceseparating subagent spend from the main loop was not asked for in the body above and is worth keeping in the output: 30% of the cost on that sample.Input changed by #684, and one constraint on the output format
The reader no longer takes its figures from the sink by default. #684 decided that it reads the files the tools already wrote (#685) and turns tokens into money through the price table (#654). The receiver stays as an opt-in path for an exact billed amount.
The constraint that has to be in the output format from the first version: every amount printed carries a provenance marker - computed or billed. They are not the same number and they are not interchangeable. A computed figure is fine for internal statistics and wrong for rebilling a client, and a reader cannot tell them apart from the digits.
This replaces the exact-versus-estimated marker the ticket needed for its per-step block. That one is now simpler than it was: with #663 landed, a step joins its cost by identifier on Claude Code (
prompt_id), so the attribution is exact and only the amount carries provenance, not the attribution.Coverage is not uniform and the report has to say so rather than print zeros. Claude Code, Codex and OpenCode are readable locally. Copilot gives
outputTokensper turn and nothing else - no per-request input figure exists, so no per-step breakdown can be built for it. Cursor writes no token count anywhere, and its export is an Enterprise setting nobody here can enable, so it is uncovered by both routes. A tool that cannot be measured must read as uncovered, never as idle.Three parts of the body above are stale — read this before implementing
Nothing here changes what this ticket is for. It changes where the data comes from, which the body describes in detail and describes wrongly now.
1. The join mechanism was replaced. The whole "How the join actually works" section —
skill_activatedlog events, carry-forward byevent.sequenceontoapi_request— was the OTLP design. #687 replaced it with two named sources, and a record now says which one it used:step_attributionWhere the step came from tool-statedThe tool's own per-message field. Claude Code writes attributionSkill, exact, nothird-partyredaction.journal-intervalA half-open interval between step_startboundaries in the run journal (#663). An inference, and marked as one.unattributedNeither source could say. Not "no step ran" — the two are indistinguishable and the stronger reading is never asserted. The carry-forward is not implemented and should not be.
aidd_docs/product/metrics-contract.mdis the shape to build against.2.
depends_on: #620, #646, #647is stale. #647 was the receiver-based sink, demoted off the critical path by #684. #654, listed as a blocker in the epic's earlier sequencing, is closed — the rates live in the SaaS, so no amount is computed here. An amount is printed only where a tool's own files already carried one, and a tool without one prints tokens and an explicit unknown, never a zero.3. "Why a skill and not a CLI command" predates the commands.
aidd telemetry on|off|receive|readall exist, andreadalready writes exactly what this reads. The computation goes beside it asaidd telemetry report; the skill calls it. A skill holding its own aggregation would compute the same figures a second way.Two things the body asked for that got better answers
"Steps sum to the task total, with a residual bucket." Reconciliation is now three-valued, not one: tool-stated + journal-interval + unattributed = total. Unattributed is printed under that name, never as a residual, because a residual reads as "work outside every step" and nothing measured supports that.
"When two skills interleave, the output says attribution is approximate." That becomes numbers rather than a sentence — the share of the total each of the three strengths accounts for. A reader sees how much of the breakdown is measured and how much is inferred, at a glance, and it is assertable in a test.
What unblocked the task join
aidd_docs/runs/README.mdstates thatfile_writtencarries a repository-relative path and deliberately notask_id, because "task identity is a derivation from the path, and derivations belong to whatever reads the log". This reader is that thing. #649 does not block this ticket — no task identity file is needed.One risk carried into phase 1
The journal hook writes
vendor_idfrompayload.session_id; the Codex rollout reader resolves a session bysession_meta.id. Measured: 124 of 330 local rollouts are resumed sessions where those two values differ. If a Codex hook ever reports the parent's identifier, the join drops those sessions and the report still looks healthy — the exact failure mode this epic exists to prevent. Phase 1 settles it by measurement before anything is built on top.Plan:
aidd_docs/tasks/2026_08/2026_08_21_cost-reporter/Shipped.
aidd telemetry reportanswers what a period or one task cost, broken down by step, model and tool, with each attribution's strength printed as a number rather than as a caveat.Proven live on a real headless Claude Code session, hand-verified against the transcript:
by step of tokens aidd-ui:01-hello 67% 78,188 tokens stated by the tool aidd-ui:01-hello 33% 38,490 tokens from a journal intervalBoth strengths, on one skill, from two independent sources — the tool did not state the skill on the message where it decided to invoke it, and the journal had already opened the interval. The two cover each other, which is why #687 refused to merge the rows.
Three corrections to this ticket's own body, recorded in a comment above: the OTLP carry-forward it described was replaced by #687,
#647in itsdepends_onwas demoted by #684, and its 'why a skill and not a CLI command' argument predates the commands.- added 4 commits that reference this issue
on Aug 22, 2026 Closed as completed; the code reverses one of its acceptance points on purpose
Found by a backlog audit on 2026-08-28.
The box asked for a sentence saying attribution is approximate.
cli/src/domain/models/cost-report.ts:119-121states the opposite decision in its own words: the figures are "printed as three figures rather than as a sentence saying attribution is approximate". That is the three-strength breakdown —tool-stated,journal-interval,unattributed— which is more precise than the sentence the box asked for, not less. The reversal looks right; it is just undocumented here.One other box reads as unmet and is not: nothing in
cli/srcmentionsresidual, because the role is filled byunattributed, which is a value the reader returns rather than an omission. Same thing under a better name.Not reopening. Recording it so the issue stops disagreeing with the code it produced.
- added a commit that references this issue
on Sep 2, 2026
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Projects
- StatusShow more project fieldsDone
As a developer or a tech lead
I want to see what a task cost, broken down by step and by model
So that I can decide where to spend optimisation effort instead of guessing
Acceptance
Expected output
How the join actually works
Measured, and not what the first design assumed.
Per-session totals come from the metrics:
claude_code.token.usageandclaude_code.cost.usageboth carrysession.id, andclaude_code.active_time.totalgives the time.Per-step breakdown cannot come from the metrics.
skill.nameon both counters reads the literal stringthird-partyfor every AIDD skill, because the docs replace third-party plugin skill names, andOTEL_LOG_TOOL_DETAILS=1does not lift it on metrics. The real name appears only on theskill_activatedlog event.So the breakdown is an exact event correlation, not a time window. Measured on a real session, both events carry the same correlation keys:
skill_activatedskill.name, withsession.id,prompt.id,event.sequenceapi_requestinput_tokens,output_tokens,cache_*,cost_usd,model,query_source, withsession.id,prompt.id,event.sequenceThe rule: within a session, order by
event.sequenceand carry the lastskill_activatedforward onto theapi_requestrecords that follow, until the next one.api_requesthas its ownskill.name, and it is redacted tothird-partyexactly like the metrics — with the flag on as well. It is not usable, and the carry-forward is what replaces it. That is not a workaround: it mirrors the provider's sticky attribution instead of fighting it.The same mechanism serves the other four tools later, with one difference: none of them emits a
skill_activatedequivalent, so there the step boundaries must be emitted by the framework itself.Two measured limits the output must respect:
skill.nameis sticky. Once activated it rides the following datapoints, including subagents launched afterwards. Correct for sequential steps, wrong for interleaved ones.active_time.totalcarries no skill attribute. Time is per session only. Any percentage in the by-step block is cost, never time.Why a skill and not a CLI command
The question is asked from inside a session, about work in progress. The plugin already carries the diagnostic; the figure belongs beside it. This also satisfies #297's requirement that at least one skill consume the data and produce something a user acts on, without sending anything anywhere.
Out of scope
Relations