Skip to content

feat(framework): a skill that reports cost per task and per step #629

Description

@blafourcade

As a developer or a tech lead
I want to see what a task cost, broken down by step and by model
So that I can decide where to spend optimisation effort instead of guessing

Acceptance

  • Steps sum to the task total, with a residual bucket so the sum reconciles exactly.
  • A task with no session prints zeros and exits 0, never an error.
  • A run whose identifier joins nothing is counted as unattributed and named, never dropped.
  • Telemetry whose session id matches no run file is also named — the mirror case, and the one the epic exists to catch.
  • A session that emitted a run file but no datapoints is distinguished from a session that was never journaled.
  • The unattached share over the period is printed, and the period is a named flag with a documented default.
  • When two skills interleave, the output says step attribution is approximate rather than presenting it as exact.
  • No prompt, code or diff content appears anywhere in the output.

Expected output

task 2026_08_14_telemetry-v1

  sessions            6
  active time         47 min          (per session; not attributable to steps)
  tokens              310,400         34% cache
  cost                $4.20

  by step
    aidd-dev:02-implement    61%   $2.56
    aidd-dev:05-review       19%   $0.80
    aidd-dev:01-plan         12%   $0.50
    residual                  8%   $0.34

  by model
    claude-opus-5            78% of cost for 31% of calls

  unattached          39% of the period

How the join actually works

Measured, and not what the first design assumed.

Per-session totals come from the metrics: claude_code.token.usage and claude_code.cost.usage both carry session.id, and claude_code.active_time.total gives the time.

Per-step breakdown cannot come from the metrics. skill.name on both counters reads the literal string third-party for every AIDD skill, because the docs replace third-party plugin skill names, and OTEL_LOG_TOOL_DETAILS=1 does not lift it on metrics. The real name appears only on the skill_activated log event.

So the breakdown is an exact event correlation, not a time window. Measured on a real session, both events carry the same correlation keys:

Event Carries
skill_activated the real skill.name, with session.id, prompt.id, event.sequence
api_request input_tokens, output_tokens, cache_*, cost_usd, model, query_source, with session.id, prompt.id, event.sequence

The rule: within a session, order by event.sequence and carry the last skill_activated forward onto the api_request records that follow, until the next one.

api_request has its own skill.name, and it is redacted to third-party exactly like the metrics — with the flag on as well. It is not usable, and the carry-forward is what replaces it. That is not a workaround: it mirrors the provider's sticky attribution instead of fighting it.

The same mechanism serves the other four tools later, with one difference: none of them emits a skill_activated equivalent, so there the step boundaries must be emitted by the framework itself.

Two measured limits the output must respect:

  • skill.name is sticky. Once activated it rides the following datapoints, including subagents launched afterwards. Correct for sequential steps, wrong for interleaved ones.
  • active_time.total carries no skill attribute. Time is per session only. Any percentage in the by-step block is cost, never time.

Why a skill and not a CLI command

The question is asked from inside a session, about work in progress. The plugin already carries the diagnostic; the figure belongs beside it. This also satisfies #297's requirement that at least one skill consume the data and produce something a user acts on, without sending anything anywhere.

Out of scope

  • The four other tools, and the price table the two that export no amount will need.
  • Any aggregation by person, team or epic.
  • Currency conversion. Costs print in USD, as exported.

Relations

Field Value
parent #631
depends_on #620, #646, #647
related #617, #632

Activity

  1. moved this from Ideation to Todo in AIDD Roadmapon Aug 14, 2026
  2. changed the title [-]feat(cli): read one number per task and per step[/-] [+]feat(framework): a skill that reports cost per task and per step[/+] on Aug 14, 2026
  3. blafourcade commented on Aug 16, 2026

    @blafourcade
    ContributorAuthor

    Simplified by #663 moving into v1, 2026-08-16

    The framework now emits its own step boundaries in the same milestone. This issue's join changes shape.

    Primary path: read the framework's step boundaries. They are exact, they need no provider flag, and they do not suffer the stickiness of skill.name.

    Secondary path, free and worth keeping: where skill_activated exists — Claude Code only — compare the two. Agreement is a correctness check that costs nothing; disagreement is a signal worth surfacing rather than hiding.

    The event.sequence carry-forward rule stays documented as the fallback for a session where boundaries are missing, and as the mechanism that will be needed on tools that emit neither.

    Two acceptance criteria relax: the warning about interleaved skills is no longer required to be approximate, since boundaries are observed rather than inferred; and the by-step block no longer depends on OTEL_LOG_TOOL_DETAILS being set.

  4. blafourcade commented on Aug 20, 2026

    @blafourcade
    ContributorAuthor

    Two things the reader must not get wrong, measured 2026-08-20

    Ran the real chain end to end — receiver up, the three captured OTLP payloads posted, day file read back.

    Metrics are delta, so metrics and logs measure the same thing twice

    claude_code.cost.usage         sum  aggregationTemporality=1  monotonic=true
    claude_code.token.usage        sum  aggregationTemporality=1  monotonic=true  pts=4
    claude_code.active_time.total  sum  aggregationTemporality=1  monotonic=true
    

    aggregationTemporality: 1 is DELTA: each export carries the increment since the previous one, every 10 s. So the metric datapoints are the same tokens and the same dollars the api_request log records already carry, sliced by time instead of by request.

    On the captured session that is visible directly: the request lines sum to $0.1605, the metric lines to $0.0151. Not a contradiction — one is every billed request, the other is one 10-second window. Adding them double counts.

    The rule for this report: cost and the four token counters come from kind: "request" lines and nowhere else. active_time_s comes from kind: "session" lines and nowhere else, because no log record carries it. Never take the same quantity from both.

    kind: "session" is one line per datapoint, never merged

    Six lines for one session in the capture — cost, then each of the four token counters separately, then active time. They are not merged on the way in, deliberately: joining them would assume an ordering no tool documents. A reader that expects one session line per session will silently read a fifth of the truth.

    What the aggregate looks like today

    Produced with a throwaway script over the stored JSONL, which is exactly the gap this ticket closes:

    modèle                        appels     entrée   sortie   coût USD
    claude-sonnet-5                    3     114485      165   0.173744
    
    query_source                  appels   coût USD
    sdk                                2   0.121843
    agent:builtin:general-purpose      1   0.051901
    
    active_time_s                                       9.714
    

    query_source separating subagent spend from the main loop was not asked for in the body above and is worth keeping in the output: 30% of the cost on that sample.

  5. blafourcade commented on Aug 20, 2026

    @blafourcade
    ContributorAuthor

    Input changed by #684, and one constraint on the output format

    The reader no longer takes its figures from the sink by default. #684 decided that it reads the files the tools already wrote (#685) and turns tokens into money through the price table (#654). The receiver stays as an opt-in path for an exact billed amount.

    The constraint that has to be in the output format from the first version: every amount printed carries a provenance marker - computed or billed. They are not the same number and they are not interchangeable. A computed figure is fine for internal statistics and wrong for rebilling a client, and a reader cannot tell them apart from the digits.

    This replaces the exact-versus-estimated marker the ticket needed for its per-step block. That one is now simpler than it was: with #663 landed, a step joins its cost by identifier on Claude Code (prompt_id), so the attribution is exact and only the amount carries provenance, not the attribution.

    Coverage is not uniform and the report has to say so rather than print zeros. Claude Code, Codex and OpenCode are readable locally. Copilot gives outputTokens per turn and nothing else - no per-request input figure exists, so no per-step breakdown can be built for it. Cursor writes no token count anywhere, and its export is an Enterprise setting nobody here can enable, so it is uncovered by both routes. A tool that cannot be measured must read as uncovered, never as idle.

  6. blafourcade commented on Aug 21, 2026

    @blafourcade
    ContributorAuthor

    Three parts of the body above are stale — read this before implementing

    Nothing here changes what this ticket is for. It changes where the data comes from, which the body describes in detail and describes wrongly now.

    1. The join mechanism was replaced. The whole "How the join actually works" section — skill_activated log events, carry-forward by event.sequence onto api_request — was the OTLP design. #687 replaced it with two named sources, and a record now says which one it used:

    step_attribution Where the step came from
    tool-stated The tool's own per-message field. Claude Code writes attributionSkill, exact, no third-party redaction.
    journal-interval A half-open interval between step_start boundaries in the run journal (#663). An inference, and marked as one.
    unattributed Neither source could say. Not "no step ran" — the two are indistinguishable and the stronger reading is never asserted.

    The carry-forward is not implemented and should not be. aidd_docs/product/metrics-contract.md is the shape to build against.

    2. depends_on: #620, #646, #647 is stale. #647 was the receiver-based sink, demoted off the critical path by #684. #654, listed as a blocker in the epic's earlier sequencing, is closed — the rates live in the SaaS, so no amount is computed here. An amount is printed only where a tool's own files already carried one, and a tool without one prints tokens and an explicit unknown, never a zero.

    3. "Why a skill and not a CLI command" predates the commands. aidd telemetry on|off|receive|read all exist, and read already writes exactly what this reads. The computation goes beside it as aidd telemetry report; the skill calls it. A skill holding its own aggregation would compute the same figures a second way.

    Two things the body asked for that got better answers

    "Steps sum to the task total, with a residual bucket." Reconciliation is now three-valued, not one: tool-stated + journal-interval + unattributed = total. Unattributed is printed under that name, never as a residual, because a residual reads as "work outside every step" and nothing measured supports that.

    "When two skills interleave, the output says attribution is approximate." That becomes numbers rather than a sentence — the share of the total each of the three strengths accounts for. A reader sees how much of the breakdown is measured and how much is inferred, at a glance, and it is assertable in a test.

    What unblocked the task join

    aidd_docs/runs/README.md states that file_written carries a repository-relative path and deliberately no task_id, because "task identity is a derivation from the path, and derivations belong to whatever reads the log". This reader is that thing. #649 does not block this ticket — no task identity file is needed.

    One risk carried into phase 1

    The journal hook writes vendor_id from payload.session_id; the Codex rollout reader resolves a session by session_meta.id. Measured: 124 of 330 local rollouts are resumed sessions where those two values differ. If a Codex hook ever reports the parent's identifier, the join drops those sessions and the report still looks healthy — the exact failure mode this epic exists to prevent. Phase 1 settles it by measurement before anything is built on top.

    Plan: aidd_docs/tasks/2026_08/2026_08_21_cost-reporter/

  7. blafourcade commented on Aug 21, 2026

    @blafourcade
    ContributorAuthor

    Shipped. aidd telemetry report answers what a period or one task cost, broken down by step, model and tool, with each attribution's strength printed as a number rather than as a caveat.

    Proven live on a real headless Claude Code session, hand-verified against the transcript:

    by step    of tokens
      aidd-ui:01-hello           67%   78,188 tokens    stated by the tool
      aidd-ui:01-hello           33%   38,490 tokens    from a journal interval
    

    Both strengths, on one skill, from two independent sources — the tool did not state the skill on the message where it decided to invoke it, and the journal had already opened the interval. The two cover each other, which is why #687 refused to merge the rows.

    Three corrections to this ticket's own body, recorded in a comment above: the OTLP carry-forward it described was replaced by #687, #647 in its depends_on was demoted by #684, and its 'why a skill and not a CLI command' argument predates the commands.

  8. blafourcade commented on Aug 28, 2026

    @blafourcade
    ContributorAuthor

    Closed as completed; the code reverses one of its acceptance points on purpose

    Found by a backlog audit on 2026-08-28.

    The box asked for a sentence saying attribution is approximate. cli/src/domain/models/cost-report.ts:119-121 states the opposite decision in its own words: the figures are "printed as three figures rather than as a sentence saying attribution is approximate". That is the three-strength breakdown — tool-stated, journal-interval, unattributed — which is more precise than the sentence the box asked for, not less. The reversal looks right; it is just undocumented here.

    One other box reads as unmet and is not: nothing in cli/src mentions residual, because the role is filled by unattributed, which is a value the reader returns rather than an omission. Same thing under a better name.

    Not reopening. Recording it so the issue stops disagreeing with the code it produced.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Fields

    Priority

    High

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions