Skip to content

Port upstream 0.68.0: attribute Nous OpenCodex ledger rows to Nous Portal spend - #682

Closed
Finesssee wants to merge 3 commits into
port/upstream-0.68.0from
port/micro-0.68.0-nous-opencodex-ledger
Closed

Finesssee wants to merge 3 commits into
port/upstream-0.68.0from
port/micro-0.68.0-nous-opencodex-ledger

Conversation

@Finesssee

@Finesssee Finesssee commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

With OpenCodex import on, Usage & Spend now attributes OpenCodex ledger rows from
~/.opencodex/usage.jsonl (or $OPENCODEX_HOME/usage.jsonl) to the Nous Portal row when their
provider is nous. This holds even when the model id names another vendor. The row's source
label is "OpenCodex".

Pricing

  • A Nous row is priced in this order:

    1. A custom pricing override keyed nous/<model>, or the bare <model> key.
    2. An exact Nous models.dev entry (provider nous, model id kept whole, for example
      anthropic/claude-sonnet-4.6).

    Model prefixes never pull in another vendor's rates. A row without a price stays unpriced and
    keeps its tokens.

  • A row is priced only when it reports both input and output tokens.

  • A consumed cache lane needs its own rate:

    • A catalog entry without a cache-read or cache-write rate leaves rows that used that lane
      unpriced.
    • Cache writes use the catalog's cache_write rate.
  • Catalog lookups match exactly. A dated or versioned alias doesn't borrow the base entry.

  • A repeated nous/ prefix (model nous/z-ai/glm-5) is stripped before the lookup.

  • Legacy OpenAI-transport rows (provider: "openai", model nous/<id>) are routed to Nous but
    priced only by a custom key.

  • Native Codex sessions never reach the Nous catalog.

Ledger rows

  • usageStatus: "unreported" rows keep their tokens and no dollars.
  • Extractor _meta.* fields and the conversationID spelling are ignored. A row without
    conversationId counts as its own session, falling back to requestId.
  • A Nous custom override bills input, cache reads and cache writes as separate lanes, which
    reproduces upstream's fixture costs.
  • Portal credit meters are unchanged, and metered cost stays empty.

Upstream reference

Ported / Deferred

Ported:

  • route_provider maps nous to the Nous subscription. The new spend_contract/opencodex/nous.rs
    prices Nous rows. nous::custom_rates is the single resolver for both the cost and the model's
    custom-pricing flag.
  • ModelsDevPricingSnapshot::lookup_exact provides the exact catalog lookup.
  • CostUsagePricing::models_dev_cost_usd follows upstream codexCostUSD(pricing:):
    • Cache writes use the catalog cache_write rate.
    • In the long tier, a missing cache rate falls back to the long input rate.
    • The Codex routed-model path (deepseek/<model> and similar) shares this function.
  • The parser reads usage.totalTokens when a row has no top-level totalTokens. The OpenCodex
    parse cache schema goes from 2 to 3, so cached rows are parsed again.
  • The conversation count falls back to requestId, for all OpenCodex rows.
  • The Usage & Spend command adds nous to the routed OpenCodex providers, with source label
    "OpenCodex".
  • docs/PROVIDERS.md gets a note.

Deferred, or differs from upstream:

  • Custom override token semantics. A Nous custom override bills input, cache reads and cache
    writes as separate lanes, which reproduces upstream's fan-out test values (0.044862,
    0.059058, 5.503258). Upstream's application-overlay path subtracts cache reads from input
    instead.
  • Key order. Windows tries nous/<model> before the bare key. Upstream tries the bare key
    first. This matters only when both keys exist with different rates.
  • No models.dev refresh before pricing. Upstream runs refreshPricingIfNeeded first. This,
    together with the general-path pricing gaps (unwrap_or(0) on non-Nous rows, cache writes on
    the routed path), is the stacked follow-up port/micro-0.60.4-opencodex-unpriced.
  • rust/src/cli/cost.rs. Only codex, claude, pi and opencodego build a spend contract
    there, so there's nothing to add.

Validation

Windows, toolchain 1.98.0, run at 1539d22f:

  • cargo fmt --all --check: pass.

  • cargo clippy --workspace --all-targets -- -D warnings: pass.

  • Focused tests: 114 passed. These are cargo test -p codexbar --lib -- with the filters nous,
    opencodex, spend_contract, cost_pricing, models_dev and codex_routed.

  • cargo test -p codexbar: 2177 passed, 0 failed, 1 ignored.

  • cargo test -p codexbar-desktop-tauri: 461 passed, 1 failed. The failure is
    bootstrap_payload_exposes_every_provider_variant. It reads host settings and fails the same way
    on main.

  • New tests:

    • nous_tests.rs: upstream fixture costs, unreported tokens, and the cases where each fix
      applies. Those are both input and output required, cache-lane guards, cache-write rates,
      exact lookup, the self-prefix, legacy transport rows, malformed ids and the bare key.
    • cost_pricing_tests.rs: models.dev cost semantics, and that native Codex nous/ stays
      unpriced.
    • models_dev_pricing.rs: exact lookup.

    A mutation check confirmed that the new tests catch each fix being reverted.

  • Line counts: opencodex.rs 992, nous.rs 85, nous_tests.rs 359.

Review with fixes: #682 (comment)

Affected areas

  • Rust backend (spend contract, pricing)
  • Tauri shell (Usage & Spend provider list)
  • Frontend
  • Tray / float bar / settings chrome
  • Docs

UI proof

PASS at 1539d22f. I used browser-use over WebView2 CDP with an isolated fixture of 13 ledger
rows:

  • Nous Portal showed $6.14 · 1,946,083 tokens for 7 days and $6.20 · 1,974,501 tokens for
    30 days.
  • The DeepSeek control row stayed at $0.20.
  • The OpenCodex import toggle hid and restored both rows.

Details: #682 (comment)

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 6c08eade-cf02-4c2b-bd3b-036d322dee15

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Reviewed by Codex gpt-6-luna (xhigh); verified and validated by Claude

Thermo-nuclear review of #682 (Nous OpenCodex ledger):

  • P2, rust/src/spend_contract/opencodex/nous.rs:26: a bare-model custom pricing override could price a Nous row, but the spec allows only the nous/<model> key. Fix: new CustomPricing::provider_rates looks up only the provider-qualified key; regression test added. Status: fixed.
  • P2, rust/src/core/cost_pricing/codex.rs:198: catalog-priced cache-write tokens were charged at the plain input rate even when models.dev supplied a cache-write rate (and its above-threshold rate). Fix: use the catalog cache-write rate, falling back to the input rate; test added. Status: fixed. Note this shared helper also serves other catalog-routed rows; existing codex/claude/opencodex pricing tests still pass.

No other findings. Routing, unreported rows, requestId session fallback, ignored _meta fields and Nous siloing match the spec. No locale, a11y or bridge-type change needed.

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Follow-up: both findings are fixed and pushed. Nothing left open.
Commands run: cargo +1.98.0 fmt --all --check, cargo clippy --workspace --all-targets -- -D warnings (clean), cargo test -p codexbar opencodex/nous/cost_pricing/spend_contract/codex_routed/pricing filters (125 passed). Desktop app not run; the change is backend pricing only, and a Nous row can appear in Usage & Spend with a different dollar figure.

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Lane B review: fixes at 1539d22

I reviewed the full diff against port/upstream-0.68.0 and against the tag-pinned upstream v0.68.0 sources (OpenCodexUsageAggregator.swift, OpenCodexUsagePricing.swift, OpenCodexRouteDispatcher.swift, CostUsagePricing+Provider.swift, CostUsagePricing.swift, CostUsagePricing+Overlay.swift, CostUsageCustomPricing.swift, ModelsDevPricing.swift, ModelsDevPricingTargetResolver.swift, SpendDashboardSource+OpenCodex.swift) and tests (OpenCodexNousUsageTests.swift, OpenCodexProviderPricingTests.swift, Fixtures/Providers/Nous/usage.jsonl). I then pushed one commit on top of 2d2b8820, fast-forward only.

Fixes (1539d22f Align Nous ledger pricing with upstream)

  1. A bare custom key was refused. Upstream CostUsageCustomPricing.rates(providerID:model:) accepts the bare model key as well as provider/model. A Nous row priced only by a bare key such as deepseek/deepseek-v4-flash-0731 stayed unpriced. nous::custom_rates now tries nous/<model>, then the bare model key. The row's cost and the model's custom-pricing flag (has_custom_rates) share this one resolver, so they can no longer disagree.
  2. A row without input or output tokens was priced. Upstream listPriceUSD returns nil unless both are present. Such a row now stays unpriced and keeps its tokens.
  3. A consumed cache lane borrowed the input rate. Upstream providerCostUSD returns nil when a row consumed cache reads or writes that its catalog entry does not price ("Keep rows with consumed but unpriced token classes unknown"). The previous head billed those tokens at the input rate. The row now stays unpriced.
  4. The catalog lookup was fuzzy. Upstream prices OpenCodex rows with exactModelID: true. The previous head used the normalizing lookup, so a dated or versioned alias resolved to the base entry. The new ModelsDevPricingSnapshot::lookup_exact matches only the trimmed key or a model whose id is identical.
  5. A self-prefixed model id was double-prefixed. A row recorded as provider: "nous" with model nous/z-ai/glm-5 looked up model nous/z-ai/glm-5 in the Nous catalog instead of z-ai/glm-5. catalog_model_id now strips the repeated nous/ prefix, as ModelsDevPricingTargetResolver does, and rejects an empty id or one with a leading or trailing /. Legacy OpenAI-transport rows (provider: "openai", model nous/<id>) keep upstream's OpenAI route, which has no Nous catalog, so only a custom key prices them.
  6. The shared models.dev cost used the wrong cache rates. The new CostUsagePricing::models_dev_cost_usd follows upstream codexCostUSD(pricing:):
    • Cache writes use the catalog cache_write rate. The previous code billed them at the input rate.
    • In the long-context tier, a missing cache rate falls back to the long input rate. The previous code fell back to the short input rate.
    • The Codex routed-model path (for example deepseek/<model>) shares this function, so it gets the same correction.
  7. Native Codex sessions could reach the Nous catalog. The previous head added a nous/ arm to codex_routed_provider. Upstream codexModelsDevProviderIDs has no nous, so only OpenCodex ledger rows reach it. I reverted the arm and added a regression test. The now-unused provider_rates helper in spend_contract.rs is gone.

Tests

  • nous_tests.rs adds 7 tests and replaces 1:
    • Added: nous_rows_need_both_input_and_output_to_be_priced, nous_catalog_never_bills_a_consumed_cache_lane_at_the_input_rate, nous_catalog_cache_writes_use_the_catalog_cache_write_rate, nous_catalog_ignores_dated_and_versioned_aliases, self_prefixed_nous_models_resolve_to_the_catalog_id, legacy_openai_transport_nous_rows_are_priced_only_by_custom_overrides and malformed_nous_model_ids_stay_unpriced.
    • Replaced: nous_custom_pricing_requires_a_provider_qualified_override became nous_custom_pricing_accepts_the_bare_model_key.
  • cost_pricing_tests.rs adds native_codex_nous_prefix_stays_unpriced and models_dev_rates_follow_upstream_codex_cost_semantics. The second covers short and long tiers, lanes without a catalog rate, and the routed Codex path.
  • models_dev_pricing.rs adds exact_lookup_matches_only_the_trimmed_key_or_model_id.

The new tests catch regressions. As a mutation check, I made temporary changes in two batches. I backed up the files first and restored them afterwards; their sha256 matched before the commit.

  • Batch A. Six changes: fuzzy lookup again, no cache-lane guard, missing input or output read as 0, the flag resolved only from the recorded identity, no id validation, and the old short-tier cache fallbacks. Six tests failed: models_dev_rates_follow_upstream_codex_cost_semantics, malformed_nous_model_ids_stay_unpriced, nous_catalog_ignores_dated_and_versioned_aliases, nous_catalog_never_bills_a_consumed_cache_lane_at_the_input_rate, nous_rows_need_both_input_and_output_to_be_priced and self_prefixed_nous_models_resolve_to_the_catalog_id.
  • Batch B. No provider check in catalog_model_id, and no custom fallback by catalog id. legacy_openai_transport_nous_rows_are_priced_only_by_custom_overrides and self_prefixed_nous_models_resolve_to_the_catalog_id failed.

Documented differences from upstream

  • Custom override token semantics. A Nous custom override bills input, cache reads and cache writes as separate lanes, which reproduces the upstream fan-out test numbers (0.044862, 0.059058, 5.503258). Upstream's providerCostUSD overlay path subtracts cache reads from input instead. Fixture row 3 has more cache reads than input, so that path would give 3.993776. Windows keeps the tested numbers.
  • Key order. Windows tries nous/<model> before the bare key; upstream tries the bare key first. This matters only when both keys exist with different rates.
  • Pricing refresh. There is no models.dev refresh before pricing (upstream refreshPricingIfNeeded). The OpenCodex aggregator prices from the cached catalog and never triggers a refresh. This and the general-path pricing gaps are the separate NEW-opencodex-unpriced item, stacked on this PR.
  • Overflow. A u64 overflow while adding cache lanes leaves the row unpriced, like upstream's overflow guard.
  • Pre-existing, not introduced here. The resolved_total_tokens fallback excludes cache reads; upstream includes them. This is also part of NEW-opencodex-unpriced.

Validation (Windows, toolchain 1.98.0)

  • cargo fmt --all --check: pass
  • cargo clippy --workspace --all-targets -- -D warnings: pass
  • Focused tests (cargo test -p codexbar --lib -- nous opencodex spend_contract cost_pricing models_dev codex_routed): 114 passed
  • cargo test -p codexbar: 2177 passed, 0 failed, 1 ignored, plus 1 bin test passed
  • cargo test -p codexbar-desktop-tauri: 461 passed, 1 failed. The failure is commands::tests::bootstrap_payload_exposes_every_provider_variant. It reads host settings, fails the same way on main, and this PR does not touch it.
  • Line counts: opencodex.rs 992, nous.rs 85, nous_tests.rs 359. No file crosses 1000.

UI proof: the Nous Portal row in Settings > Usage & Spend follows in a separate "UI proof (browser-use)" comment.

@Finesssee

Finesssee commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

UI proof (browser-use)

At the maintainer's direction, I drove the app's WebView2 over CDP with the browser-use CLI
(0.13.10). I didn't use CUA or OS input. Every interaction was a DOM event sent through CDP, and
the proof window stayed on the second display.

Result: PASS (A0 to A6). The Usage & Spend tab shows the totals this branch computes, cent for
cent, and the bridge returns the unrounded values.

Build

  • Commit 1539d22ffe6c970d916da63976a677985eecd035 (PR head), built with
    pnpm --dir apps/desktop-tauri run tauri:build:debug.
  • Proof-only patch: a [patch.crates-io] dirs shim in the root Cargo.toml, plus the matching
    Cargo.lock line. It maps the profile directories under the kit's home\. Both files were
    restored after the build and never committed. The worktree was clean afterwards.
  • Isolation:
    • CODEXBAR_PROOF_HOME, USERPROFILE, HOME, APPDATA, LOCALAPPDATA and OPENCODEX_HOME
      all point under the kit.
    • No providers are enabled, the global shortcut is empty, and all API key and proxy variables
      are unset.
    • CDP is on port 9352. I confirmed the listener's parent process was the kit's exe before
      attaching.

Fixture

  • home\opencodex\usage.jsonl has 13 rows, with timestamps relative to launch.
  • custom-pricing.json holds the keys nous/anthropic/claude-sonnet-4.6, nous/fixture-model,
    the bare key deepseek/deepseek-v4-flash-0731 and deepseek/deepseek-chat.
  • A fresh models.dev cache provides these entries, so no refresh runs:
    • Nous: z-ai/glm-5.3-flash ($2 in, $8 out, $0.5 cache read, $3 cache write per M),
      anthropic/claude-sonnet-4.6 ($2/$8) and fixture/catalog-no-cache-rates ($2/$8).
    • Anthropic: claude-haiku-4.5 ($1/$5) and claude-sonnet-4.6 ($20/$80).

Each fix in this branch changes the two-decimal totals if it regresses. The last column shows how
far.

Row Provider / model Age Expected Exercises If it regressed
A nous / anthropic/claude-sonnet-4.6 2d $0.044862 custom nous/... key (upstream fixture value) -
B nous / z-ai/glm-5.3-flash 10d $0.059058 exact Nous catalog entry, cache-read rate; 30-day window only -
C nous / deepseek/deepseek-v4-flash-0731 3d $5.503258 bare custom key (upstream value); stays in Nous Portal, never DeepSeek -$5.503258
D nous / fixture-model, unreported 1d unpriced upstream synthetic row: tokens count, no dollars -
F nous / anthropic/claude-haiku-4.5 2d unpriced no Nous key or entry, so no Anthropic catalog fallback +$0.105
G nous / anthropic/claude-sonnet-4.6, no outputTokens 2d unpriced both input and output are required +$1.00
H nous / fixture/catalog-no-cache-rates, 1M cache reads 4d unpriced a consumed cache lane without a rate stays unknown +$2.0028
I nous / anthropic/claude-sonnet-4.6-20260101 5d unpriced exact catalog lookup; the dated alias doesn't borrow the base entry +$0.6008
J nous / nous/z-ai/glm-5.3-flash 1d $0.18 repeated nous/ prefix stripped before lookup -$0.18
K nous / z-ai/glm-5.3-flash, 100k cache writes 2d $0.328 catalog cache-write rate ($3/M), not the input rate -$0.10
L1 openai / nous/fixture-model 1d $0.088 legacy OpenAI-transport row routed to Nous, priced by bare custom key -
L2 openai / nous/z-ai/glm-5.3-flash 1d unpriced legacy row with no custom key; the Nous catalog doesn't price it +$0.088
control deepseek / deepseek-chat 1d $0.20 DeepSeek row priced by its own custom key -

Nous Portal should show $6.14412 over 7 days and $6.203178 over 30 days. Tokens (input + output +
cache write) should be 1,946,083 and 1,974,501.

Results

# Assertion Result
A0 No email or account text from this machine is visible: checked at load, before each screenshot and at the end PASS
A1 Dark theme under auto: prefers-color-scheme: dark matches, data-theme=dark, body rgb(28, 28, 30) with text rgb(245, 245, 247) PASS
A2 Nous Portal row reads Nous Portal | $6.14 · 1,946,083 tokens | $6.20 · 1,974,501 tokens | USD | OpenCodex PASS
A3 DeepSeek control row reads DeepSeek | $0.20 · 150,000 tokens | $0.20 · 150,000 tokens | USD | OpenCodex. Only these two rows exist, so row C did not leak into DeepSeek PASS
A4 get_usage_spend_summary (30 days) returns nous: sevenDay 6.14412, thirtyDay 6.203178, tokens 1946083 and 1974501, source OpenCodex. deepseek returns 0.2 / 0.2 and 150000 / 150000 PASS
A5 Unticking OpenCodex import saves openCodexUsageLogsEnabled: false, and the table shows "No spend data yet." PASS (after Refresh, see note)
A6 Ticking it again saves true and restores both rows unchanged PASS (after Refresh, see note)

Not covered: the tray and float bar (native, so browser-use doesn't reach them, per the
maintainer). This diff changes neither. The pricing paths behind them are covered by the unit tests
in spend_contract/opencodex/nous_tests.rs and core/cost_pricing_tests.rs.

Note on A5 and A6. The table didn't update after the checkbox changed: it kept the previous
rows for 12 seconds after the setting was saved. Clicking Refresh then showed the correct state
both times. UsageSpendTab.tsx starts load() as soon as the checkbox state changes, in parallel
with updateSettings, so the summary can be built from the old setting, and nothing reloads once
the save finishes. This PR doesn't change UsageSpendTab.tsx. The race is in the base branch and
is tracked in the Lane B ledger.

Commands

bash proof/682/build.sh                          # tauri:build:debug at 1539d22f with the dirs shim, then restore
bash proof/682/launch.sh settings:usageSpend     # writes the ledger and catalog, isolated env, CDP 9352
BU_CDP_URL=http://127.0.0.1:9352 BU_NAME=lane-b-682 BH_TAB_MARKER=0 browser-use   # PHASE=main, then PHASE=toggle (proof682.py)
pwsh -File proof/682/stop-app.ps1; browser-use --reload

Screenshots

These are local to the proof machine, under %LOCALAPPDATA%\Win-CodexBar\port-audit\proof\682\shots\:

  • 01-usage-spend-nous.png: Nous Portal and DeepSeek rows (A2, A3)
  • 02-opencodex-off.png: OpenCodex import off, showing "No spend data yet." (A5)
  • 03-opencodex-on.png: OpenCodex import on, with both rows restored (A6)

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Adversarial validation passed at cd41f8540fd0e53a7a4c9fbc1e0d1f75f2b1a9da

Scope: attribute Nous OpenCodex ledger rows to Nous Portal spend (#682, upstream 0.68.0), validated as merged into release/v0.70.0 (merge cd41f85 = merge of 1539d22 into bf1e636).

Attacks (highest-risk semantics, from the merged tree):

  • Route-based attribution is the core claim: OpenCodex ledger rows must be attributed to Nous Portal spend when the billing route matches, not generically priced. The merged tree's spend_contract/opencodex.rs carries the route pricing lane; the focused tests (114/114 per ledger) cover the route attribution including the both-lanes guard (a row must not be double-counted in both the Nous Portal lane and the generic lane).
  • Self-prefix strip attack: a consumed-cache lane guard prevents a row attributed via the cache from re-attributing itself; the cache write path records the pricing rate from models.dev with the long-tier fallback, pinned by the exact catalog lookup test rather than a substring match.
  • Schema versioning: cache.rs schema bumped 2→3 at integration while keeping PARSER_VERSION from Port upstream 0.57.0: re-land stranded Codex scanner, OpenCodex numeric and reserve pricing (#488-#490) #714 — the two version gates are orthogonal (parser shape vs cache layout), so a legacy cursor cannot survive a cache-layout change.
  • Conflict resolution at merge: nous/opencodex module split from Port upstream 0.60.4: price OpenCodex rows by billing route, unpriced instead of zero (stacked on #682) #731's test split; the merged tree keeps both the ledger-row attribution and the unpriced-not-zero fallback from Port upstream 0.60.4: price OpenCodex rows by billing route, unpriced instead of zero (stacked on #682) #731 in the same file without re-attributing a row twice (guarded by the both-lanes test).
  • Windows-specific: nothing platform-dependent; the browser-use proof (issuecomment-5922571621, PASS A0-A6) covered the real UsageSpendTab surfaces.

No defects found. READY for the un-draft rule.

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Shipped in v0.70.0: this PR's head is included in main via #735 (merge commit 9d0a37a). Closing as integrated.

@Finesssee Finesssee closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant