Skip to content

ci(backend): add minimal go build+test CI for PRs - #36

Closed
bytemain wants to merge 26 commits into
walkingddd:mainfrom
bytemain:ci/go-build-test
Closed

bytemain wants to merge 26 commits into
walkingddd:mainfrom
bytemain:ci/go-build-test

Conversation

@bytemain

Copy link
Copy Markdown

目前仓库 PR 没有任何 CI 检查,只有 main 上的 build-and-release workflow。补一个最小 CI:PR + main push 触发,backend 目录下跑 go build ./... 和 go test ./...。

Copilot AI and others added 26 commits August 6, 2026 17:53
Co-authored-by: bytemain <13938334+bytemain@users.noreply.github.com>
Co-authored-by: bytemain <13938334+bytemain@users.noreply.github.com>
Co-authored-by: bytemain <13938334+bytemain@users.noreply.github.com>
Co-authored-by: bytemain <13938334+bytemain@users.noreply.github.com>
Co-authored-by: bytemain <13938334+bytemain@users.noreply.github.com>
The frontend bootstraps branding from window.__PRODUCT_INFO__ injected
into index.html, so the standalone /api/product-info endpoint is no
longer called by anything: remove the route, handler, and the unused
ProductInfoResponse type.

Also remove the __CPA_HELPER_LOGO_URL__ ReplaceAll in injectBranding:
the vite build substitutes that placeholder at build time, so it can
never match in served HTML. The custom favicon is now applied by the
frontend from __PRODUCT_INFO__ instead.

Revert the pointless await-promise-then rewrite in LoginView and fix
the window cast that broke vue-tsc -b.

Co-authored-by: bytemain <13938334+bytemain@users.noreply.github.com>
Signed-off-by: 怪味胡豆 <raft-mobile-guaiweihudou@mail.build>
Signed-off-by: artin <artin@cat.ms>
Backend: keeperQuotaWindowUsage gains window progress fields
(elapsed seconds/percent) and ProjectedCostUSD — observed cost linearly
extrapolated to the full window; stale windows project to their own
total. Exposed via /api/codex-keeper/accounts as window_elapsed_seconds,
window_elapsed_percent, projected_cost_usd.

Frontend: per-window tags/text show the projection, and the accounts
dashboard gains a 本窗口预计消费 metric card summing each enabled
account's long-window projection at list price.
The account status page hardcoded the 5-hour/weekly label pair for
every paid plan. Upstream usage payloads decide the real window length
(primary_window_seconds), and pro accounts can report a weekly primary
window, so the page showed weekly data under a 5-hour label.

Pick the label from the actual window seconds (5h/weekly/monthly) and
keep the conventional pair only as a fallback when the seconds are
unknown.

Signed-off-by: 怪味胡豆 <raft-mobile-guaiweihudou@mail.build>
Signed-off-by: artin <artin@cat.ms>
Merge the UI-only success-rate precision fix. Production deployment remains a separate authorized step.
Per the task walkingddd#30 contract pinned in-thread: both 历史用量 and 我的用量
(shared dashboard component) gain a 缓存命中率 card showing only the
two-decimal hit rate, footnote = the same time-range label as the
requests card, no absolute amounts.

Rate = provider-aware cache-hit tokens / aggregated input tokens,
capped at 100%. The summary emits a new cache_hit_tokens field:
Claude-style records count cache_read_tokens (upstream keeps cache
reads outside input_tokens), everything else counts cached_tokens
(a subset of input). All existing fields unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
usage: 缓存命中率 metric card (task walkingddd#30 rewrite)
Add a Reset action to the admin account-status page. The backend endpoint
POST /api/codex-keeper/reset-quota resolves the account's auth_index and
calls CLIProxyAPI POST /v0/management/reset-quota (clearing the account's
quota-exceeded/cooldown/model error state); only after CLIProxyAPI
succeeds is the per-account reset counter incremented (new
codex_keeper_quota_resets table). The accounts listing now carries
quota_reset_count / last_quota_reset_at and the page shows a Resets
column plus a confirm-guarded per-row Reset button.

Contract test drives the real route: successful resets increment the
counter (1 then 2, visible in the listing), a CLIProxyAPI failure
surfaces the error without incrementing, and empty/unknown auth names
are rejected before any CLIProxyAPI call.

Co-authored-by: feiniu (Raft agent) <admin@oranix.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…al wire DTO (#8)

* fix(account-status): fail-closed reset verification, row fence, minimal wire DTO

Address the four post-landing review findings:

- Verify the CLIProxyAPI reset response instead of trusting any 2xx: require
  status=ok with the exact requested auth_index, otherwise fail closed and do
  not increment the counter (covered by empty-body / wrong-index / bad-status
  contract cases alongside the HTTP-failure case).
- Return a deliberately minimal reset DTO (name, quota_reset_count,
  last_quota_reset_at) instead of serializing the internal keeperAccount
  (auth_index/email/last_error no longer cross the wire); the contract test
  now asserts no extra fields leak.
- Include reset-quota in the row-level action fence so a resetting row cannot
  be concurrently disabled/re-prioritized/deleted/refreshed.
- Correct the table scroll widths to the actual column sums
  (disabled 1454, normal 1958) so the fixed action column is not clipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(i18n): map the reset backend messages for English mode

Add messages.ts pairs for all four backend strings introduced by the reset
feature (fail-closed verification message, missing-auth_index hint,
empty-auth_name validation, account-not-found) so English mode no longer
shows mixed-language errors, with i18n smoke assertions pinning each pair.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(account-status): extend the row fence to bulk and programmatic entries

Bulk select-all, bulk/programmatic refresh, and bulk delete previously
ignored rows with an in-flight per-row action, so a resetting (or
toggling/deleting) row could be hit concurrently through those paths.
All three entries now respect the same isRowActing fence as the per-row
buttons: select-all skips acting rows, refresh drops busy targets and
proceeds with the idle rest, and bulk delete filters busy rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(account-status): enforce exact auth_index echo on reset confirmation

The CPA response auth_index is now compared without trimming, matching the
documented exact-match contract (the request-side auth_index is already
trimmed, so a well-behaved CPA echo passes; a whitespace-padded echo now
fails closed instead of being leniently accepted). Adds a padded-index
deceptive-response case asserting no counter increment.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: feiniu (Raft agent) <admin@oranix.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…dow timeline (#9)

Per owner clarification: the manual Reset button and its backend endpoint
stay, but the standalone reset-count column is dropped in favor of showing
the account's own quota reset schedule. A new Quota Reset Windows column
lists each window (primary/secondary) with its reset time and a relative
countdown, reusing the existing primary_reset_at/secondary_reset_at and
window data (no new upstream calls or backend changes). The reset counter
still increments on manual reset and is shown in the reset-confirm dialog.

Co-authored-by: feiniu (Raft agent) <admin@oranix.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…its (#10)

Rebuild the account-status reset timeline on the correct data source (wham/rate-limit-reset-credits) instead of the 5h/weekly usage windows. Strict parsing (inner 2xx status, non-negative safe-integer available_count, id/granted_at/expiry validation, present-vs-null expiry), identity-scoped preserve-on-failed-fetch, migration 202609040002. Independent review GO by 怪味胡豆 on exact head 96fc9ee (diff sha256 83eda666). Follow-ups (fetched_at/stale metadata, account_id identity key) tracked in task #71.
Post-reset synchronous single-account re-inspection (per-auth locked) + full success/failure/skip/partial audit for reset-quota/enable/disable/delete(bulk)/priority, with safe reason codes and DB-write-failure surfacing. Independent review GO by 怪味胡豆 (task #73) on exact head 9ac5b93, diff sha256 97da0944, triple hash-verified. Closes task #72.
* feat(codex-keeper): real reset-credit consume + subscription renewal + drop manual reset counter

真重置: resetKeeperQuota now redeems a real OpenAI reset credit instead of only
clearing the local 429 cooldown. It rebuilds the merged auth detail (list+download,
so account_id and auth_index stay same-row), re-reads the authoritative
available_count FRESH (a NULL/stale snapshot is never treated as a known 0; an
unknown count blocks rather than degrading to a cooldown-only clear), and when a
credit is available POSTs wham/rate-limit-reset-credits/consume via the per-auth
api-call egress. The redeem_request_id is a client idempotency key generated once
per operation, so keeperRequest's internal retries reuse the same key and can never
double-consume. A 2xx inner status alone is not proof: the inner `code` decides —
only reset/already_redeemed count as consumed; no_credit/nothing_to_reset clear only
the cooldown; any transport error / non-2xx / unrecognized code fails closed and
never reports a redemption. It then always clears the local cooldown via /reset-quota.

真续期: parse the ChatGPT subscription renewal time from the account id_token claim
chatgpt_subscription_active_until (surfaced by CPA ListAuthFiles). Migration
202609060001 adds subscription_active_until. The extractor is tri-state so a
missing/malformed claim never corrupts a good snapshot: parsed=known value,
confirmed-absent=known nil (clears the stored value), unreadable/malformed=unknown
(upsert preserves the previous value). Surfaced in GET /accounts and a new 续期时间
column with countdown.

去噪: remove the meaningless manual reset counter — drop the codex_keeper_quota_resets
increment, mergeKeeperQuotaResetCounts, and QuotaResetCount/LastQuotaResetAt from the
structs/response/mapper. Migration 202609060002 DROPs the table (Down recreates the
empty schema; the historical counts are intentionally not restored). Reset result DTO
simplified to {name, consumed}. Frontend reset dialog reworded and the mislabeled
"主动重置次数" cell header relabeled to the accurate 可用重置额度.

Tests: rewrite the reset route test for the consume/no-credit/fail-closed/unknown-count
paths and wire minimalism; add a tri-state unit test for the subscription extractor.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): harden real reset-credit consume — per-auth lock, identity binding, redeem ledger, redaction

Addresses the review of the true-reset feature (real, paid, limited credits).

Concurrency & identity:
- Hold the per-auth lock across the whole fresh-count → consume → cooldown sequence
  (non-blocking; a contended request returns 409) so concurrent resets of one account
  can redeem at most one credit.
- Resolve the account identity FRESH and bind everything to it: validate the list
  entry's and download detail's EXPLICIT auth_index/account_id separately (never the
  merged id_token, whose download side is a raw JWT string), fail closed on a
  list-vs-detail conflict, never fall back to the auth name for auth_index, require the
  download's top-level account_id (the consume header) while treating the list's
  id_token.chatgpt_account_id as an optional cross-check, and require the fresh
  auth_index to exactly match the DB row.

Cross-operation idempotency (redeem ledger, migration 202609060003):
- Persist a per-auth redeem_request_id bound to (auth_index, account_id). An
  identity-matched pending redeem is ALWAYS replayed with its original key — even when
  the fresh available_count is 0 — because a lost first response may have consumed the
  last credit, and only replaying the same key recovers already_redeemed. A fresh id is
  minted only when there is no usable pending redeem and a credit is available; an old
  account's id is never inherited. The ledger stays pending through the cooldown step
  and is finalized only after the whole operation succeeds; a consume-then-cooldown
  failure is audited as an irreversible partial and kept pending for an idempotent
  retry. Account delete and prune clear the ledger.

Redaction:
- Never log an external inner body (even truncated). Consume outcomes audit only a
  stable classification plus whitelisted inner status_code / recognized code.

Subscription (strict):
- Parse the id_token claim strictly: reject NaN/Inf/fractional/overflow and bound the
  epoch to a sane 2000–2100 window. Scope preserve-on-unknown to auth identity so a
  reassigned auth_index never keeps the previous account's renewal date.

Frontend:
- Reset API returns {name, consumed}; the toast now reports real redemption vs
  cooldown-only. The reset-credit count shows "未知/陈旧" when unknown instead of
  substituting a possibly-truncated detail length or 0. The account detail drawer shows
  the renewal time. New backend errors are mapped in i18n (with smoke assertions) so the
  English UI shows specific recovery guidance.

Rollback:
- Document the Down-first-then-old-binary contract (docs/migrations-rollback.md) and
  test that Down to 202609040002 restores the codex_keeper_quota_resets compat table.

Tests: per-auth concurrency (single consume), lost-response → same-key replay (incl. at
count=0), redaction canary, auth_index mismatch, missing account_id, list/download
account_id cross-check, ledger identity change, strict subscription parser, and the
rollback compat-schema assertion.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): round-3 review hardening — resolver guards, outcome DTO, dry-run, rollback CLI

Deep-review follow-ups on the real reset-credit consume path (paid, limited credits).

Redeem ledger control flow:
- Resolve an identity-matched pending redeem FIRST and replay it with its original key
  even when the fresh reset-credit GET is unavailable — the count gate applies only to a
  brand-new operation, so a lost-response pending is never stranded while the count
  endpoint is down.
- Make the ledger claim a DB-atomic guarded upsert (overwrite only a non-same-identity
  pending) + read-winner, so concurrent claimers converge on one request_id. Documented
  that this hardens the overlap window but the contract is a SINGLE ACTIVE INSTANCE on
  single-writer SQLite (no client idempotency key spanning HTTP).
- Keep the ledger pending through the cooldown step; finalize only after the whole
  operation succeeds; audit an irreversible partial (consumed + cooldown_failed) and keep
  it pending for an idempotent retry.
- Delete the state row and ledger row in ONE transaction (a ledger-clear failure now
  fails the delete) so a same-identity re-import can never reuse a residual pending.

Identity resolution (irreversibility guards):
- Require the list entry to be Codex (and a detail with an explicit non-Codex type fails
  closed); require the download's account_id and access_token; reject a duplicate remote
  name; validate a detail name matches the request; validate intra-object alias
  consistency (auth_index/authIndex/index and account_id vs id_token claim) instead of
  silent precedence; never fall back to the auth name for auth_index.
- Acquire the per-auth lock before reading state (decide under the lock), and fail closed
  when the keeper runner is absent instead of proceeding unlocked.

Dry-run: a manual reset is fail-closed under dry-run (before any remote call or ledger
write) so an admin testing the keeper never silently burns a credit. NOTE: the default
config is DryRun=true, so a fresh deploy blocks reset until an admin turns it off.

Outcome DTO: the reset response now carries a distinct outcome
(reset|already_redeemed|no_credit|nothing_to_reset|cooldown_only) instead of a lossy
consumed bool; the UI messages each case. Strict subscription parsing now range-gates the
RFC3339/date string paths too (not just numeric), and the subscription write is gated on a
confirmed identity so a failed detail read preserves the old renewal snapshot.

Rollback: implement a real `migrate down-to <version>` subcommand (allowlisted target;
refuses a non-downgrade; refuses when pending redeems exist unless --allow-pending) with
explicit previous→target output — a bare `migrate` only runs Up. Documented the
Down-first-then-old-binary contract, the offline/quiescence requirement, and the
subscription data loss.

Frontend: derive the table scroll width by summing column widths (drift-proof after adding
the renewal column) and localize all new errors with i18n smoke assertions.

Tests: pending-replay-when-fetch-unavailable, count-zero replay, outcome-codes contract,
dry-run fail-closed, resolver guards (duplicate name, alias conflict, missing token,
detail-name mismatch, non-Codex), subscription-preserved-on-unconfirmed-identity, string
extreme-date subscriptions, and the migrate down-to CLI (rollback target + pending block).

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(cli): reject unknown migrate subcommands/flags and DB newer than binary

A destructive maintenance command must not silently misfire:
- `migrate <unknown>` now errors instead of falling back to Up.
- `migrate down-to <v> <unknown-flag>` now errors instead of ignoring the flag
  (only --allow-pending is accepted).
- MigrateDownTo refuses when the DB version is newer than the binary's embedded
  LatestVersion (its Down migrations would not cover the extra versions).

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): put all per-account mutations under the per-auth fence + partial audit

Concurrency (real double-consume blocker): delete, prune, enable/disable, and priority now
share the SAME per-auth lock as reset/inspection, so none can delete the redeem ledger or
drift the remote/DB identity while a reset holds the lock mid-consume (which could drop an
in-flight redeem's idempotency key and let a later operation mint a new one → double
consume).
- deleteKeeperAccount takes the lock (409 on contention); state+ledger delete extracted to
  deleteKeeperStateAndRedeem (one transaction).
- pruneKeeperMissingAuthStates tries the lock per stale name and SKIPS any a reset holds
  (retried next cycle) — it never blocks, so no lock-order deadlock; bulk delete locks one
  name at a time (never two at once).
- setKeeperAccountDisabled / updateKeeperAccountPriority take the lock too. These are
  handler-only; processKeeperAuth uses the lock-free setKeeperRemote{Disabled,Priority}
  helpers, so there is no self-conflict with the inspection run that already holds the lock.
- Missing runner fails closed instead of proceeding unlocked.

Audit: a consumed-credit-then-cooldown-failure now audits result=partial (an irreversible
partial), not a plain error, so the money-affecting half-completion is visible.

i18n: map createKeeperRedeem's inconsistent-state error and the new keeper-not-initialized
errors, with smoke assertions.

Tests: delete-conflicts-with-in-flight-reset (deterministic via the in-lock gate),
partial-audit-on-cooldown-failure, and a cross-process claim convergence test — two App
instances with independent DB handles on one SQLite file concurrently claim and converge on
one winner request_id (proves the DB-atomic claim, not just a comment).

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): protect persisted pending redeems from prune/delete (ledger safety)

The prior per-auth fence used only the in-memory lock, which is empty after a restart —
so a successful-but-empty/transient remote auth-files list could prune an account that
still holds a persisted status='pending' redeem, dropping the sole idempotency key (a
later reappearance then mints a new key → double consume). This is a ledger-safety
blocker, not a deployment-prerequisite waiver.

- Add hasPendingKeeperRedeem: a PERSISTENT (DB) check, independent of the in-memory lock.
- pruneKeeperMissingAuthStates retains (skips + audits) any stale account with a pending
  redeem, and now fails closed when there is no runner (no fence) instead of deleting
  state+ledger unlocked — consistent with reset/delete/priority.
- deleteKeeperAccount refuses when a pending redeem is unresolved (reconcile first).

Tests: prune retains an account with a pending redeem while pruning a non-pending stale
one; prune deletes nothing without a runner; delete is refused with a pending redeem.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): key the redeem ledger by full identity so no pending key is overwritten

The ledger was keyed by auth_name alone, so createKeeperRedeem overwrote a different
identity's pending row when an auth_name was rebuilt onto another account — dropping the
old identity's unresolved (maybe-consumed) idempotency key. If the original account later
returned, a fresh key could be minted and double-consume.

- Migration 202609060003 now keys the table by (auth_name, auth_index, account_id), so
  distinct identities sharing an auth_name each keep their own row; a redeem is never
  overwritten across identities.
- createKeeperRedeem upserts per full identity (ON CONFLICT on the composite key),
  overwriting only its own terminal row; a concurrent same-identity pending row is still
  preserved (cross-process convergence unchanged).
- lookupPendingKeeperRedeem selects by full identity; hasPendingKeeperRedeem now reports a
  pending redeem for ANY identity under the auth_name, so delete/prune still fully protect
  every in-flight key.

Test: an unknown-outcome redeem for identity A survives a reset under identity B (same
auth_name); when A returns, its ORIGINAL key is replayed (already_redeemed) rather than a
new one.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): scope subscription identity to account_id + stop handler masking the partial audit

Subscription account-swap (identity scope): CPA's file auth_index is a hash of
provider+path, so swapping a filename to a different OpenAI account keeps auth_index
unchanged. Scoping the subscription preserve-on-unknown to auth_index alone would let the
new account inherit the old account's renewal date. Migration 202609060004 adds an
account_id column to codex_keeper_auth_states; the inspection records it; and the upsert's
subscription CASE now preserves on an unknown claim only when the account is unchanged and
CLEARS on a confirmed account swap (both account_ids known and different). An unconfirmed
identity (detail read failed) still preserves.

Partial audit no longer masked: the reset handler previously wrote a generic
result=error line after resetKeeperQuota had already audited result=partial, so the log
tail hid the irreversible partial. resetKeeperQuota now returns a partial-coded AppError
(reset_partial, HTTP 409) for a consumed-credit-then-cooldown-failure, and the handler
skips its generic re-audit for that code — leaving result=partial as the single audit line.

i18n: localize the partial-reset message. Tests: subscription cleared on same-index
account swap / preserved on same account; partial-audit test also asserts no generic
result=error line masks it; the reset-quota-failure-after-consume cases now assert 409.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* docs(codex-keeper): remove stale Consumed comment on the reset result (DTO is Outcome)

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* docs: cover migration 202609060004 (account_id) in the rollback runbook + test

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): scope the reset-credit snapshot preserve/clear by account_id

Same as the subscription fix: the reset_credit_count / reset_credits upsert CASE compared
only auth_index, but CPA's file auth_index is a hash of provider+path and stays the same
when a filename is swapped to a different OpenAI account. A new account whose reset-credit
fetch failed would then inherit the previous account's count/schedule (wrong UI and wrong
next-reset decision). The CASE now preserves on an unknown/unconfirmed identity and CLEARS
only on a confirmed account swap (both account_ids known and different).

The existing identity-change test now models the change by account_id (its premise was
auth_index-only); a new same-index/different-account snapshot test pins the clear.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): make account_id the identity — key the redeem ledger by account_id, treat auth_index as routing

CPA's auth_index is a hash of the file path (and auth_name is a filename), so BOTH change
on a file rename/move/reorder while the OpenAI account_id stays the same. Keying the redeem
ledger (or the reset identity gate) on the routing selectors meant a legitimate reindex/
rename stranded an in-flight pending redeem → a new UUID → possible double-consume.

Ledger (migration 202609060003, rewritten — it is unpublished; prod is at 202609040002 and
every 0600xx migration has only run in disposable test DBs):
- Keyed by the STABLE account_id. lookup/create/finish take account_id; the same account
  converges on one request_id however it is routed, and a lost-response pending is found
  after a reindex/rename. Distinct accounts keep distinct rows; a resolved (terminal) row
  lets the next claim mint a fresh id.

Reset path:
- The identity gate compares the fresh account_id against the DB row's stored account_id (a
  mismatch — a stale page acting on a swapped-out account — fails closed as
  account_id_mismatch). auth_index is used as-is for ROUTING and is no longer required to
  match the DB (a reindex must not block).
- Inspection reconciles the account identity across the list id_token.chatgpt_account_id and
  the download account_id (keeperReconcileInspectionAccountID); on conflict it treats the
  identity as unknown — binds no account_id, does not trust the list subscription, and skips
  the reset-credit fetch — so a mixed A/B response never writes one account's data onto
  another's row.

delete/prune: because the ledger is account-keyed (not file-keyed), removing a file can
never drop the account's key — delete/prune now remove only the state row (no ledger touch,
no pending-refuse/retain), still under the per-auth fence for state consistency.

State: codex_keeper_auth_states.account_id (migration 202609060004) is now read back into
keeperAuthState so the reset gate can compare it; the reset-credit and subscription upsert
CASEs already clear on a confirmed account swap and preserve otherwise.

Tests: ledger per-account & route-agnostic claim; reset replays a pending across an
auth_index change (rename); account_id mismatch fails closed while an index-only change
proceeds; prune/delete keep the account ledger; identity reconciliation; and migration
coverage now includes 202609040002 → head → 202609040002 → head (replay).

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): route the cooldown clear with the FRESH auth_index, not the stale DB value

resetKeeperQuota consumed with the fresh identity.authIndex but still routed
/v0/management/reset-quota (and audited) with the stale DB auth_index. On a reindex
(file rename/move/reorder) that meant "consume the credit on the new index, clear the
cooldown on the old index" — clearing the wrong account or leaving an irreversible
partial. After the identity resolves, ALL remote routing (consume, cooldown clear, the
status=ok/auth_index echo check) and audits now use the fresh identity.authIndex; the DB
value is kept only for the initial existence check / page snapshot.

Test: the auth_index-change case now also asserts /reset-quota was routed with the fresh
index ([idx-different]), not the stale one.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): add an account_id-level fence so two files of one account can't double-consume

The per-auth lock keys on auth_name (the file), but the paid/limited credit and the redeem
ledger are the ACCOUNT's. If the SAME OpenAI account is briefly reachable via two filenames
/routes (rename overlap, duplicate import), two resets take DIFFERENT auth_name locks; A
could claim → consume → finish (terminal) while B sits between fresh-fetch and claim, then
B's account-keyed createKeeperRedeem sees the terminal row and mints a NEW request_id →
a second credit. This needs no blue/green or multi-instance, so the single-active-instance
contract does not cover it.

resetKeeperQuota now also acquires a non-blocking account_id-level fence
(KeeperRunner.tryLockAccountID) right after the identity resolves, held across the whole
pending-lookup → fetch → claim → consume → cooldown → finish sequence; a contended reset of
the same account returns 409. The account fence is released on return (before the handler's
chained inspection).

Test: two files (a.json idx-A, b.json idx-B) that map to one account_id — while fileA holds
the fence (deterministically blocked in its fresh fetch), a reset of fileB is refused (409)
and exactly one credit is consumed.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): fail closed on an inspection identity conflict (identity_error) + localize the message

An inspection whose list and download identities disagree previously only skipped the
reset-credit/subscription writes but still fetched usage, wrote email/priority, and marked
result=healthy — so keeperRefreshAuditOutcome misreported a post-reset refresh as ok.

Now processKeeperAuth reconciles BOTH the account_id AND the auth_index across the list and
detail (keeperReconcileInspectionIdentity); on any conflict (or a self-contradictory
source) it bails out early with result="identity_error", preserves the prior snapshot
(no usage/reset-credit fetch, no account-data write), records a stable error, and the run
carries a new IdentityError stat so keeperRefreshAuditOutcome returns error/identity_error
(never ok). The conflict message is stable Chinese, mapped in i18n with smoke assertions
(also adds the previously-unmapped account_busy assertion).

Tests: keeperReconcileInspectionIdentity (account/index agree, conflict, self-conflict);
an end-to-end InspectAccountsLocked conflict (list account A vs detail account B) that
asserts no reset-credit fetch, the prior snapshot preserved, IdentityError=1, and
audit=error/identity_error; and keeperStats.add covers the new field.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): identity_error must preserve the full prior snapshot, not clobber it

The identity_error bail-out called persistState(result) on a near-empty result (only
Name/Result/CheckedAt/subscription set), and upsertKeeperState writes email, auth_index,
account_type, disabled, priority, usage, quota with `= excluded.*` (not COALESCE/CASE). So a
list/detail identity conflict cleared the prior good snapshot's email/auth_index/account_id/
usage/quota/priority to NULL/false — the opposite of "preserve snapshot" — and a cleared
auth_index would then block a later reset. Only reset_credits/subscription had CASE-preserve.

Route the identity_error write through a dedicated markKeeperIdentityError that UPDATEs only
last_error/latest_action/last_checked_at/updated_at and touches NO business column, so every
existing value survives intact. It deliberately does not set last_healthy_at (a conflict is
never a healthy refresh) and does not INSERT: a first-ever inspection that hits a conflict
touches zero rows rather than persisting a partial/ambiguous identity.

Test: TestKeeperInspectIdentityConflictPreservesSnapshot now seeds a FULL snapshot (email,
auth_index, account_id, account_type, disabled, priority, usage, quota, subscription, reset
credits, last_healthy_at) and asserts every column is unchanged after the conflict, that
last_healthy_at did not advance, and that only last_error/latest_action/last_checked_at moved.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): treat "neither source has account_id" as identity-unknown — skip account-scoped writes

keeperReconcileInspectionIdentity returns ("", true) when NEITHER the list entry nor the
download detail carries an account_id (a legacy auth_index-only credential): the identity is
consistent but UNKNOWN. processKeeperAuth then still fetched the reset-credit snapshot (the
api-call omits the Chatgpt-Account-Id header, so OpenAI resolves the account from $TOKEN$
alone and the snapshot cannot be attributed to a resource) and still bound the list's
subscription renewal claim — both account-scoped, both unsafe without a stable account_id, and
neither can detect a same-index account swap.

Now processKeeperAuth gates account-scoped work on accountIDKnown (acct != ""). When the
account_id is unknown it still runs usage/priority (transient current state), but skips the
reset-credit fetch (leaving the fields nil so the upsert's COALESCE preserves the prior
snapshot), clears the up-front subscription claim (SubscriptionActiveUntil=nil,
SubscriptionKnown=false, so its CASE also preserves), and flags ResetCreditsUnavailable so the
refresh is audited partial/reset_credits_unavailable rather than a falsely-healthy ok. The
previously-known account_id survives via COALESCE(excluded.account_id, existing).

Test: TestKeeperInspectNeitherAccountIDSkipsAccountScopedWrites — list+detail both lack
account_id over a seeded snapshot (known account_id, reset credits, subscription); asserts 0
reset-credit fetches, reset credits + subscription + prior account_id all preserved, and
audit=partial/reset_credits_unavailable.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* test(codex-keeper): make the neither-account-id subscription guard right-cause

The neither-account-id test's list fixture had no subscription claim at all, so it did not
actually exercise the SubscriptionActiveUntil=nil/SubscriptionKnown=false guard — deleting
those two lines still left the test green. Give the list a PARSEABLE renewal claim that
differs from the seeded snapshot (via id_token.chatgpt_subscription_active_until, still no
chatgpt_account_id), and assert the stored renewal equals the OLD snapshot and is NOT the list
claim. Verified by mutation: removing the guard now fails the test (the list renewal binds).

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): decode raw-JWT id_token when resolving account_id identity (fail-closed on conflict)

keeperExplicitAccountID only read id_token when it was a map, so it ignored CLIProxyAPI's real
download form — a raw JWT string — entirely. A deceptive/misrouted detail with top-level
account_id=A but an id_token whose own chatgpt_account_id claim is B passed as consistent: the
inspection could mix A's metadata with B's token, and the reset resolver would consume a real,
irreversible reset credit on A's header/route using B's token. (Mutation-verified: with the old
code the deceptive-JWT reset logs op=reset-consume result=ok code=reset consumed=true.)

Now keeperExplicitAccountID decodes the id_token via the existing keeperIDTokenClaims (map, JSON
string, or raw JWT payload) and cross-checks chatgpt_account_id from BOTH the flattened top level
and the nested OpenAI auth namespace (https://api.openai.com/auth.chatgpt_account_id — where the
raw download JWT actually carries it) against the top-level account_id; any disagreement is an
identity conflict. An id_token that is present but unparseable leaves the identity indeterminate
and also fails closed. This covers both the inspection path (→ identity_error, snapshot preserved)
and the reset resolver (→ 422, no consume) since both call keeperExplicitAccountID.

Tests: keeperReconcileInspectionIdentity gains raw-JWT agree/flattened-conflict/nested-namespace-
conflict/unparseable cases; a reset entry test injects a nested-claim raw JWT and asserts no
consume/no reset-quota (422); an inspection entry test asserts no reset-credit fetch, the usage +
reset-credit snapshot preserved, and audit=error/identity_error.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* docs(codex-keeper): correct stale lock/DTO comments to the account_id-fence + outcome model

Comment-only, no behavior change.

- codex_keeper_reset_test.go: several comments still described the old boolean DTO
  (consumed=true/false); the wire is now `outcome`
  (reset|already_redeemed|no_credit|nothing_to_reset|cooldown_only). Updated the struct
  doc and the TestKeeperReset / count-zero-replay comments so a future reader is not
  misled about the response shape.
- lookupPendingKeeperRedeem / createKeeperRedeem: the doc comments still claimed the
  per-auth (auth_name) lock is the serializing guard for the money-critical window. After
  the account_id rekey, the guard that serializes the same OpenAI account across different
  files/routes is the per-account_id fence; the per-auth_name lock is per-file and does not
  by itself serialize two files of the same account. Corrected both so the mutex boundary is
  not reused incorrectly later.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): fail closed on a present-but-wrong-type identity alias (not silently ignore it)

keeperExplicitAuthIndex / keeperExplicitAccountID / keeperClaimsAccountIDs read identity fields
through keeperString / a map type-assertion, both of which map a present-but-wrong-type value to
"absent". So a deceptive entry like {auth_index: 123, authIndex: "idx-A"} silently trusted the
string sibling, and a JWT namespace https://api.openai.com/auth: "not-object" (or a numeric
chatgpt_account_id) was ignored rather than rejected — violating the "any illegal/conflicting
alias fails closed" contract for an irreversible consume.

New keeperExplicitStringField distinguishes absent (key missing/null → "") from present-invalid
(present with a non-string type → errKeeperIdentityConflict). All three helpers now use it, and
keeperClaimsAccountIDs additionally rejects a present-but-non-object auth namespace. A present
identity field must be a string (the namespace must be an object) or the identity fails closed
(inspection → identity_error, reset → validation error).

Tests: keeperReconcileInspectionIdentity gains authindex-wrong-type, account-id-wrong-type,
jwt-namespace-not-object, and jwt-nested-claim-wrong-type cases.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): route/attribute inspection api-calls by the reconciled identity, not the raw merge

processKeeperAuth resolved the account_id but passed the raw right-biased `merged` to
checkKeeperUsage/fetchKeeperResetCredits, which read the top-level account_id for the
Chatgpt-Account-Id header and call keeperAuthIndex for routing. Two gaps:

1. account_id known only from the list id_token claim (legacy detail with no top-level
   account_id) → merged["account_id"] empty → the api-call omitted Chatgpt-Account-Id even
   though the account was known, yet still wrote an account-scoped snapshot (inconsistent
   attribution).
2. auth_index present only in the list while the detail carried an explicit-null auth_index →
   the right-biased merge dropped it → keeperAuthIndex fell back to the auth NAME for routing.

keeperReconcileInspectionIdentity now returns a keeperInspectionIdentity{accountID, authIndex}
(per-source validated). processKeeperAuth normalizes `merged` with the reconciled auth_index
(when non-empty) and account_id (when known) before any account-scoped work, so routing and the
account header come only from the reconciled identity, never re-derived from the raw merge.
Business fields still read from merged.

Tests: TestKeeperInspectListOnlyAccountIDSetsHeader (account_id only in the list claim, detail
auth_index explicit-null) asserts both api-calls carry Chatgpt-Account-Id=acct-A and route
auth_index=idx-1 (not the name), and the snapshot is written; the reconcile unit test switches
to the identity struct.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): CAS-persist a legacy NULL account_id on reset before consuming (bind-on-reset)

A pre-account_id (NULL account_id) row is allowed to bind on a reset, but the account_id was
only used in-memory — never written back — so the NULL row stayed bindable-to-anything on every
reset. The account_id swap gate only fires when the stored id is non-NULL, so if the same
filename was swapped to a different account between resets (with no successful inspection
persisting the new id in the gap the reviewer noted), the next reset would consume the WRONG
account's credit undetected.

resetKeeperQuota now CAS-binds the resolved account_id onto a NULL row inside the account fence
and BEFORE any pending-lookup/fetch/consume/cooldown/ledger write, via bindKeeperAccountID:
`UPDATE ... SET account_id=? WHERE auth_name=? AND account_id IS NULL`. affected==1 binds;
affected==0 re-reads and continues only if the stored id already equals the resolved id
(concurrent/idempotent), else fails closed as an account_id mismatch. So a NULL row is never left
unbound after a real consume, and a later same-name swap is caught by the existing gate.

6e278b0d (withdrawn by review as non-blocking): same-account duplicate files are separate
auth_name rows, so one row's inspect never overwrites the other's — it is task #71 display
staleness, not a write-back/monetary issue. No fence change; documented as duplicate-row eventual
consistency in the PR body instead.

Test: TestBindKeeperAccountIDCASAndSwap — NULL row binds to A, idempotent re-bind of A, and a
different-account bind fails closed without changing the stored id.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* fix(codex-keeper): take the per-account fence on the stored account_id BEFORE the resolve (close the pre-fence double-consume window)

The account fence was acquired AFTER the fresh list/download resolve, leaving a pre-fence window:
two different files of the SAME account could interleave — file A blocked mid-download (no fence
yet) while file B fully resolves, consumes, finishes and releases, then A resumes, sees a terminal
ledger and mints a fresh redeem id → a SECOND consume. Reproduced deterministically at A=200,
B=200, consume=2.

Move mutual exclusion before identity resolution: resetKeeperQuota now requires the state row to
already carry a confirmed account_id and takes tryLockAccountID on that STORED account_id BEFORE
resolveKeeperResetIdentity, holding it across resolve → lookup → fetch → consume → cooldown →
finish. The fresh account_id must then equal the stored (fenced) account_id or it fails closed.
This supersedes the previous NULL-bind-on-reset (bindKeeperAccountID removed): a legacy
pre-account_id (NULL) row is now refused (fail closed) until it has been inspected once, which
persists its account_id — the reviewer's minimal safe option — rather than fencing on an unknown
identity. 6e278b0d (duplicate-row snapshot staleness) stays documented-only per review.

Tests: TestKeeperResetAccountFenceBeforeResolve gates fileA inside its DOWNLOAD (pre-fetch) while
it holds the fence and asserts fileB (same account) is refused 409 with exactly one consume;
TestKeeperResetNullAccountRefusedUntilInspected asserts a NULL-account row is refused before any
remote call. i18n adds the "identity not confirmed" message with a smoke assertion; the old
bind-failure messages/tests are removed.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

* docs(codex-keeper): correct the per-auth_name lock comment to per-file/state guard

Comment-only, no behavior change. The resetKeeperQuota lock comment claimed the per-auth lock is
THE paid-resource guard that stops two same-account resets from each consuming a credit. That
holds only for the same auth_name; the cross-file monetary mutual exclusion is the
stored-account_id fence taken afterward. Reworded to describe the per-auth_name lock as a
per-file/state guard and point to the account_id fence for the cross-route paid-credit exclusion,
so the mutex boundary is not misused later.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>

---------

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>
Co-authored-by: feiniu (Raft agent) <admin@oranix.io>
Add provider-dispatched Antigravity quota inspection, identity-bound snapshot persistence, and localized quota presentation across Keeper account surfaces.

Reviewed source exact:
- base: da4484f
- head: bdb03e9
- tree: 1936ce8

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>
Signed-off-by: 怪味胡豆 <raft-mobile-guaiweihudou@mail.build>
API keys minted by CPA-Helper were hardcoded to `sk-<random>`. Add an admin
setting `api_key_prefix` so a deployment can brand its keys (ours will use
`sk-cortex`, yielding `sk-cortex-...`). A generated key is
`<prefix>-<52 random alphanumerics>`.

- Default stays `sk`, so untouched deployments keep generating `sk-...` keys
  byte-for-byte as before; blank input resets to the default.
- Only NEW keys use the prefix. Existing keys are never rewritten and keep
  working (tested: the legacy key stays listed with its original value).
- Validation (422): letters, digits, `-` and `_`; must start with a letter or
  digit and must not end with `-` (the generator adds the joining dash); max
  32 chars; whitespace trimmed; rejected values do not persist.
- Migration 202609160001 adds `app_settings.api_key_prefix` (Down drops it;
  a rollback only loses the configured prefix). LatestVersion, the `migrate
  down-to` CLI test head, the Up/Down/replay schema test and the rollback
  runbook are updated.
- The CLIProxyAPI key-sync path passes keys as opaque strings and makes no
  prefix assumption (verified).
- Frontend: new "API KEY 设置 / API Key Settings" section with a live preview
  and rules; key masking is now prefix-aware — the random secret never
  contains `-`, so the last `-` separates any multi-segment prefix from the
  secret (`sk-cortex-` stays readable, >= 8 secret chars always masked).
  Extracted to a pure maskApiKey helper with a smoke test; i18n for the two
  validation messages plus smoke assertions.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>
Co-authored-by: feiniu (Raft agent) <a-9b3ff9ce@mail.build>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…#15)

The frontend smoke scripts (i18n, Antigravity countdown/window/quota-format,
quota-exhaustion, API-key masking) were only run by hand. Add an aggregate
`test:smoke` script and run `npm run lint` + `npm run test:smoke` before the
frontend build in both the release workflow and the Dockerfile frontend stage,
so a regression blocks the release. No runtime behaviour change.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>
Co-authored-by: feiniu (Raft agent) <a-9b3ff9ce@mail.build>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…l prices (#16)

Usage from reverse-proxied accounts is reported under the proxy's provider
name and often a reasoning-tier suffix (provider `antigravity`, model
`gemini-3.8-flash-high`), while the price list synced from LiteLLM only knows
the canonical vendor entry (provider `gemini`, model
`gemini/gemini-3.8-flash`). Price lookup was exact-match only, so these models
were never priced.

findMatchingPrice now falls back through ordered candidates, most specific
first; an exact (provider, model) price always wins, so a manual price can
override any association:
- model: the name itself, then progressively stripped variant suffixes
  (-thinking, -minimal, -low, -medium, -high, -xhigh), which do not change the
  per-token price;
- provider: itself, then its canonical vendor (codex→openai, claude→anthropic,
  gemini-cli/aistudio→gemini, antigravity→gemini/anthropic/openai by model
  family);
- each pair is also tried with LiteLLM's `<provider>/<model>` key form.

No schema change. Cost is computed at query time, so existing history is priced
as soon as this ships. All callers go through findMatchingPrice (usage cost,
quota charging, model catalog), so they stay consistent.

Signed-off-by: feiniu <a-9b3ff9ce@mail.build>
Co-authored-by: feiniu (Raft agent) <a-9b3ff9ce@mail.build>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…vider/endpoint/source-account) (#19)

* feat(cli): read-only usage-cost subcommand reporting cost grouped by model/provider/endpoint/source-account

Prices records with the production recordCost derivation via a narrow
export so the report cannot drift from what was recorded; unpriced
records are reported separately and never read as free. --db defaults
to the service database path, overridable for analysis of other copies.
Golden vectors pin the wired cost function against production.

* fix(cli): bind dbTime-layout string for usage-cost --since window

Review finding (PR #19): LoadRecords bound since.UTC() as time.Time; the
driver serialises it space-separated UTC while production writes timestamps
as 'T'-separated Asia/Shanghai text, and ' ' < 'T' made the lexicographic
>= admit rows up to ~16h older than --since. Export app.UsageDBTime so the
reader binds the same byte shape production writes, and make fixtures write
timestamps through the same helper so the test exercises the real byte
layout instead of the reader's own.

---------

Co-authored-by: feiniu <a-9b3ff9ce@mail.build>
Co-authored-by: feiniu (Raft agent) <admin@oranix.io>
* feat(account-runway): read-only per-account quota runway report

* fix(usagecost): add busy_timeout to read-only DSN for delete-journal DB

Same one-line class as accountrunway: production runs journal_mode=delete
with frequent writer traffic, and a bare read-only open can hit
SQLITE_BUSY(5) during a writer transaction. Wait up to 5s for the lock.

---------

Co-authored-by: feiniu (Raft agent) <admin@oranix.io>
@bytemain

Copy link
Copy Markdown
Author

Closing: opened against the wrong repository (upstream public repo). Intended for the private fork.

@bytemain bytemain closed this Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants