Skip to content

Commit be7a36e

Browse files
feat(model): per-model context limits and pricing, generated from Cursor's docs (#89)
* feat(model): per-model context window limits from Cursor docs * feat(model): per-model pricing from Cursor docs for TUI cost display * feat(model): per-model output limits for high-output frontier models * fix(model): emit limits and cost on the config channel opencode reads * feat(model): generate context limits and pricing from Cursor docs * docs(model): point MODEL_OUTPUT_LIMITS edits at the generator template * fix(model): make the model-data drift check unable to pass without checking The drift mechanism had three ways to report success without having verified anything, and the function holding every hard contract had no test coverage. - Remove the fail-open CLI guard. `isInvokedDirectly()` compared `process.argv[1]` against `import.meta.url`; on any invocation where they differ, `main()` never ran and the process exited 0 having done nothing. The symlink fix addressed one instance, not the class. The generator module is now import-pure and `scripts/sync-model-limits-cli.mjs` calls `main()` unconditionally. Verified through a symlink, an absolute path from another cwd, and `npm exec`. - The CI step now captures stdout and requires the run summary line, so "exited 0 having done nothing" fails instead of passing for free. - An empty price cell no longer becomes $0. `parseDocsTable` leaves an absent cell unset (distinct from present-but-empty) and throws when a requested column is missing from a row; `parsePrice("")` throws. `"-"` still means $0, which is what Cursor documents it as. - `generate()` takes injectable `modelIds`/`overrides`, so strict overrides, ambiguity on both tables, docs-over-override precedence and the deterministic sort are exercised against the fixtures rather than resting on one-time manual probes. `normalizeForComparison` is covered too: it is what keeps `--check` from failing daily. - Matching consults the `Provider` column. `claude`/`gpt` are dropped from both sides to survive word-order differences, which also discarded vendor identity; an id naming a dropped vendor now matches only a row whose Provider agrees. Counts unchanged: 29 exact / 0 ambiguous on context, 27 exact / 6 overridden on pricing. - Fixtures gain the verbatim `Claude Opus 4.7 (fast mode)` row ($30/$150, 6x Opus's rate) and assert `claude-opus-4-7` refuses it. - Write-mode I/O failure exits 2, not 1: a failed write does not establish that the committed file is stale. - `Synced:` relabelled `Data last changed:`. Write mode short-circuits when only the date would move, so the committed date was never a verification date and read as though the file were a year stale. - `schedule`/`workflow_dispatch` are workflow-wide triggers, so `build` and `integration` are now scoped to push/pull_request. - `SOURCES.context` points at the URL that currently serves markdown. The `.md` form began returning HTTP 404, which made every run exit 2 — the scheduled job's exit-2-is-a-warning branch would have swallowed that indefinitely. No emitted value changes. Zero numeric values change in the regenerated `src/model-limits.ts`: key sets, key order, `DEFAULT_*`, `MODEL_OUTPUT_LIMITS` and all three `resolve*` functions are byte-identical. The only diff is the header. Tests 416 -> 442. Each new test was mutation-checked against the pre-change behaviour to confirm it can fail. * fix(model): use .md doc URLs and fail the drift job on unverified Two corrections to the drift check: - SOURCES.context pointed at the extensionless docs URL, justified by a comment claiming the .md form returned 404. Not reproducible: measured against both URL shapes, the .md form returns the markdown table under both a markdown-preferring and a wildcard Accept header, while the extensionless form returns a ~110KB HTML page under a wildcard Accept. The .md form does not depend on the Accept header staying markdown-preferring, so both sources now use it. Comment corrected to the measured behavior. - Exit 2 (docs unreachable/unparseable, or a model id matching nothing) annotated a warning and exited 0, so a permanently broken check would pass every week indefinitely. It now fails the job: unverified is not the same as verified-clean. Regenerated output changes one line, the source URL recorded in the generated header. No numeric value changed.
1 parent 4f49156 commit be7a36e

13 files changed

Lines changed: 1730 additions & 5 deletions

‎.github/workflows/ci.yml‎

Lines changed: 57 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,16 +1,25 @@
11
name: CI
22

3+
# `schedule` and `workflow_dispatch` exist for the model-data-drift job only.
4+
# Triggers are workflow-wide in Actions, so every job below carries an `if:`
5+
# that scopes it to the events it is actually for.
36
on:
47
push:
58
branches: ["**"]
69
pull_request:
10+
schedule:
11+
- cron: "0 9 * * 1"
12+
workflow_dispatch:
713

814
permissions:
915
contents: read
1016

1117
jobs:
1218
build:
1319
name: typecheck · test · build (node ${{ matrix.node-version }})
20+
# Code CI only. The weekly cron and manual dispatch exist for
21+
# model-data-drift; re-running the matrix on them buys nothing.
22+
if: github.event_name == 'push' || github.event_name == 'pull_request'
1423
runs-on: ubuntu-latest
1524
strategy:
1625
fail-fast: false
@@ -45,8 +54,56 @@ jobs:
4554
- name: Build
4655
run: npm run build
4756

57+
model-data-drift:
58+
name: model limits · drift vs Cursor docs
59+
# Deliberately not on pull_request: this job reaches cursor.com, and PR CI
60+
# stays hermetic.
61+
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
62+
runs-on: ubuntu-latest
63+
steps:
64+
- uses: actions/checkout@v7
65+
66+
- name: Set up Node
67+
uses: actions/setup-node@v7
68+
with:
69+
node-version: "24.x"
70+
cache: npm
71+
72+
- name: Install dependencies
73+
run: npm ci
74+
75+
- name: Check committed model limits against Cursor docs
76+
run: |
77+
log="$RUNNER_TEMP/drift.log"
78+
set +e
79+
npm run sync:model-limits -- --check 2>&1 | tee "$log"
80+
code=${PIPESTATUS[0]}
81+
set -e
82+
if [ "$code" = "1" ]; then
83+
echo "::error::src/model-limits.ts is stale. Run 'npm run sync:model-limits' and commit."
84+
exit 1
85+
fi
86+
if [ "$code" = "2" ]; then
87+
echo "::error::Could not verify model data (docs unreachable, unparseable, or a model id matched nothing). The drift check did not run — treat this as unverified, not as passing."
88+
exit 1
89+
fi
90+
if [ "$code" != "0" ]; then
91+
echo "::error::Unexpected exit code $code from the drift check."
92+
exit "$code"
93+
fi
94+
# Exit 0 is only trustworthy if the run also reported its summary. A
95+
# bare 0 is what "did nothing at all" looks like, and that has
96+
# happened: an entry-point guard once skipped main() entirely and the
97+
# job passed for free. Demand the evidence, not just the code.
98+
if ! grep -qE 'sync-model-limits: src/model-limits\.ts is up to date \(context: [0-9]+ from docs, [0-9]+ overridden \| cost: [0-9]+ from docs, [0-9]+ overridden \| [0-9]+ model ids\)' "$log"; then
99+
echo "::error::Drift check exited 0 without printing a run summary — it did not actually verify anything."
100+
exit 1
101+
fi
102+
48103
integration:
49104
name: e2e · opencode loads plugin & lists models
105+
# Code CI only, same reason as `build`.
106+
if: github.event_name == 'push' || github.event_name == 'pull_request'
50107
runs-on: ubuntu-latest
51108
steps:
52109
- uses: actions/checkout@v7

‎package.json‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -53,6 +53,7 @@
5353
"test": "vitest run",
5454
"test:watch": "vitest",
5555
"test:e2e": "vitest run --config vitest.e2e.config.ts --passWithNoTests",
56+
"sync:model-limits": "node scripts/sync-model-limits-cli.mjs",
5657
"prepublishOnly": "npm run typecheck && npm test && npm run build"
5758
},
5859
"dependencies": {

‎scripts/sync-model-limits-cli.mjs‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,17 @@
1+
#!/usr/bin/env node
2+
/**
3+
* CLI entry point for `scripts/sync-model-limits.mjs`.
4+
*
5+
* This file exists so the generator module stays import-pure (the tests import
6+
* it) without needing an "am I the entry point?" guard inside it. Such a guard
7+
* — comparing `process.argv[1]` against `import.meta.url` — is fail-open: on
8+
* any invocation where the two differ (a symlinked path, a wrapper, an exec
9+
* shim) `main()` never runs and the process exits 0 having done nothing, which
10+
* makes the scheduled drift check permanently green and permanently useless.
11+
* That already happened once, via a symlinked `/tmp` path on macOS.
12+
*
13+
* There is no guard here. Running this file always runs `main()`.
14+
*/
15+
import { main } from "./sync-model-limits.mjs";
16+
17+
process.exitCode = await main(process.argv.slice(2));

‎scripts/sync-model-limits.d.mts‎

Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,56 @@
1+
/**
2+
* Types for `sync-model-limits.mjs`. The script is plain ESM JavaScript (it
3+
* runs via `node` with no build step), but `tsconfig.json` includes `test`, so
4+
* `test/sync-model-limits.test.ts` needs a declaration to import it.
5+
*
6+
* This mirror is hand-maintained. `allowJs: true` would remove the need for it,
7+
* but it does not work in this repo: it pulls `src/sidecar/agent-host.mjs` into
8+
* the program, and that file assigns to `console.log`, which strips `log`,
9+
* `debug`, `info`, `warn`, and `error` off the global `Console` type and breaks
10+
* 30 checks in existing `.ts` files. Measured, not assumed — see
11+
* `.superpowers/sdd/task-6-report.md`. Two things keep this file honest in the
12+
* meantime: `test/sync-model-limits.test.ts` asserts the module's runtime
13+
* export names match the list declared here, and the tests call every declared
14+
* signature, so a parameter that is declared but missing (or vice versa) fails
15+
* `npm run typecheck`.
16+
*/
17+
18+
export type DocsRow = Record<string, string>;
19+
20+
export type ModelCost = {
21+
input: number;
22+
output: number;
23+
cacheRead: number;
24+
cacheWrite: number;
25+
};
26+
27+
export type MatchResult =
28+
| { row: DocsRow; ambiguous?: never }
29+
| { row?: never; ambiguous: string[] };
30+
31+
export declare const SOURCES: { context: string; pricing: string };
32+
export declare const MODEL_IDS: string[];
33+
export declare const OVERRIDES: Record<
34+
string,
35+
{ context?: number; cost?: ModelCost; why: string }
36+
>;
37+
38+
export declare function parseDocsTable(md: string, columnNames: string[]): DocsRow[];
39+
export declare function parseTokens(text: string): number | undefined;
40+
export declare function parsePrice(text: string): number;
41+
export declare function matchModelId(id: string, docRows: DocsRow[]): MatchResult | undefined;
42+
export declare function normalizeForComparison(text: string): string;
43+
export declare function generate(input: {
44+
contextMd: string;
45+
pricingMd: string;
46+
modelIds?: readonly string[];
47+
overrides?: Record<string, { context?: number; cost?: ModelCost; why?: string }>;
48+
date?: string;
49+
}): {
50+
text: string;
51+
stats: {
52+
context: { matched: number; overridden: number };
53+
cost: { matched: number; overridden: number };
54+
};
55+
};
56+
export declare function main(argv: string[]): Promise<number>;

0 commit comments

Comments
 (0)