You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
tests: the styling gate covers run, trace and metrics (#126)
Closes#17. The contract — strip the escapes from what a terminal receives and
it equals what a pipe receives, exactly — was gated for four commands: `plan`,
`models`, `models --check`, `demo stage0`. The commands most likely to be piped
were the uncovered ones: `run --check-only` is a linter in this repository's own
CI, `trace | grep` and `metrics` in a script are the obvious uses, and `run`'s
REFUSED block tints the key *and* the value of its verdict line, a shape no
covered case reached.
Four cases added, and the gap is demonstrated rather than asserted. Painting
`trace`'s step column unconditionally leaves the old four-case gate **fully
green** while both new `trace` cases fail. That is the whole claim of the issue,
reproduced.
A case is a builder rather than a literal argv now, because these commands take
a file argument. The `workspace` fixture writes an admitted topology, a refused
one carrying two objections so the per-objection row runs more than once, and a
trace from one `demo stage0` run.
The hermetic argv matters more than it looks. `run` resolves its registry and
policy from the working directory, and this repository's root carries a
`registry.py` and a `grapharc.toml` that are **gitignored** — dogfooding
residue. A case leaning on those would read one registry locally and another in
CI and then compare output that differs for reasons unrelated to styling. So the
registry is named explicitly and `--config` points at an empty file.
One normaliser added, under protest. `run`'s fingerprint is not stable across
two loads of the same topology: `Subgraph.proposal_id` defaults to a fresh
`uuid4` and `fingerprint()` hashes the whole model, that field included. So the
same topology file yields a different fingerprint every invocation, while
`graphrun.py` prints it under the comment "the fingerprint is what a later run
is compared against". Normalising it keeps the ADMITTED block's styling in the
comparison instead of dropping `run` over one token; the underlying problem is
filed separately, and when it is fixed this normaliser should go and the
comparison gets stricter for it.
Verified: 2205 selected, 13 deselected, ruff clean.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copy file name to clipboardExpand all lines: docs/deep-dive.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -254,7 +254,7 @@ A stable system is not one that claims to have no edges — it is one whose edge
254
254
-**`.env` and `grapharc.toml` follow the same discovery rule: the working directory, and nowhere else.** Neither searches parent directories — a run must not be governed by a file you did not know about, and must not be *billed* to one either. **This is a behaviour change:** the credential loader used to walk up to `/`, so a `.env` in an ancestor directory (a `$HOME` one on a shared box, a client project one above a demo checkout) was picked up silently. If you relied on that, move the file into the directory you run from, `export` the variable, or pass `env_file=` to name it explicitly. A real environment variable still beats any file.
255
255
-**`grapharc run` has no budget unless you give it one.** Set any of `--max-tokens`, `--max-iterations`, `--max-seconds`, or `--max-concurrency`; without them each dimension is unlimited and the gate admits a topology of any worst-case cost.
256
256
257
-
**Verified this pass:**`pytest` → green, 2,197 selected and 13 deselected (the live ones); `ruff check .` clean; all eight `grapharc demo` stages green, plus the `trace` / `metrics` / `viz` / `replay` tour against a freshly recorded demo trace; the wheel builds and imports all submodules in a clean virtualenv with `[all]`, and `0.1.8` on PyPI is that wheel. The counts are a snapshot, not a property of the project — `pytest` re-derives them in one command, which is the only reason they are quoted, and `tests/test_deep_dive.py` fails this line rather than letting it drift.
257
+
**Verified this pass:**`pytest` → green, 2,205 selected and 13 deselected (the live ones); `ruff check .` clean; all eight `grapharc demo` stages green, plus the `trace` / `metrics` / `viz` / `replay` tour against a freshly recorded demo trace; the wheel builds and imports all submodules in a clean virtualenv with `[all]`, and `0.1.8` on PyPI is that wheel. The counts are a snapshot, not a property of the project — `pytest` re-derives them in one command, which is the only reason they are quoted, and `tests/test_deep_dive.py` fails this line rather than letting it drift.
258
258
259
259
[ROADMAP.md](../ROADMAP.md) tracks what is built and what is not, item by item.
0 commit comments