Status: experimental result; not a completeness claim
Date: 2026-08-03
This follow-up adds SACP-0.2 adaptive context to the same 30 pinned
Multi-SWE-bench Rust tasks and openai/gpt-4.1-mini file-localization profile
used by the v0.1 study. It retains the original control observations and adds
one new adaptive response per task.
| Strategy | Mean prompt tokens | Token savings | Evidence recall | Fix-path recall | Recorded localization spend |
|---|---|---|---|---|---|
| Full context | 380,856 | 0.00% | 100.00% | 53.32% | $4.5755 |
| Lexical BM25 | 7,677 | 95.82% | 58.55% | 46.87% | $0.0974 |
| Fixed compiled | 5,138 | 97.58% | 51.85% | 41.76% | $0.0668 |
| Adaptive compiled | 4,345 | 97.80% | 51.85% | 45.09% | $0.0573 |
Relative to fixed compiled context, the adaptive arm selected 15.44% fewer mean prompt tokens, improved exact fix-path recall by 3.33 percentage points, and reduced recorded localization spend by 14.23%. Evidence recall was unchanged.
Adaptive context still trailed lexical BM25 fix-path recall by 1.78 percentage points and full context by 8.23 points. The result therefore improves the efficiency/quality tradeoff of the compiled arm but does not establish quality parity.
The preregistered stopping policy used:
- 50% query-term coverage;
- at least two evidence excerpts;
- token ceilings from 1,000 through 8,000 in 1,000-token steps; and
- node ceilings from 2 through 16 in two-node steps.
Long queries whose packet kernel exceeded 1,000 tokens began at the minimum kernel size while retaining the 8,000-token hard maximum.
| Stop outcome | Tasks |
|---|---|
| Estimated sufficient | 12 |
| Hard limit reached | 18 |
| Candidate exhausted | 0 |
The mean selected round was 7.03. Four tasks stopped at round 4, three at round 5, two at round 6, and twenty-one selected round 8. Every observation records the complete ladder, source paths, prompt counts, context digests, matched and missing terms, and stop reason.
The 50% threshold was selected by the exploratory sweep documented in
adaptive-calibration.md. Because calibration and
evaluation use the same 30 tasks, this is not an independent estimate of
generalization. The model was sampled once per task and arm, and the benchmark
measures file localization rather than patch correctness.
This result supports a narrower statement:
On this fixed run, adaptive stopping improved the fixed compiled arm's token, cost, and localization-recall measurements without changing its evidence recall.
It does not show that the sufficiency estimator detects complete evidence, that adaptive context is universally better than BM25, or that answer quality is preserved on other corpora or models.
The raw observations and report are checked into
benchmarks/results/eceb-multiswe-rust-adaptive-v0.2/.
Replay all schemas, metrics, context-trace invariants, and negative
selected-round and premature-hard-limit tests without API access:
scripts/eceb-adaptive-study-check.shA fresh provider run requires OpenRouter credentials and spends API credit:
cargo run --features openrouter --bin symgliph-eceb-study -- run \
benchmarks/eceb-multiswe-rust-v0.1.json \
--strategies full,lexical,compiled,adaptive \
--budget-tokens 8000 \
--top-k 16 \
--adaptive-min-query-term-coverage-bps 5000 \
--localizer-model openai/gpt-4.1-mini \
--localizer-output-tokens 1024 \
--env-file /path/to/openrouter.env