Skip to content

Latest commit

 

History

History
94 lines (72 loc) · 3.61 KB

File metadata and controls

94 lines (72 loc) · 3.61 KB

ECEB adaptive context study v0.2

Status: experimental result; not a completeness claim
Date: 2026-08-03

Result

This follow-up adds SACP-0.2 adaptive context to the same 30 pinned Multi-SWE-bench Rust tasks and openai/gpt-4.1-mini file-localization profile used by the v0.1 study. It retains the original control observations and adds one new adaptive response per task.

Strategy Mean prompt tokens Token savings Evidence recall Fix-path recall Recorded localization spend
Full context 380,856 0.00% 100.00% 53.32% $4.5755
Lexical BM25 7,677 95.82% 58.55% 46.87% $0.0974
Fixed compiled 5,138 97.58% 51.85% 41.76% $0.0668
Adaptive compiled 4,345 97.80% 51.85% 45.09% $0.0573

Relative to fixed compiled context, the adaptive arm selected 15.44% fewer mean prompt tokens, improved exact fix-path recall by 3.33 percentage points, and reduced recorded localization spend by 14.23%. Evidence recall was unchanged.

Adaptive context still trailed lexical BM25 fix-path recall by 1.78 percentage points and full context by 8.23 points. The result therefore improves the efficiency/quality tradeoff of the compiled arm but does not establish quality parity.

Adaptive behavior

The preregistered stopping policy used:

  • 50% query-term coverage;
  • at least two evidence excerpts;
  • token ceilings from 1,000 through 8,000 in 1,000-token steps; and
  • node ceilings from 2 through 16 in two-node steps.

Long queries whose packet kernel exceeded 1,000 tokens began at the minimum kernel size while retaining the 8,000-token hard maximum.

Stop outcome Tasks
Estimated sufficient 12
Hard limit reached 18
Candidate exhausted 0

The mean selected round was 7.03. Four tasks stopped at round 4, three at round 5, two at round 6, and twenty-one selected round 8. Every observation records the complete ladder, source paths, prompt counts, context digests, matched and missing terms, and stop reason.

Method and claim boundary

The 50% threshold was selected by the exploratory sweep documented in adaptive-calibration.md. Because calibration and evaluation use the same 30 tasks, this is not an independent estimate of generalization. The model was sampled once per task and arm, and the benchmark measures file localization rather than patch correctness.

This result supports a narrower statement:

On this fixed run, adaptive stopping improved the fixed compiled arm's token, cost, and localization-recall measurements without changing its evidence recall.

It does not show that the sufficiency estimator detects complete evidence, that adaptive context is universally better than BM25, or that answer quality is preserved on other corpora or models.

Reproduction

The raw observations and report are checked into benchmarks/results/eceb-multiswe-rust-adaptive-v0.2/. Replay all schemas, metrics, context-trace invariants, and negative selected-round and premature-hard-limit tests without API access:

scripts/eceb-adaptive-study-check.sh

A fresh provider run requires OpenRouter credentials and spends API credit:

cargo run --features openrouter --bin symgliph-eceb-study -- run \
  benchmarks/eceb-multiswe-rust-v0.1.json \
  --strategies full,lexical,compiled,adaptive \
  --budget-tokens 8000 \
  --top-k 16 \
  --adaptive-min-query-term-coverage-bps 5000 \
  --localizer-model openai/gpt-4.1-mini \
  --localizer-output-tokens 1024 \
  --env-file /path/to/openrouter.env