LoopFormer studies recurrent computation and adaptive inference depth in Qwen/Qwen2.5-0.5B-Instruct.
Can a recurrent language model learn useful iterative computation, and can inference determine when additional recurrent computation is useful, unnecessary, or harmful?
Reusing a learned block lets inference vary computational depth without adding a separate set of weights for every step. The project first tests whether recurrence executes an interpretable algorithm, then studies depth generalization, overscaling dynamics, and eventually compute-quality tradeoffs through stopping. Pointer chasing is the controlled starting environment: fresh rules live in each prompt, and exact intermediate targets make errors and progress checkable.
prompt
↓
frozen prelude
↓
shared recurrent block × T
↓
frozen coda
↓
readout
One recurrent block is reused across all loops. The current experiment trains all weights in that block while keeping the prelude and coda frozen; older experiments used LoRA. The executor’s prompt context stays fixed across loops; the frozen coda reads each loop's state for supervision and evaluation. The next loop consumes the recurrent hidden state, rather than the coda output or a decoded answer. See architecture.
The seed-61 executor audit finds strong execution beyond trained depth on a limited set of 26-state graphs. The controller diagnostic isolates poor stopping despite accessible initial count information and successful tiny fits. The completed controller-only run reaches 95.3% exact stopping on trained counts, but 45.8% on held-out counts 7/9/11 and zero beyond count 13. R remains frozen; controller generalization is still unresolved.
The matched remaining-work comparison improves the controller's numerical readout but leaves stopping generalization essentially unchanged. Unfamiliar counts are already misestimated before looping; the next step is to audit training fit and supervision before changing the model.
Run the read-only training audit on the desktop with the existing cached features and saved controller weights:
git pull --ff-only
bash audit_controller.shThis evaluates all six runs, compares training and validation errors, and measures
loss gradients without updating weights or loading Qwen. Results and logs go to
eval/pointer_diagnostics/ and W&B. See the audit guide
for required artifacts, metrics and options.
Start with the research plan, current status, and documentation index. Historical experiments remain in the linked reports; development results are not independent confirmation.
Use Python 3.11 and run commands from the repository root. Windows/NVIDIA users should follow the WSL2/CUDA setup guide first; other environment details are in setup.
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python -m pip checkDownload the base revision used for the reproducible experiments:
hf download Qwen/Qwen2.5-0.5B-Instruct --revision 7ae557604adf67be50417f59c2c2f167def9a775Check ordinary model loading and generation:
python scripts/smoke_test_qwen.pyActivate .venv in each new terminal. The commands below default to CPU unless a device is supplied. Use --device mps on Apple Silicon or --device cuda on the configured NVIDIA desktop. For deterministic CUDA runs, set:
export CUBLAS_WORKSPACE_CONFIG=:4096:8Dataset code lives in scripts/dataset/. Preview five examples without writing files:
python -m scripts.dataset --seed 17 --dry-runUse --dry-run 10 for ten examples. Reproduce the existing dataset with master seed 17, the pinned tokenizer above, and this complete configuration:
python -m scripts.dataset \
--seed 17 \
--train-count 10000 \
--validation-count 1000 \
--test-count 1000 \
--depth-test-count 1000 \
--min-depth 1 \
--max-train-depth 8 \
--max-eval-depth 16 \
--output data/pointer/seed-17
python -m scripts.dataset --verify data/pointer/seed-17The existing seed-17 dataset already reaches depth 16 for development evaluation. The current executor uses the separate seed-61 dataset created by bash prepare_executor.sh; see its run instructions. Generated files are excluded from Git, so use the reproduction command above only if the dataset is missing on a new machine.
File under data/pointer/seed-17/ |
Examples | Depths | Examples per depth |
|---|---|---|---|
train.jsonl |
10,000 | 1–8 | 1,250 |
validation.jsonl |
1,000 | 1–8 | 125 |
test.jsonl |
1,000 | 1–8 | 125 |
depth_test.jsonl |
1,000 | 9–16 | 125 |
The current seed-61 run has 36,000 training questions at requested counts 1–6, 8, 10 and 12; its deep development queries span 13–64. See dataset and training details.
Inspect three ordinary-model questions before running the full test:
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-Instruct --test
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-InstructThe ordinary baseline uses three-shot examples at depths 1, 2, and 3, defined in prompts/pointer_task.txt. --model also accepts a saved model directory. Models load locally by default; --download permits missing downloads.
Progress and throughput appear in the terminal. Predictions and summaries go to eval/pointer_task/. See baseline evaluation for loading and scoring options.
On the CUDA desktop, prepare the new depth-independent dataset and check the isolated executor pipeline:
bash prepare_executor.sh --dry-run
bash prepare_executor.sh
bash train_executor.sh --dry-run
bash train_executor.sh --smoke-testAfter inspecting the smoke results, run bash train_executor.sh, then
bash eval_executor.sh. The default trains all recurrent-block weights with an
isolated controller and per-loop supervision. Full-model and LoRA controls are
also available. See the current run instructions
for configuration, exact data reproduction and output locations.
A compact Rich dashboard shows progress, ETA, losses and memory. W&B uses project
loopformer. Checkpoints and logs go to models/stage1_pointer/; model binaries
are excluded from Git. CUDA fit, speed and quality need the new desktop smoke/run.
bash eval_executor.sh runs matched-count/precision diagnostics, full-loop evaluation and actual learned stopping on development data. Results and terminal logs go under eval/pointer_diagnostics/. The older depth-12 scripts remain available for historical runs.
Pass a complete saved step directory, including its adapter weights and tokenizer:
python -m scripts.eval.loop_test \
--model models/stage1_pointer/executor_r-seed61/step-000750 \
--data data/pointer/seed-61-independent/validation.jsonl \
--device cuda --loops 12 --testThis requires the complete checkpoint on the training desktop; local metadata alone is insufficient. Remove --test for all 1,536 validation queries. Outputs go to eval/pointer_loops/. This reads the model after every recurrent loop; the ordinary three-shot prompt is not used. See full-loop evaluation for commands and metrics.
Preview the separate absorbing-terminal task variant without loading a model or writing files:
python -m scripts.eval.overscaling_test --dry-runCheckpoint sweeps, repair/damage scoring, and survival exports are available but running them is deferred until further Stage 1 training and review. See overscaling usage.