Skip to content

Latest commit

 

History

84 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LoopFormer

LoopFormer studies recurrent computation and adaptive inference depth in Qwen/Qwen2.5-0.5B-Instruct.

Can a recurrent language model learn useful iterative computation, and can inference determine when additional recurrent computation is useful, unnecessary, or harmful?

Reusing a learned block lets inference vary computational depth without adding a separate set of weights for every step. The project first tests whether recurrence executes an interpretable algorithm, then studies depth generalization, overscaling dynamics, and eventually compute-quality tradeoffs through stopping. Pointer chasing is the controlled starting environment: fresh rules live in each prompt, and exact intermediate targets make errors and progress checkable.

prompt
  ↓
frozen prelude
  ↓
shared recurrent block × T
  ↓
frozen coda
  ↓
readout

One recurrent block is reused across all loops. The current experiment trains all weights in that block while keeping the prelude and coda frozen; older experiments used LoRA. The executor’s prompt context stays fixed across loops; the frozen coda reads each loop's state for supervision and evaluation. The next loop consumes the recurrent hidden state, rather than the coda output or a decoded answer. See architecture.

Current result and diagnostics

The seed-61 executor audit finds strong execution beyond trained depth on a limited set of 26-state graphs. The controller diagnostic isolates poor stopping despite accessible initial count information and successful tiny fits. The completed controller-only run reaches 95.3% exact stopping on trained counts, but 45.8% on held-out counts 7/9/11 and zero beyond count 13. R remains frozen; controller generalization is still unresolved.

The matched remaining-work comparison improves the controller's numerical readout but leaves stopping generalization essentially unchanged. Unfamiliar counts are already misestimated before looping; the next step is to audit training fit and supervision before changing the model.

Run the read-only training audit on the desktop with the existing cached features and saved controller weights:

git pull --ff-only
bash audit_controller.sh

This evaluates all six runs, compares training and validation errors, and measures loss gradients without updating weights or loading Qwen. Results and logs go to eval/pointer_diagnostics/ and W&B. See the audit guide for required artifacts, metrics and options.

Start with the research plan, current status, and documentation index. Historical experiments remain in the linked reports; development results are not independent confirmation.

Setup

Use Python 3.11 and run commands from the repository root. Windows/NVIDIA users should follow the WSL2/CUDA setup guide first; other environment details are in setup.

python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python -m pip check

Download the base revision used for the reproducible experiments:

hf download Qwen/Qwen2.5-0.5B-Instruct --revision 7ae557604adf67be50417f59c2c2f167def9a775

Check ordinary model loading and generation:

python scripts/smoke_test_qwen.py

Activate .venv in each new terminal. The commands below default to CPU unless a device is supplied. Use --device mps on Apple Silicon or --device cuda on the configured NVIDIA desktop. For deterministic CUDA runs, set:

export CUBLAS_WORKSPACE_CONFIG=:4096:8

Original pointer dataset: preview and reproduce

Dataset code lives in scripts/dataset/. Preview five examples without writing files:

python -m scripts.dataset --seed 17 --dry-run

Use --dry-run 10 for ten examples. Reproduce the existing dataset with master seed 17, the pinned tokenizer above, and this complete configuration:

python -m scripts.dataset \
  --seed 17 \
  --train-count 10000 \
  --validation-count 1000 \
  --test-count 1000 \
  --depth-test-count 1000 \
  --min-depth 1 \
  --max-train-depth 8 \
  --max-eval-depth 16 \
  --output data/pointer/seed-17

python -m scripts.dataset --verify data/pointer/seed-17

The existing seed-17 dataset already reaches depth 16 for development evaluation. The current executor uses the separate seed-61 dataset created by bash prepare_executor.sh; see its run instructions. Generated files are excluded from Git, so use the reproduction command above only if the dataset is missing on a new machine.

File under data/pointer/seed-17/ Examples Depths Examples per depth
train.jsonl 10,000 1–8 1,250
validation.jsonl 1,000 1–8 125
test.jsonl 1,000 1–8 125
depth_test.jsonl 1,000 9–16 125

The current seed-61 run has 36,000 training questions at requested counts 1–6, 8, 10 and 12; its deep development queries span 13–64. See dataset and training details.

Evaluate a model

Inspect three ordinary-model questions before running the full test:

python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-Instruct --test
python -m scripts.eval.naive_test --model Qwen/Qwen2.5-0.5B-Instruct

The ordinary baseline uses three-shot examples at depths 1, 2, and 3, defined in prompts/pointer_task.txt. --model also accepts a saved model directory. Models load locally by default; --download permits missing downloads.

Progress and throughput appear in the terminal. Predictions and summaries go to eval/pointer_task/. See baseline evaluation for loading and scoring options.

Train pointer execution

On the CUDA desktop, prepare the new depth-independent dataset and check the isolated executor pipeline:

bash prepare_executor.sh --dry-run
bash prepare_executor.sh
bash train_executor.sh --dry-run
bash train_executor.sh --smoke-test

After inspecting the smoke results, run bash train_executor.sh, then bash eval_executor.sh. The default trains all recurrent-block weights with an isolated controller and per-loop supervision. Full-model and LoRA controls are also available. See the current run instructions for configuration, exact data reproduction and output locations.

A compact Rich dashboard shows progress, ETA, losses and memory. W&B uses project loopformer. Checkpoints and logs go to models/stage1_pointer/; model binaries are excluded from Git. CUDA fit, speed and quality need the new desktop smoke/run.

Inspect recurrent checkpoints

bash eval_executor.sh runs matched-count/precision diagnostics, full-loop evaluation and actual learned stopping on development data. Results and terminal logs go under eval/pointer_diagnostics/. The older depth-12 scripts remain available for historical runs.

Pass a complete saved step directory, including its adapter weights and tokenizer:

python -m scripts.eval.loop_test \
  --model models/stage1_pointer/executor_r-seed61/step-000750 \
  --data data/pointer/seed-61-independent/validation.jsonl \
  --device cuda --loops 12 --test

This requires the complete checkpoint on the training desktop; local metadata alone is insufficient. Remove --test for all 1,536 validation queries. Outputs go to eval/pointer_loops/. This reads the model after every recurrent loop; the ordinary three-shot prompt is not used. See full-loop evaluation for commands and metrics.

Deferred overscaling experiments

Preview the separate absorbing-terminal task variant without loading a model or writing files:

python -m scripts.eval.overscaling_test --dry-run

Checkpoint sweeps, repair/damage scoring, and survival exports are available but running them is deferred until further Stage 1 training and review. See overscaling usage.

Documentation

About

Studying recurrent transformer blocks for iterative reasoning, error correction, and depth generalization.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages