Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Stable Because It Stopped Computing

Loop Stability

Code and data for a study of looped (weight-tied) transformers: models that apply one shared block r times instead of stacking r distinct layers.

The question the project answers is why a deeply looped model is stable when it is run for more steps than it was trained for, and the answer is not an architectural trick. The steps of a looped transformer are a power iteration on the Jacobian of its own block, and a power iteration suppresses every direction but the dominant one. A deep loop is stable because it has stopped computing.

Everything here is CPU-reproducible except the training runs themselves.


The claim, in three lines

Write the loop as h_{k+1} = F(h_k) with a constant injection, and set delta_k = h_{k+1} - h_k. Differentiability alone gives

delta_{k+1} = J(h_k) delta_k + O(||delta_k||^2)

That is a power iteration. Three consequences follow, and they are what the data tests:

consequence where it is measured
T1 the direction of delta_k converges consistency check only
T2 the ratio ||delta_{k+1}||/||delta_k|| tends to rho(J), not ||J|| analysis/exp102_collapse.py
T3 the trajectory spans a bounded number of directions analysis/exp128_power.py

T2 is the one that carries the argument, and it is falsifiable rather than decorative because the two candidates make different predictions. Measured against the observed step-to-step gain, rho(J) is closer than ||J|| by 5.4x on phase C, 6.7x on phase D, 14.6x on phase F and 23.4x on phase G. A bound built on ||J|| is not merely looser; it names the wrong quantity.

Theorem D and the measure Sigma

Under the hypothesis

(H)   C := sup_{n >= r} rho_{n+1} / rho_n  <  1

telescoping plus the triangle inequality give a bound on the drift D_k that the state accumulates when the loop is extended past its training depth:

D_k  <=  Sigma (1 - C^k)  <=  Sigma,      Sigma = S / (1 - C)

(H) is an assumption, verified empirically and not proved. It is per-input, and an adversarial input breaks it. Both facts are stated rather than hidden: see analysis/exp141_hypothesis_H.py.

C_med is an estimator, C_max is a certificate. They are never interchangeable. With the median the bound is violated on 16 of 72 models; with the certificate it is not violated at all.

Reproducing the headline numbers

No GPU needed, and no checkpoints: these read results/ only.

cd analysis
python exp198_postnorm.py            # the chain, pre-norm against post-norm
python exp201_bootstrap_postnorm.py  # why the difference is NOT called "better"
python exp180_phaseO_sigma.py        # Sigma and Theorem D at d = 256
python exp141_hypothesis_H.py        # hypothesis (H), estimator against certificate

Then, from verify/:

cd verify
python check_abstract.py             # 133 reported numbers against their source files
python check_total.py                # model count, deduplicated by hash

check_abstract.py recomputes every number that appears in the write-up directly from the JSON that produced it. If a number was edited by hand and the data was not, it fails.

Layout

directory what is in it
runners/ the training runs, one file per experimental phase
analysis/ the scripts that turn checkpoints and JSON into the reported numbers
lib/ the shared numerical core: block, Jacobian, resolvent, Kreiss
results/ one JSON per phase, the raw record of every trained model
checkpoints/ a small sample of the trained weights, 18 of 628
verify/ scripts that check the write-up against results/

Each directory has its own README explaining every file in it. STRUCTURE.md is the map of the whole repository.

Data

results/ holds one record per trained model: configuration, validation loss at r, 2r and 4r steps, the residual sequence after warm-up, and the derived S, C and Sigma. 628 distinct models, deduplicated by weight hash rather than by table row, 127 of them at d = 256 and 57 post-norm.

checkpoints/ holds a small sample of the trained weights: 18 models out of 628, at d = 48, spanning six depths on three phases. The complete set runs to 439 MB, which is too large for a git repository.

The full set of checkpoints is available on request, and will be released publicly soon. Open an issue or get in touch.

The sample is enough to run three of the ten weight-dependent scripts. The other twelve analysis scripts, and every correlational result in the study, reproduce from the JSON records alone.

Numbers recomputed from the sample will not match the reported ones, and should not be expected to. The reported figures come from twenty cells or sixty models per arm; the sample holds six cells at one seed and one carry. As an example, the median second-order remainder at d = 48 is 0.336% over the full twenty cells and 1.145% over the six here. Neither is wrong: they are medians of different sets. The sample shows the shape of the effect, which is what the argument rests on, not the precision of any single figure. For that, use results/, which was computed over the full set.

Three traps, paid for in errors

These are the conventions that differ between runners. Reading a number without reading the runner that produced it is how this project lost four days.

  1. The drift exponents are not what the names suggest. In cluster_run4.py the field D2 is a drift over r steps and D4 over 3r, not 2r and 4r. In cluster_run12.py and cluster_run13.py, val_loss_4r is evaluated at 4r literal steps.

  2. Budget conventions differ between phases. cluster_run4.py counts block applications (steps = budget // r); phase G and later count forward-equivalents (budget = steps * 3r).

  3. A results file holding several experiments does not label the one you want. phaseK.json contains three variants tagged full, and two of them have the injection deleted. Filter on inj_mode == "const" for any baseline comparison. A median over every row tagged full returns a perplexity of 1.8630, which belongs to a broken model.

What the data does not support

Reported here because the negative results carry as much of the argument as the positive ones.

  • Sigma ranks between design cells, not within one. Spearman +0.923 between configurations; +0.016 at p = 0.83 after demeaning within (family, recipe, depth). It cannot compare two models from the same recipe at the same depth.
  • Report S and C, not Sigma. Sigma = S x G with G = 1/(1-C), and G carries the design variables rather than the signal. Pooling across carries reverses the comparison, a Simpson effect.
  • A budget scaling law was retracted the day it was made. One optimum out of seven candidates survived the check that an argmin must rise on both sides in units of seed noise. The retraction is kept deliberately.
  • Early-exit schemes rest on a false premise. A model trained at k beats every deeper model stopped at k, so intermediate states are not approximate answers.

Concurrent work

Viakhirev et al., arXiv:2608.18222, ask the same question independently and reach a related result: they classify a trained operator as settling, marginal or drifting, and prove that a decoded answer cannot change once the remaining path length falls below the decoder's decision margin. That path length is the quantity Theorem D bounds, and both accounts agree on its degenerate case, an infinite sum certifying nothing.

The differences are what this repository adds. They bound the decoded answer on algorithmic tasks, where a discrete answer and a margin exist; Sigma bounds the state, with no decoder involved. They assume the path length known; six forward passes estimate it here to within 1.15% at ninety-six steps. And on the population in results/, their three-way label does not order extension cost, at 0.491 pairwise accuracy against 0.828 for Sigma, because a small settling ratio is bought at small r, where the residual is an order of magnitude larger.

Requirements

Python 3.10+, torch, numpy. The analysis scripts use nothing else. The Jacobian work uses torch.func (jacrev, jvp, vjp).

Citation

@article{cofone2026stable,
  title   = {Stable Because It Stopped Computing},
  author  = {Cofone, Leonardo},
  journal = {arXiv preprint},
  year    = {2026}
}

The arXiv identifier will be added here once the preprint is posted.

License

Released under CC BY-NC-SA 4.0: you may share and adapt this work, including the data and the weights, provided you give attribution, do not use it commercially, and license any derivative under the same terms. See LICENSE. For commercial use, or for other terms, get in touch.

About

Code and data for the paper "Stable Because It Stopped Computing"

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages