Code and data for a study of looped (weight-tied) transformers: models that
apply one shared block r times instead of stacking r distinct layers.
The question the project answers is why a deeply looped model is stable when it is run for more steps than it was trained for, and the answer is not an architectural trick. The steps of a looped transformer are a power iteration on the Jacobian of its own block, and a power iteration suppresses every direction but the dominant one. A deep loop is stable because it has stopped computing.
Everything here is CPU-reproducible except the training runs themselves.
Write the loop as h_{k+1} = F(h_k) with a constant injection, and set
delta_k = h_{k+1} - h_k. Differentiability alone gives
delta_{k+1} = J(h_k) delta_k + O(||delta_k||^2)
That is a power iteration. Three consequences follow, and they are what the data tests:
| consequence | where it is measured | |
|---|---|---|
| T1 | the direction of delta_k converges |
consistency check only |
| T2 | the ratio ||delta_{k+1}||/||delta_k|| tends to rho(J), not ||J|| |
analysis/exp102_collapse.py |
| T3 | the trajectory spans a bounded number of directions | analysis/exp128_power.py |
T2 is the one that carries the argument, and it is falsifiable rather than
decorative because the two candidates make different predictions. Measured
against the observed step-to-step gain, rho(J) is closer than ||J|| by
5.4x on phase C, 6.7x on phase D, 14.6x on phase F and 23.4x on phase G.
A bound built on ||J|| is not merely looser; it names the wrong quantity.
Under the hypothesis
(H) C := sup_{n >= r} rho_{n+1} / rho_n < 1
telescoping plus the triangle inequality give a bound on the drift D_k that
the state accumulates when the loop is extended past its training depth:
D_k <= Sigma (1 - C^k) <= Sigma, Sigma = S / (1 - C)
(H) is an assumption, verified empirically and not proved. It is
per-input, and an adversarial input breaks it. Both facts are stated rather
than hidden: see analysis/exp141_hypothesis_H.py.
C_med is an estimator, C_max is a certificate. They are never
interchangeable. With the median the bound is violated on 16 of 72 models;
with the certificate it is not violated at all.
No GPU needed, and no checkpoints: these read results/ only.
cd analysis
python exp198_postnorm.py # the chain, pre-norm against post-norm
python exp201_bootstrap_postnorm.py # why the difference is NOT called "better"
python exp180_phaseO_sigma.py # Sigma and Theorem D at d = 256
python exp141_hypothesis_H.py # hypothesis (H), estimator against certificate
Then, from verify/:
cd verify
python check_abstract.py # 133 reported numbers against their source files
python check_total.py # model count, deduplicated by hash
check_abstract.py recomputes every number that appears in the write-up
directly from the JSON that produced it. If a number was edited by hand and
the data was not, it fails.
| directory | what is in it |
|---|---|
runners/ |
the training runs, one file per experimental phase |
analysis/ |
the scripts that turn checkpoints and JSON into the reported numbers |
lib/ |
the shared numerical core: block, Jacobian, resolvent, Kreiss |
results/ |
one JSON per phase, the raw record of every trained model |
checkpoints/ |
a small sample of the trained weights, 18 of 628 |
verify/ |
scripts that check the write-up against results/ |
Each directory has its own README explaining every file in it.
STRUCTURE.md is the map of the whole repository.
results/ holds one record per trained model: configuration, validation loss
at r, 2r and 4r steps, the residual sequence after warm-up, and the
derived S, C and Sigma. 628 distinct models, deduplicated by weight
hash rather than by table row, 127 of them at d = 256 and 57 post-norm.
checkpoints/ holds a small sample of the trained weights: 18 models out of
628, at d = 48, spanning six depths on three phases. The complete set runs
to 439 MB, which is too large for a git repository.
The full set of checkpoints is available on request, and will be released publicly soon. Open an issue or get in touch.
The sample is enough to run three of the ten weight-dependent scripts. The other twelve analysis scripts, and every correlational result in the study, reproduce from the JSON records alone.
Numbers recomputed from the sample will not match the reported ones, and
should not be expected to. The reported figures come from twenty cells or
sixty models per arm; the sample holds six cells at one seed and one carry.
As an example, the median second-order remainder at d = 48 is 0.336% over
the full twenty cells and 1.145% over the six here. Neither is wrong: they
are medians of different sets. The sample shows the shape of the effect,
which is what the argument rests on, not the precision of any single figure.
For that, use results/, which was computed over the full set.
These are the conventions that differ between runners. Reading a number without reading the runner that produced it is how this project lost four days.
-
The drift exponents are not what the names suggest. In
cluster_run4.pythe fieldD2is a drift overrsteps andD4over3r, not2rand4r. Incluster_run12.pyandcluster_run13.py,val_loss_4ris evaluated at4rliteral steps. -
Budget conventions differ between phases.
cluster_run4.pycounts block applications (steps = budget // r); phase G and later count forward-equivalents (budget = steps * 3r). -
A results file holding several experiments does not label the one you want.
phaseK.jsoncontains three variants taggedfull, and two of them have the injection deleted. Filter oninj_mode == "const"for any baseline comparison. A median over every row taggedfullreturns a perplexity of 1.8630, which belongs to a broken model.
Reported here because the negative results carry as much of the argument as the positive ones.
- Sigma ranks between design cells, not within one. Spearman
+0.923between configurations;+0.016atp = 0.83after demeaning within (family, recipe, depth). It cannot compare two models from the same recipe at the same depth. - Report
SandC, notSigma.Sigma = S x GwithG = 1/(1-C), andGcarries the design variables rather than the signal. Pooling across carries reverses the comparison, a Simpson effect. - A budget scaling law was retracted the day it was made. One optimum out of seven candidates survived the check that an argmin must rise on both sides in units of seed noise. The retraction is kept deliberately.
- Early-exit schemes rest on a false premise. A model trained at
kbeats every deeper model stopped atk, so intermediate states are not approximate answers.
Viakhirev et al., arXiv:2608.18222, ask
the same question independently and reach a related result: they classify a
trained operator as settling, marginal or drifting, and prove that a decoded
answer cannot change once the remaining path length falls below the decoder's
decision margin. That path length is the quantity Theorem D bounds, and both
accounts agree on its degenerate case, an infinite sum certifying nothing.
The differences are what this repository adds. They bound the decoded answer on
algorithmic tasks, where a discrete answer and a margin exist; Sigma bounds
the state, with no decoder involved. They assume the path length known; six
forward passes estimate it here to within 1.15% at ninety-six steps. And on
the population in results/, their three-way label does not order extension
cost, at 0.491 pairwise accuracy against 0.828 for Sigma, because a small
settling ratio is bought at small r, where the residual is an order of
magnitude larger.
Python 3.10+, torch, numpy. The analysis scripts use nothing else. The
Jacobian work uses torch.func (jacrev, jvp, vjp).
@article{cofone2026stable,
title = {Stable Because It Stopped Computing},
author = {Cofone, Leonardo},
journal = {arXiv preprint},
year = {2026}
}The arXiv identifier will be added here once the preprint is posted.
Released under CC BY-NC-SA 4.0: you may share and adapt this work,
including the data and the weights, provided you give attribution, do not use
it commercially, and license any derivative under the same terms. See
LICENSE. For commercial use, or for other terms, get in touch.