A diagnostic control paradigm for activation measurements in transformer language models. Cross-replay separates text-bound from architecture-bound components by replaying generated sequences through intact and perturbed model variants.
language-models pythia interpretability variance-decomposition linear-probing permutation-test detrended-fluctuation-analysis gpt-2 activation-analysis representational-similarity transformer-interpretability gpt-j centered-kernel-alignment mechanistic-interpretability attention-analysis cross-replay procrustes-distance
-
Updated
May 17, 2026 - Jupyter Notebook