Evidence and Its Limits for Position Bias Across Sequence Mixers
Aman Behera1, Namit Solanki2, Mehul Anand3
1IIT Roorkee 2AISSMS IOIT, Pune 3Independent Researcher
Project page · Paper · Dataset · Evidence ledger · Artifacts · Runbooks
TL;DR. "Lost in the middle" was measured on Transformers, but it is often quoted as a fact about long context itself. We move one answer-bearing passage through ten positions of a fixed document set and compare Transformer, state-space, and hybrid models on identical questions. Pythia favours the opening of the context while two Mamba checkpoints do not, and all of them favour the end. The measured curves differ by model family, but the study does not isolate architecture as the cause.
1 2 3 4 5 6 7 8 9 10
Position 1 [■] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] Question
Position 5 [ ] [ ] [ ] [ ] [■] [ ] [ ] [ ] [ ] [ ] Question
Position 10 [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [■] Question
└─ start ─┘ └─ middle ─┘ └── end ──┘
The released evaluation contains 2,655 multi-document questions, each with one answer-bearing passage (■) and nine fixed distractors. We use 800 questions for exploratory analysis and leave 1,855 questions untouched for confirmation. For each question, we move the answer-bearing passage through positions 1 to 10 while holding the question, distractors, prompt template, and decoding configuration fixed. The protocol is derived from the Lost in the Middle evaluation.
- Primacy is mean accuracy at positions 1 and 2 minus mean accuracy at positions 5 and 6.
- Recency is mean accuracy at positions 9 and 10 minus mean accuracy at positions 5 and 6.
- Inference uses 10,000 bootstrap resamples of complete question bundles and Holm correction across the two edge tests.
| Primary comparison at 2.8B scale | Matched 8B pure and hybrid checkpoints |
Primary comparison.
Pythia-2.8B has a +5.19 percentage-point primacy edge, while Mamba-2.8B and Mamba-2 2.7B have edges of -0.13 and -1.81 points.
The paired Pythia-minus-Mamba primacy differences are +5.31 and +7.00 points, both with Holm p < 0.0001.
All three models show positive recency edges.
See the Phase 2 summary.
Matched 8B comparison.
The hybrid model has a larger primacy estimate than pure Mamba-2, but the paired hybrid-minus-pure effect is +1.88 points with a 95% confidence interval of [-0.56, +4.44] and Holm p = 0.1442.
The direction is consistent with the primary comparison, but the paired effect is statistically uncertain and does not establish that attention caused the difference.
This released-checkpoint contrast also changes the models' MLP composition, so it is not an attention-only intervention.
The original run recorded a dirty producing tree, while a later clean rerun at commit 33d6bb5 reproduced the summary byte-for-byte.
See the Phase 3 summary, clean-rerun report, and the matched model release of Waleffe et al. (2024).
| Calibration and positive control | Scale across five size pairs |
Calibration and scale. The end-to-end calibration and key-value control show that the harness can detect a known position effect. Across five approximate size pairs, the family gap is near zero at the two smallest scales and appears in the three larger pairs, but capability and architecture remain confounded. See the Phase 1 summary, Phase 4 summary, and the original Pythia, Mamba, and Mamba-2 papers.
| Corpus control | Task check on synthetic retrieval |
Corpus and task checks. Changing the Mamba-2.8B pretraining corpus changes overall accuracy but produces no detectable primacy or recency shape change in this executed contrast. On RULER at 2K tokens, Pythia reproduces a primacy edge, while both Mamba models saturate at perfect accuracy and therefore do not support a mixer comparison on that task. See the Phase 5 summary and Phase 6 summary.
Mechanistic evidence and production systems
Mechanistic evidence. Late-layer attention-sink mass tracks Pythia primacy across scale, position remains linearly decodable in both model families, and prompt variants show substantially different per-condition edges. These results are correlational and diagnostic, not evidence of a causal mechanism. See the Phase 7 summary.
Production systems. Nemotron-H-8B, Llama-3.1-8B, and Qwen2.5-7B all show positive primacy edges, but their many architectural and training differences make this a descriptive prevalence check rather than an architecture test. See the Phase 8 summary.
Important
- All reported ten-document QA results are exploratory, and the 1,855-question confirmatory split remains unopened.
- The committed sham-gold and distractor-order controls passed on a 200-question exploratory Pythia sample, but they were not run on the 8B Megatron checkpoints, and the prepared manual audit has no human labels.
- The matched 8B paired effect is statistically uncertain, and the mechanism analyses do not support a causal attention claim.
- Confidence intervals quantify variation across questions for fixed checkpoints, prompts, and decoding settings, not variation across training seeds or model checkpoints.
The committed summaries regenerate every canonical SVG and PDF figure deterministically.
uv sync --extra test
uv run pytest -q
uv run python paper/generate_figures.pyGPU execution, checkpoint validation, and phase-specific analysis commands are documented in the runbooks. The figure tests check deterministic regeneration, expected labels, and source-summary provenance.
The released dataset flattens 229,700 selected committed model generations across 17 pinned checkpoints, 10 evidence positions, 4 prompt variants, and 2 tasks into one schema in dataset/, together with 280,000 per-layer attention-sink measurements.
It excludes the uncommitted Phase 6 synthetic-retrieval generations, the later Pythia certification-control runs, and the duplicate clean Phase 3 rerun.
Field documentation and collection details are in the datasheet.
| File | Contents |
|---|---|
generations.jsonl.gz |
One row per model generation |
position_accuracy.csv |
Accuracy by model, condition, and evidence position |
attention_sink.jsonl.gz |
Per-layer attention-sink measurements |
runs.csv |
Pinned checkpoints and run metadata |
uv run mixing-matters build-dataset --output datasetThe builder reads only committed artifacts, so any clone reproduces the same files.
The project page presents these results interactively.
Its source is in web/, and every number it renders is generated from the committed phase summaries.
uv run mixing-matters build-site-data --output web/data/results.json
python3 -m http.server 8123 --directory webPushing to main deploys it through the Pages workflow, which regenerates the page data and stages the paper PDF and the dataset alongside it.
src/mixing_matters/ evaluation harness, phase analyses, dataset and site builders
artifacts/ committed per-phase summaries and reports
dataset/ released generations, accuracies, and datasheet
paper/ paper source, figures, and evidence ledger
web/ project page
docs/ GPU runbooks for each phase
tests/ end-to-end and figure regeneration tests
The paper builds on the position-intervention protocol of Liu et al. (2024) and evaluates models introduced by Biderman et al. (2023), Gu and Dao (2024), Dao and Gu (2024), and Waleffe et al. (2024). The complete scholarly bibliography is included in the paper.
If you use the study, the released dataset, or the evaluation harness, please cite the paper.
@inproceedings{behera2026mixing,
title = {Mixing Matters? Evidence and Its Limits for Position Bias
Across Sequence Mixers},
author = {Behera, Aman and Solanki, Namit and Anand, Mehul},
booktitle = {New in ML Workshop at NeurIPS},
year = {2026},
url = {https://beingamanforever.github.io/Mixing-Matters/}
}We thank the Indian Institute of Technology Roorkee (IIT Roorkee) for providing the computational resources that supported this research.
This project is released under the Apache License 2.0.