Skip to content

Latest commit

 

History

110 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Mixing Matters? Evidence and Its Limits for Position Bias Across Sequence Mixers. Accepted at New in ML, NeurIPS 2026.

Mixing Matters?

Evidence and Its Limits for Position Bias Across Sequence Mixers

Aman Behera1, Namit Solanki2, Mehul Anand3

1IIT Roorkee   2AISSMS IOIT, Pune   3Independent Researcher

Accepted at New in ML @ NeurIPS 2026 Project page Paper Dataset License Python

Project page · Paper · Dataset · Evidence ledger · Artifacts · Runbooks

TL;DR. "Lost in the middle" was measured on Transformers, but it is often quoted as a fact about long context itself. We move one answer-bearing passage through ten positions of a fixed document set and compare Transformer, state-space, and hybrid models on identical questions. Pythia favours the opening of the context while two Mamba checkpoints do not, and all of them favour the end. The measured curves differ by model family, but the study does not isolate architecture as the cause.

Study design

            1     2     3     4     5     6     7     8     9     10
Position 1  [■]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   Question
Position 5  [ ]   [ ]   [ ]   [ ]   [■]   [ ]   [ ]   [ ]   [ ]   [ ]   Question
Position 10 [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [ ]   [■]   Question
            └─ start ─┘             └─ middle ─┘             └── end ──┘

The released evaluation contains 2,655 multi-document questions, each with one answer-bearing passage (■) and nine fixed distractors. We use 800 questions for exploratory analysis and leave 1,855 questions untouched for confirmation. For each question, we move the answer-bearing passage through positions 1 to 10 while holding the question, distractors, prompt template, and decoding configuration fixed. The protocol is derived from the Lost in the Middle evaluation.

  • Primacy is mean accuracy at positions 1 and 2 minus mean accuracy at positions 5 and 6.
  • Recency is mean accuracy at positions 9 and 10 minus mean accuracy at positions 5 and 6.
  • Inference uses 10,000 bootstrap resamples of complete question bundles and Holm correction across the two edge tests.

Main results

Accuracy by evidence position for Pythia 2.8B, Mamba 2.8B, and Mamba-2 2.7B Accuracy by evidence position for matched pure and hybrid Mamba-2 8B models
Primary comparison at 2.8B scale Matched 8B pure and hybrid checkpoints

Primary comparison. Pythia-2.8B has a +5.19 percentage-point primacy edge, while Mamba-2.8B and Mamba-2 2.7B have edges of -0.13 and -1.81 points. The paired Pythia-minus-Mamba primacy differences are +5.31 and +7.00 points, both with Holm p < 0.0001. All three models show positive recency edges. See the Phase 2 summary.

Matched 8B comparison. The hybrid model has a larger primacy estimate than pure Mamba-2, but the paired hybrid-minus-pure effect is +1.88 points with a 95% confidence interval of [-0.56, +4.44] and Holm p = 0.1442. The direction is consistent with the primary comparison, but the paired effect is statistically uncertain and does not establish that attention caused the difference. This released-checkpoint contrast also changes the models' MLP composition, so it is not an attention-only intervention. The original run recorded a dirty producing tree, while a later clean rerun at commit 33d6bb5 reproduced the summary byte-for-byte. See the Phase 3 summary, clean-rerun report, and the matched model release of Waleffe et al. (2024).

Controls and scope

End-to-end calibration and key-value positive-control results Primacy effects and paired Pythia-minus-Mamba differences across model scales
Calibration and positive control Scale across five size pairs

Calibration and scale. The end-to-end calibration and key-value control show that the harness can detect a known position effect. Across five approximate size pairs, the family gap is near zero at the two smallest scales and appears in the three larger pairs, but capability and architecture remain confounded. See the Phase 1 summary, Phase 4 summary, and the original Pythia, Mamba, and Mamba-2 papers.

Mamba 2.8B position curves after changing the pretraining corpus Pythia and Mamba primacy and recency effects on multi-document QA and 2K-token synthetic needle retrieval
Corpus control Task check on synthetic retrieval

Corpus and task checks. Changing the Mamba-2.8B pretraining corpus changes overall accuracy but produces no detectable primacy or recency shape change in this executed contrast. On RULER at 2K tokens, Pythia reproduces a primacy edge, while both Mamba models saturate at perfect accuracy and therefore do not support a mixer comparison on that task. See the Phase 5 summary and Phase 6 summary.

Mechanistic evidence and production systems

Attention-sink, position-probe, and prompt-sensitivity analyses

Mechanistic evidence. Late-layer attention-sink mass tracks Pythia primacy across scale, position remains linearly decodable in both model families, and prompt variants show substantially different per-condition edges. These results are correlational and diagnostic, not evidence of a causal mechanism. See the Phase 7 summary.

Position curves for Nemotron-H 8B, Llama 3.1 8B, and Qwen 2.5 7B

Production systems. Nemotron-H-8B, Llama-3.1-8B, and Qwen2.5-7B all show positive primacy edges, but their many architectural and training differences make this a descriptive prevalence check rather than an architecture test. See the Phase 8 summary.

Evidence boundary

Important

  • All reported ten-document QA results are exploratory, and the 1,855-question confirmatory split remains unopened.
  • The committed sham-gold and distractor-order controls passed on a 200-question exploratory Pythia sample, but they were not run on the 8B Megatron checkpoints, and the prepared manual audit has no human labels.
  • The matched 8B paired effect is statistically uncertain, and the mechanism analyses do not support a causal attention claim.
  • Confidence intervals quantify variation across questions for fixed checkpoints, prompts, and decoding settings, not variation across training seeds or model checkpoints.

Quick start

The committed summaries regenerate every canonical SVG and PDF figure deterministically.

uv sync --extra test
uv run pytest -q
uv run python paper/generate_figures.py

GPU execution, checkpoint validation, and phase-specific analysis commands are documented in the runbooks. The figure tests check deterministic regeneration, expected labels, and source-summary provenance.

Released dataset

The released dataset flattens 229,700 selected committed model generations across 17 pinned checkpoints, 10 evidence positions, 4 prompt variants, and 2 tasks into one schema in dataset/, together with 280,000 per-layer attention-sink measurements. It excludes the uncommitted Phase 6 synthetic-retrieval generations, the later Pythia certification-control runs, and the duplicate clean Phase 3 rerun. Field documentation and collection details are in the datasheet.

File Contents
generations.jsonl.gz One row per model generation
position_accuracy.csv Accuracy by model, condition, and evidence position
attention_sink.jsonl.gz Per-layer attention-sink measurements
runs.csv Pinned checkpoints and run metadata
uv run mixing-matters build-dataset --output dataset

The builder reads only committed artifacts, so any clone reproduces the same files.

Project page

The project page presents these results interactively. Its source is in web/, and every number it renders is generated from the committed phase summaries.

uv run mixing-matters build-site-data --output web/data/results.json
python3 -m http.server 8123 --directory web

Pushing to main deploys it through the Pages workflow, which regenerates the page data and stages the paper PDF and the dataset alongside it.

Repository layout

src/mixing_matters/   evaluation harness, phase analyses, dataset and site builders
artifacts/            committed per-phase summaries and reports
dataset/              released generations, accuracies, and datasheet
paper/                paper source, figures, and evidence ledger
web/                  project page
docs/                 GPU runbooks for each phase
tests/                end-to-end and figure regeneration tests

Citation

The paper builds on the position-intervention protocol of Liu et al. (2024) and evaluates models introduced by Biderman et al. (2023), Gu and Dao (2024), Dao and Gu (2024), and Waleffe et al. (2024). The complete scholarly bibliography is included in the paper.

If you use the study, the released dataset, or the evaluation harness, please cite the paper.

@inproceedings{behera2026mixing,
  title     = {Mixing Matters? Evidence and Its Limits for Position Bias
               Across Sequence Mixers},
  author    = {Behera, Aman and Solanki, Namit and Anand, Mehul},
  booktitle = {New in ML Workshop at NeurIPS},
  year      = {2026},
  url       = {https://beingamanforever.github.io/Mixing-Matters/}
}

Acknowledgments

We thank the Indian Institute of Technology Roorkee (IIT Roorkee) for providing the computational resources that supported this research.

License

This project is released under the Apache License 2.0.

About

[NeurIPS'2026] Position Bias in Mamba and Hybrid Language Models | Evidence Position Bias Across Sequence Mixers in Long-Context Question Answering

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages