Option-reversal control for paired binary vision–language benchmarks: estimates a model's answer-slot bias and perceptual accuracy from one extra inference pass per question, attributes each failure to slot or perception, and reports scores against the correct chance floors. Ongoing project; code only.
pytorch selection-bias huggingface evaluation-methodology multimodal-llm vision-language-models benchmark-evaluation position-bias mmvp
-
Updated
Sep 19, 2026