SSIT: a label-free, gold-free test for whether LLM-judge position x verbosity bias corrections actually compose. No human labels, no model of the judge. Code, 7-judge/6-family pilot data, and pre-registered protocol for the NeurIPS 2026 JUDGe workshop paper.
bias-correction difference-in-differences ssit llm-evaluation llm-as-judge neurips-2026 position-bias verbosity-bias
-
Updated
Sep 16, 2026 - Python