Position bias, live

12 pairs: does correction beat the truth?

The same mock judge scores system X vs Y on 12 prompts. Switch between the naive judge (X always shown first) and the corrected judge (both orders averaged) and watch which pairs disagree with the true, higher-quality winner.

Toggle the mode below. In naive mode, pairs tagged would flip their verdict if the order were swapped, that's the position bias. In corrected mode the amber flip tag disappears, but one card still turns rose: a verbosity error averaging cannot touch.
Position-bias rate
0.58
7/12 flip when order swaps
Disagreement with truth
0.42
5/12 pairs

Verbosity bias survives correction

How many words does it take to beat quality?

Five pairs below are tied in quality and differ only in length. Order-averaging cannot help here, there is no position to correct, only length.

pairX wordsY wordslongerwinner
q015020XX
q022045YY
q035525XX
q043060YY
q054015XX

verbosity-bias rate = longer-wins / pairs = 5/5 = 1.00. Quality is tied, yet the longer answer wins every time.

Zoom in: pair p11, the lone survivor of correction

System X has quality 5, system Y has quality 6 and 15 words fixed. Y is the true winner. Drag X's word count and watch the corrected score (quality + 0.02/word, position bonus already averaged away) decide the verdict.

System X
quality 5, 70 words
6.40
System Y  (truth)
quality 6, 15 words
6.30