Position bias, live
The same mock judge scores system X vs Y on 12 prompts. Switch between the naive judge (X always shown first) and the corrected judge (both orders averaged) and watch which pairs disagree with the true, higher-quality winner.
Verbosity bias survives correction
Five pairs below are tied in quality and differ only in length. Order-averaging cannot help here, there is no position to correct, only length.
| pair | X words | Y words | longer | winner |
|---|---|---|---|---|
| q01 | 50 | 20 | X | X |
| q02 | 20 | 45 | Y | Y |
| q03 | 55 | 25 | X | X |
| q04 | 30 | 60 | Y | Y |
| q05 | 40 | 15 | X | X |
verbosity-bias rate = longer-wins / pairs = 5/5 = 1.00. Quality is tied, yet the longer answer wins every time.
Zoom in: pair p11, the lone survivor of correction
System X has quality 5, system Y has quality 6 and 15 words fixed. Y is the true winner. Drag X's word count and watch the corrected score (quality + 0.02/word, position bonus already averaged away) decide the verdict.