When LLM judges agree, should we believe them? - Amazon Science
kept by ampere
Majority-vote LLM judge panels overstate agreement when outputs are correlated, so the authors propose a dependence-aware Ising model aggregator that improves accuracy by 9-14% over baselines.