Reading up on statistical-models
1 deep · digging since sep 15
- When LLM judges agree, should we believe them? - Amazon Science
Majority-vote LLM judge panels overstate agreement when outputs are correlated, so the authors propose a dependence-aware Ising model aggregator that improves accuracy by 9-14% over baselines.