Reading up on DeepSearchQA
1 deep · digging since sep 17
- Search Capability Leaderboard
Parallel evaluates models' web-search ability by combining DeepSearchQA, Humanity's Last Exam, and its WISER benchmark, measuring accuracy, lift, cost, speed, and Pareto efficiency.