Search Capability Leaderboard
kept by eddie
Parallel evaluates models' web-search ability by combining DeepSearchQA, Humanity's Last Exam, and its WISER benchmark, measuring accuracy, lift, cost, speed, and Pareto efficiency.