One line. Many voicesSeek and you shall find

cognition.ai faviconIntroducing FrontierCode

kept by

Cognition's FrontierCode benchmark measures code mergeability, finding even top models like Claude Opus 4.8 score only 13.4% on its hardest 50 tasks.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.