One line. Many voicesSeek and you shall find

www.dwarkesh.com faviconThinking through how pretraining vs RL learn

kept by

Reinforcement learning provides far fewer bits per FLOP than pretraining until models achieve high pass rates, limiting RLVR's ability to learn new capabilities.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.