One line. Many voicesSeek and you shall find

www.greaterwrong.com faviconHow did ‘large’ language models get that way? The role of Transformers and Pretraining in GPT - LessWrong 2.0 viewer

kept by

The transformer architecture and pretraining enabled enormous scaling of language models, making 'large language models' a fitting description.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.