Reading up on The Pile
1 deep · digging since sep 09
- Pretraining progress is mostly coming from data
From 2019 to 2025, data improvements contributed 3.24× more compute‑efficiency gains than model changes, delivering 12× versus 3.7× efficiency gains at 1e19 FLOPs.