Compression is prediction
Compression and LLMs both minimize surprise: better next-token prediction → shorter codes, via classic info theory.
Brutally short: Training LLMs ≈ optimizing compressors.
🗣️ HN commenters add:
- Point to 3Blue1Brown series and Hutter Prize as deep dives.
- Note limits: compressor↔LM only holds when train/test match; generalization can still fail.














