Key ML Papers: 2020 to 2025
I’m in SF this week, have a pretty packed schedule but reach out if you want to say hi!
Nearly two years ago I set out to read and review the papers that Ilya Sutskever (then Chief Scientist at OpenAI) had given to John Carmack (software engineer extraordinaire, of DOOM fame). The full list of reviews is here:
That process took, end to end, a year and a half to do. O sure, I was doing a bunch of other things in the meantime, like getting married and going to SF a bunch. I probably finished reading all the papers in the first 3 or 4 months, it was the writing that took a while. But still, a lot has changed since I started that series, and more still has changed since I finished it.
Now that we’re nearing the end of 2026, what are the key papers of the last few years?
Ilya has, of course, been living his best hermit life working on SSI. I don’t expect him to come down from his mountain cave with a new stack of papers to share. So, in lieu of Ilya kindly providing me my paper reading list, I’ve gone ahead and put together my own list of key papers (and blog posts and, yes, one textbook) from 2020 - 2025. As with Ilya’s original set of papers, this new set was selected with an eye for learning. If you go through all of these papers, you ideally will have a pretty up to date understanding of the different important threads in deep learning research.
- (2020) RAG -- Harness Engineering
- (2020) An Image is Worth 16x16 Words (ViT) -- Vision
- (2021) CLIP -- Vision (but really important for a whole bunch of things)
- (2021) LoRA -- Scaling
- (2021) AlphaFold 2 -- Science
- (2021) Classifier-Free Guidance -- Diffusion
- (2021) Latent Diffusion -- Diffusion
- (2022) Chain-of-Thought Prompting -- Harness Engineering
- (2022) RLHF (InstructGPT) -- RL
- (2022) Chinchilla -- Scaling
- (2022) FlashAttention -- Scaling
- (2022) Toy Models of Superposition -- Interpretability
- (2022) ReAct / Voyager -- Harness Engineering
- (2022) Speculative Decoding -- Scaling
- (2022) Constitutional AI -- RL
- (2023) I-JEPA -- Vision
- (2023) Direct Preference Optimization -- RL
- (2023) PagedAttention -- Scaling
- (2023) Towards Monosemanticity -- Interpretability
- (2024) Model Context Protocol -- Harness Engineering
- (2025) DeepSeek-R1 -- RL
- (2025) How to Scale Your Model -- Scaling
- (2026) RYS / LLM Neuroanatomy -- Interpretability
- (2025) Welcome to the Era of Experience -- Capstone
Obviously, other smarter people may disagree with this list. Would love to hear other opinions! My hope is that ‘Key ML Papers’ will be my review series over the next year. I plan on reviewing one of these papers every few weeks, and I’ll keep this page up to date as time goes on.
12 Grams of Carbon is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.