tl;dr: Collecting some resources.
Papers
- Attention Is All You Need, by Vaswani et al., 2017
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, by Devlin et al., 2018
- Language Models are Unsupervised Multitask Learners, by Radford et al., 2019
Courses
Learning resources
- Transformer Explainer, by Polo Club of Data Science, Georgia Tech
- colah’s blog
- Distill
Concepts
- Feed-forward neural net (FFNN)
- Recurrent neural net (RNN)
- Convolutional neural net (CNN)
- Deep neural net (DNN)
- Softmax
- Activation
- Attention
- Multi-head attention
- Transformers
Journey
I’ve (partially-)looked at these resources, (somewhat) in this (approximate) order:
- My own lecture notes from 2013
- “Probabilistic Machine Learning: An Introduction”, by Kevin Murphy
- Sebastian Raschka
- 100-page ML book
- 100-page LM book
- Illustrated transformer
- 3blue1brown’s playlist
- “Attention? Attention!”, by Lilian Weng
References
For cited works, see below 👇👇