
Introduction to large language models
Dormant Last read · last published · next check
Read 8 days ago and current, but nothing has been published for 20 months.
Latest videos
Saves to your Watch queue, to pick up on another day or another device.


L14: Causal language modelling | decoder only transformers & autoregressive text generation

L13: Transformers for language modelling | encoder only decoder only & encoder decoder architectures

L11: Language modelling | pre training foundation for large language models

L12: Introduction to language modelling - motivation | fine tuning in GPT & transformers

L10: Layer normalization | normalization in transformers encoder decoder architecture explained

L8: Batch normalization | residual connections and layer normalization in transformers

L5: Sinusoidal encoding & sequence order

L6: Zooming into decoder layer | decoding transformers masked self attention &cross attention

L7: Positional encoding motivation methods & limitations

L9: Teacher forcing & masked attention | autoregressive decoding with masking in transformers

L4: Multi-headed attention in transformers explained

L3: Self-attention in transformers encoder & contextual word embeddings

L2: Attention is all you need transformer architecture explained

