The Twenty Percent · Mar 15, 2023
Language modeling journey: From bigram prediction and DIY transformers to LLaMA 65B
0Sign in to vote or save
This page cannot be shown here. You can still read it on the original site — the toolbar below keeps your place in the directory.
With all the hype surrounding chatGPT (and now GPT-4), it really bothered me that I don’t have the faintest idea of how language models or transformers work. Fortunately, the Neural Networks: Zero to Hero lecture series that helped me understand backpropagation in my previous post, also covers multiple language modeling techniques. I found that I spend too much time in my last post explaining…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.