NickLothian.com · Mar 17, 2023
Large Language Models do Gradient Descent at runtime
0Sign in to vote or save
This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.
Researchers show that LLMs secretly perform gradient descent as meta-optimizers during in-context learning
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.