RSS Amplifier

Sorry Dave · Mar 17, 2024

The Mystery of Deep Learning

0
Sign in to vote or save

Todd Moses · Sorry Dave

Despite being a few years in production, people have yet to learn precisely how large language models (LLMs) work.

Why it matters: LLMs have accomplished much more than their creators realized was possible, leaving many unanswered questions.

  • Turing Award winner Judea Pearl suggests that these models are just curve-fitting.

  • Developers of LLMs designed them to work by predicting the next word based on what came previously.

  • However, how they are suppose to work does not align with the actual mechanisms at play.

The big picture: Apple Researcher Hattie Zhou explains, "Can we ever be confident that models have stopped learning?"

  • According to MIT Technical Review, in certain cases, LLMs could fail to learn a task and then suddenly get it, as if a lightbulb had switched on.

Zoom in: MIT reports that LLMs behave in ways textbook math says they shouldn't.

  • University of California, San Diego computer scientist Mikhail Belkin explains that their theoretical analysis of LLMs cannot explain how these models learn language.

  • LLMs learn lots of things that they have not been trained for.

  • Harvard University computer scientist Boaz Barak describes, "The model can learn math problems in English, then see some French literature, and from that generalize to solving math problems in French."

Between the lines: The current theory suggests a hidden mathematical pattern in human language that LLMs have discovered.

  • This is not surprising when one realizes that each word learned by the LLM is first converted to a numeric vector that captures various characteristics of that word concerning the training text.

Reality check: A comment on Hacker News describes the idiocracy at play as "the blind leading the blind."

  • Wikipedia concludes that this phrase "is used to describe a situation where a person ignorant of a given subject is getting advice and help from another person who is just as ignorant of the subject."

  • What is more of a mystery is why organizations are racing to adopt LLMs in 2024.

  • MIT warns that the risk of these models could become a real issue shortly.

Go deeper: Reach out for a free copy of the State of Software report for 2024.

No posts

Read the original on sorrydave.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.