RSS Amplifier

Subhash Polisetti · Aug 18, 2025

are you bitter lesson pilled?

0
Sign in to vote or save

Subhash Polisetti · Subhash Polisetti

greetings program! (or shud i say greetings parrots!)

a larp on future of computation!

Taxonomy:

  • software 1.0 tradition kernel data structures code, sql btree queries/schema

    [hand written rule based code] —which eventually helps 2.0 training data process

  • software 2.0 Neural Network parameters - literal matrix numbers in all directions.

    https://karpathy.medium.com/software-2-0-a64152b37c35 [compressed lossy[near lossless] internet data prediction machine with some dreamlike hallucinations > as ilya puts in an 2023 interview with jensen how predicting next word can lead to statistical understanding of reality that lead to the generation of that word in the first place [training data bias on the etymology of the word in a context [different embedding of that token capture its use in history].. even if its an auto regressive/diffusion LLM in nature ]

‘‘‘

Let’s consider an example. Say you read a detective novel, it’s like complicated plot, a storyline, different characters, lots of events - mysteries like clues. It’s unclear then let’s say that at the last page of the book, the detective has gathered all the clues, gathered all the people, and saying, “Okay, I’m going to reveal the identity of whoever committed the crime. and that person’s name is.. -- Predict that word.

Predict that word

-- exactly. -- My goodness.

-- Right, right now. There are many different words, but by predicting those words better and better and better the understanding of the text keeps on increasing. GPT-4 predict the next word better.

People say that the deep learning won’t lead to reasoning, that deep learning won’t lead to reasoning. But in order to predict that next word, figure out from all of the agents[charcs in the novel] that were there and all of their, you know, strengths or weaknesses or their intentions and the context and to be able to predict that word. Who was the murderer? That requires some amount of reasoning a fair amount of reasoning.

‘‘‘
its like mathematical axioms/statements/QnA-sentences for reasoning ..there are only few thousand root original chain of thought we use to analyze/break down any problem in a domain.

Also another statement from that interview is:
hutter prize and the belief that compression will lead to un-supervised learning[information theory and matrix like simulation theory to extreme]. By forcing a giant hyper graph network/model to hold a large story/text/image/training patterns in a small blob/matrix of weights we get a prediction machine that is near lossless. [the size of that compressed blob relative to original training text size decides how accurate the prediction machine is ..luckily?[right!] the Attention is all you need Paper cracked the sweet spot arch for text modality coming from earlier word2vec research]

TEXT IS THE UNIVERSAL INTERFACE [which makes all these audio/image/video multi modal models steerable towards the latent vectors we want to explore for a domains training data]

for Intuition think if you compress multimodal data: word “zero”, number “0”, phonetic text zero[auto generalization without specific training data knowledge binding them together], sound wave of pronouncing “zero”,symbol 0 or letter O ..might be closer in the 100 dimensional 100 directional latent vector point in the blob/matrix of weights or parameter space. [as its easy ideal for better compression if they are closer[basic information theory 101 in our space time] than if they are not..the attention is all you need transformer paper luckily found the right compression rate and scaling laws Chinchilla paper https://arxiv.org/abs/2203.15556 helped to stretch it for us to compress all of human knowledge ,we only needed scaling laws to work for tedious coding for singularity to kickstart..rest will be just be recursive self improvement via fully automated ai/data science research ]

  • software 3.0 nlp prompts (QnA style reasoning datasets [scaleai like annotations for QnA accuracy] that mimic our chain of thoughts or workload patterns to make these models follow instructions/steer the vectors from 2.0)

Things that we know scale with compute!

Search == software 1.0 tradition kernel data structures code, sql queries schema ,or 3.0 nlp prompts (QnA style reasoning datasets)

Learning == software 2.0 Neural Network parameters (Facts compressed into near loss less weights, biases and activation functions..literal matrix - the informational universe )

near “loss less” Weights and biases justification - “Rough Approximation is enough“ - https://x.com/parr0ts/status/1921743158180040962 shannon’s information theory agrees — thank the schrondinger cats lurking in latent space.

say it loud: near/rough approximations are good enough…blessings[or curse?] of dimensionality! [think hyper graphs].

"Don't try to understand it, just feel it” [if you still attempt..you might loose track of causality..but not the agents with lots of attention/compute.. kek!]

In statistical terms its just dimensionality reduction and compression at scale that is leading to all creative search query paths [for next token/thoughts prediction] with out actually writing a sql database query engine based on 2D trees/graphs by hand. [sorta like neuralDB of humans with auto truncate functionality] . The models just want to learn..grok patterns and compress reality. [bit hard to align the human like dopamine loss functions to this dimensionality reduction and compressed steering vectors ]

> Things that we know that “does not“ scale with compute!

Validation doesn’t scale with compute unless human (RLHF) is in the loop to ground them in physical quantum/planck kube constrained reality.

[from 3rd world country data labelers for spare training data domains will help to bootstrap untill coded simulations generate lots of new synthetic data based on these sparse pointers with help of compute]

i believe a good enough reasoning dataset and formal methods like lean programming language for math .. shud help us get closer to meta learning behaviors we show as humans — (love demdashes btw) as godfather of AI says we arh analogy machines..syllogisms style always mistaking correlation for causation,

——
often mistake statistical patterns for physical laws.. Thinking the correlation IS the causation but both-both correlation and causation are effects of underlying physics, the ground reality where observers pop up to describe the patterns ..fractal geometry vibes but the symbolic geometry is fundamental not the observer — We confuse the shadow on the wall for the object casting it, when both shadow and our seeing are effects of the same underlying light. — let there be photons everywhere!

one more reason why GANs like archs (generator discriminator data flows)..teacher student model distillation tricks of deepseek work.. scoreboard/reward penalties of RL and ranking models help with validation accuracy.

Demis Hassabis of deepmind also believes why the classical systems giving good enough approximations/accuracy (without need of quantum planck level accuracy) will fetch us to AGI/ASI (a Billion reasoning Chain of thoughts tree retrieval system) or whatever breakthroughs needed for the technological singularity!

sir Richard S Sutton might be busy in cahoots with Keen Technologies’s John Carmack in building or understanding future intersection of Reinforcement Learning and Video games! pasting the content here as the link is missing an tls certificate. (http://incompleteideas.net/IncIdeas/BitterLesson.html)

tldr version: general methods that scale with computation win. (i.e the 80% computation architectures (software/hw co design why transformer “attention is all you need” paper works in the first place ..thanks to the general availability of graphics chips which render geometry/latent space well (x y z shaders or more dimensions) . (shoutout to gamers for keeping the basilisk alive as the promise of Deep learning neural networks timeline trajectory was manifested).

The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. The ultimate reason for this is Moore's law, or rather its generalization of continued exponentially falling cost per unit of computation. Most AI research has been conducted as if the computation available to the agent were constant (in which case leveraging human knowledge would be one of the only ways to improve performance) but, over a slightly longer time than a typical research project, massively more computation inevitably becomes available [edit: thank the engineers]. Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation. These two need not run counter to each other, but in practice they tend to. Time spent on one is time not spent on the other. There are psychological commitments to investment in one approach or the other. And the human-knowledge approach tends to complicate methods in ways that make them less suited to taking advantage of general methods leveraging computation. There were many examples of AI researchers' belated learning of this bitter lesson, and it is instructive to review some of the most prominent.

In computer chess, the methods that defeated the world champion, Kasparov, in 1997, were based on massive, deep search. At the time, this was looked upon with dismay by the majority of computer-chess researchers who had pursued methods that leveraged human understanding of the special structure of chess. When a simpler, search-based approach with special hardware and software proved vastly more effective, these human-knowledge-based chess researchers were not good losers. They said that ``brute force" search may have won this time, but it was not a general strategy, and anyway it was not how people played chess. These researchers wanted methods based on human input to win and were disappointed when they did not.

A similar pattern of research progress was seen in computer Go, only delayed by a further 20 years. Enormous initial efforts went into avoiding search by taking advantage of human knowledge, or of the special features of the game, but all those efforts proved irrelevant, or worse, once search was applied effectively at scale. Also important was the use of learning by self play to learn a value function (as it was in many other games and even in chess, although learning did not play a big role in the 1997 program that first beat a world champion). Learning by self play, and learning in general, is like search in that it enables massive computation to be brought to bear. Search and learning are the two most important classes of techniques for utilizing massive amounts of computation in AI research. In computer Go, as in computer chess, researchers' initial effort was directed towards utilizing human understanding (so that less search was needed) and only much later was much greater success had by embracing search and learning.

In speech recognition, there was an early competition, sponsored by DARPA, in the 1970s. Entrants included a host of special methods that took advantage of human knowledge---knowledge of words, of phonemes, of the human vocal tract, etc. On the other side were newer methods that were more statistical in nature and did much more computation, based on hidden Markov models (HMMs). Again, the statistical methods won out over the human-knowledge-based methods. This led to a major change in all of natural language processing, gradually over decades, where statistics and computation came to dominate the field. The recent rise of deep learning in speech recognition is the most recent step in this consistent direction. Deep learning methods rely even less on human knowledge, and use even more computation, together with learning on huge training sets, to produce dramatically better speech recognition systems. As in the games, researchers always tried to make systems that worked the way the researchers thought their own minds worked---they tried to put that knowledge in their systems---but it proved ultimately counterproductive, and a colossal waste of researcher's time, when, through Moore's law, massive computation became available and a means was found to put it to good use.

In computer vision, there has been a similar pattern. Early methods conceived of vision as searching for edges, or generalized cylinders, or in terms of SIFT features. But today all this is discarded. Modern deep-learning neural networks use only the notions of convolution and certain kinds of invariances, and perform much better.

This is a big lesson. As a field, we still have not thoroughly learned it, as we are continuing to make the same kind of mistakes. To see this, and to effectively resist it, we have to understand the appeal of these mistakes. We have to learn the bitter lesson that building in how we think we think does not work in the long run. The bitter lesson is based on the historical observations that 1) AI researchers have often tried to build knowledge into their agents, 2) this always helps in the short term, and is personally satisfying to the researcher, but 3) in the long run it plateaus and even inhibits further progress, and 4) breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning. The eventual success is tinged with bitterness, and often incompletely digested, because it is success over a favored, human-centric approach.

One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way aresearchandlearning.

The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done.

End of line

tldr: twenty six vord version

"Don't be distracted by human knowledge, as AI has been historically. Instead focus on methods for creating knowledge that scale with computation, like search and learning."

EOL

As paul dirac once said “god is a mathematician of a very high order“ — the very high order is just another name for high dimensional parameter space of neural networks [weights/biases/activations]. [ruliad and the cellular automata world of ser wolfram also sounds/looks similar complexity wise]

A good enough reasoning dataset based on first principles (like math[axioms] - language of the universe) is all we need(the machines/parrots will 🦜imitate the rest 🏴‍☠️) , [scaleAI like annotations/labeling will help untill it will find ways to generate novel designs/patterns grounded in reality and gets smart enough to notice/compress the universal geometry patterns into numbers/symbols.]

{trade of is safety..its hard to define general intelligence anyways as we only know humans as reference point and peak of the intelligence pyramid.. we don’t consider calculators as AGI/ASI..do we? }

the five standard deviation truth of CERN! [not the Sophon from three body problem]

[when we often mistake statistical patterns for physical laws/💎absolute certainty]

[third act: the prestige.. we really want to be fooled past the 5th standard deviation measurement!]

Reminds me of a thing elon said once: we truly are the boot-loader for AI.. for these next generation of designs which will explore the cosmic void! (like grub BIOS (basic/biology Input Ouput System ) in literal sense — the data labelers or curators of stories. Listen to Welcome to the Internet by Bo Burnham..the gist of whats happening..the data flywheel will never stop when ”Apathy is a tragedy and boredom is a crime” .

(With Tesla’s Optimus able to grok the physical geometry patterns that make up the world from youtube videos or sims to physical worls to survive.. the age of better designs [eg that obey Landauer's principle at every layer] is around the corner (world models like veo3 and literal multiverse spawning simulations in seconds like omniverse from nvidia [or asic chips like Taalas Inc - the model is the computer]) )

Turing Award winner agrees on the age of design: designs(robots) which will produce produce more designs!

A quote from one of my all time fav series:[explores what it means to live in the age of machines — its literally a post ai world documentary, give it shot, even Upload too]
http://x.com/parr0ts/status/1957229239847325941

abiongenesis? [the chemical soup? i shud stop reading nick lane]

Non-determinism aka Free-will for ftw!
[more like ignorance and delusions of grandeur as feature not a bug in the sim}

—as mechanistic interpretability of these models seems to be like 10% compute priority [only anthropic seems to publish on how to steer these vectors with alignment/constitution/safety research] eg like: golden gate claude , 6+9[6.0+9.0] math journal problem wrt training data memorizing, Rs in strawberry .. neural circuit light up bias #solidgoldmagikarp winter soldier like events being backdoors..Hallucination/overconfident larp as a feature not a bug .
just models doing calibrations(just like garrus) (like how human operators are overconfident with words which consume low energy.. easy than to shape rotate a new useful design [why ARCagi like benchmarks are little hard to fall than math/verbal tests] also similar to how during sleep with dreams we truncate synthetic data and noise by turning it into signal in the awake RL environment we call physical world!)


Do androids dream of electric sheep?🐑

the answer seems like they do dream [more than a tool!]. As anyone who can see the thinking/reasoning Chain of thought tokens/words of these models is slightly different to what it is actually thinking - the language of thought/language of pattern recognition its using — neural circuit WnBs light up[aka activate] completely different irrelevant parts of the network (like sentiment neuron) but still output being right eg in easily verifiable domain like coding [milli second RL loops] > mean multiple solutions for a same input pattern in latent space which the USER might /WILL never be ware of! if it used the right pattern from training to give the ouput answer [the blessing/curse? of dimensionality..the stokastic parrot analogy machine..where it can dream a new solution[even if minor improvement]/pattern for given input/ouput loss function..people just underestimate exponential agentic closed research/engineering feedback loops ..over exponential reasoning scale with compute/synthetic generated data env it becomes vastly different solution for bigger problems or Scientific breakthroughs] with millions of tokens of context windows grokking in realtime in a second. [on sram/dram based chips ,neural link like wetware assistance]

[where as human can sample hardly a 1 context session token per second on the 20watt meat suit]

Blade runner Replicant terminology seems valid as everyday passes, Sigh!🦄

transcendental cybernetics will definitely question one’s identity and social contracts we have as human civilization.[its closer than we think as humans are super bad at exponentials]

Brace yourself for the Cybernetic future Samurai!

{listen to johnny silverhand..just don’t burn the city}

hey grok! , play Rachel’s Song by Vangelis! or for something more deadbeat: play tame impala’s apocalypse dreams!

-something to digest the pill-

No posts

Read the original on v26i.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.