you guys know how much i love a good recursive structure and self-assembling emergence, so this one will be a bit fun (and will get a little theory heavy by the end). frankly, the whole thing got spooky, maybe more like meta-spooky, in a hurry. man, did it lead to weird dreams.
i ran down a massive rabbit hole with GPT and (unintentionally, at least initially) explored a whole space on intelligence generation that, amusingly, we predicted from first principles but that turns out to have quite a substantial literature behind it. it felt a little bit like discovering neptune by having noticed that, because of the way the other planets moved, there had to be something there and then going to fetch the telescope. the sheer galle of some people… (sorry)
it started with the simple matter of vocabulary, that most misunderstood and yet massively G loaded characteristic long poo-poohed as “rote” and “environmental.” it’s not. it correlates to and predicts general intelligence more accurately than just about any other sub factor. it’s not because of the lazy behavioralist answers like “smart people read more” or “you had some environmental advantages”; it’s a whole level deeper, rooted in relational reasoning capacity which, in turn, is accentuated and accelerated by functional working memory, which, in turn, is a function of how well you can compress concepts while still entailing nuance and maintaining fidelity. that compression is, wait for it, in turn, a function of vocabulary.
it’s a sort of run away intelligence excursion.
this became our model. (i say “our” because neither entity, gato nor chat, could have produced this alone. we filled in gaps for the other. i’m honestly not sure if and to what extent we discovered anything truly new here. more on this soon.)
thus, it would seem that vocabulary stands not just as a result of G (general intelligence) but an enhancer of it, part of the machinery by which intelligence recursively creates more intelligence.
welcome to emergent evo-devo and the idea that it’s not “environment” per se but rather who is in the environment, not what is spoken but who has heard it.
imagine 4 people in a room where this sentence is spoken:
“the committee rejected the proposal as otiose because an existing rule already accomplished everything it was intended to do.”
same environment, but some will learn the word, others will not. it comes down to contextual mapping. you have likely never heard or seen “otiose” before. (it’s a pretty obscure word.) no one has defined it for you. but the context has, it’s just a matter of whether you can see it or not. in an AI sense, the sentence has drawn a shape around the missing token.
now the question is how much and how nuanced a meaning can you extract?
the dull might pull none. the semi-smart might get some low res mapping like “bad.” the smarter will see that it must mean some very specific things because nothing else would fit and retain fidelity.
AI posed this to me, i did not know the word. i guessed “somewhere around either redundant or superfluous.”
this was pretty close to correct, but not spot on.
and that’s where it gets interesting. like borges describing his love of english because of its incredibly fine semantic mapping (like the difference between “brotherly” and “fraternal”) otiose means something a shade different from either of my first-pass guesses. it takes more run ins with the word to learn them and see where it does not quite map to the other ideas. you need to see the venns where it overlaps with what you saw before and what you see later. the new intersections grow smaller and your conceptualization of the idea grows more precise. the negative spaces where “otiose” does not overlap with “redundant” become educational.
given enough run ins with a word, the useful effect of this is that you can now deduce that “otiose” means “without useful effect,” a distinct and differentiated concept.
you just learned a word and deep nuance by context. no one explained it, the negative space revealed it, the intersections of possibility collapsed around it.
relational intelligence begat vocabulary.
and the really interesting bit is that memorizing vocab lists cannot do this in anything like the same way. the tokens you build are small and poorly linked. they lack the richness of context and the power of cross association. you’ll never get the full inclusion of anguish, ennui, and pointlessness you would from really knowing at a deep level what attaches when one uses the word “sisyphean.”
and this is where things get interesting because the word you acquired is not merely another item to remember. it is itself a memory device, it’s grist for the relational mill.
words are compression.
and compression is intelligence.
humans of average and up general intelligence have working memories of fairly similar size limited to a few independently maintained “chunks.” the number of concepts various people can hold differs too little to explain the wide range of human cognition.
what matters is what a “concept” is.
if objects and ideas like “cat” or “educated” are about as far as you can map, working memory is small. if you can compress taxonomies and add nuance to hold an object like “big cats” that now includes lions, tigers, leopards, and jaguars you have more functional working memory.
where this really blows out is when you get past these “objects” and into things like “what does the manifold where 4 objects intersect and interrelate look like.” that’s where the goody room starts and the ability to add huge piles of effective working memory explodes. (i have a working theory that you can blow this out again by moving up even further to “the manifold representing the intersection of manifolds” and make your information density go vertical.)
but here’s the thing: when you compress like that, you need to keep extremely high fidelity in each compressed component because error bars compound and if the manner in which you are mapping concepts gets too lossy, the whole relational system becomes GIGO.
and GIGO may actually constitute a charitable description because garbage in, garbage out implies the garbage just passes through. this is more like inertial navigation: be half a degree wrong standing in your driveway and no biggie. be half a degree wrong crossing the atlantic and use that position to calculate the next bearing and then use that position as the origin for the one after it and pretty soon you’re in guyana looking around wondering what the hell happened to portugal. (or the dominican republic wondering what happened to india…)
compressed concepts become coordinates for other compressed concepts. the little errors do not merely persist, they profoundly and cumulatively move the map and this would seem to make lexical precision far more important than “knowing lots of words.” (though obviously, lots of lexical precision is the dominant outcome)
a fuzzy token is not merely a fuzzy description of one thing, it’s a slightly bent axis upon which you will later hang 40 other things. then you try to use the whole crooked compilation to work out where the next concept ought to fit. but the geometry fails to run true. the cost of imprecision compounds because yesterday’s conclusion becomes tomorrow’s premise.
to make this concrete: mis-mapping otiose as redundant is a small error, but make 20 like it and you’re going to wind up semantically lost.
lack a word for a concept altogether, and you’re going to lack scale in your functional working memory because concepts are (multiplicatively) too expensive to map. every time you need the concept, you have to unpack the whole thing.
this is the conundrum and the trade off:
if you fail to compress, you wind up with low relational capacity, if you compress badly, high relational capacity just makes junk.
you need both and both build the other.
seeing negative spaces teaches words and nuance and words and nuance render one better able to see negative spaces.
to a certain extent, the idea of relational intelligence (raven test etc) and vocabulary as distinct concepts starts to get a little fraught.
this is where the questions start to take on new implications:
what if vocabulary and relational intelligence are largely just different measurements of the same underlying machinery at different stages of compilation?
hoo ah.
raven matrix tests give you a novel relational geometry and ask whether you can solve it now as a prospective matter of inference.
vocabulary presents you with the accumulated residue of thousands of previous occasions on which you successfully located a novel token inside a relational geometry. it’s measuing a history of successful inference. it’s a relational intelligence scaling factor that scales with use and experience. i’m actually starting to wonder about the role this plays in rising functional intelligence as children age.
there seems to be an assumption along the lines of “children get older, brains mature, reasoning improves” and obviously, some of this is physical and chemical, but what if we’re also seeing a feedback cycle?
relational ability improves so successful abstractions accumulate. as a result, vocabulary deepens and the compression library grows so effective working memory expands, higher-order relations become tractable and, you guessed it, still more sophisticated concepts (like vocabulary and manifold based conceptual maps) can be acquired.
what if it’s literally “inferring what a new word means is basically the exact same thing as figuring out the intersectional dimensionality of “suicidal” and “empathy”?
it’s just “how do the salients around this thing limit what the thing could be?”
what if this is all basically just forms of geometry?
if i deeply understand “arbitrage,” i do not need to hold 14 separate propositions about equivalent assets, pricing discrepancies, convergence, simultaneous transactions, financing and risk in working memory. i load one token and the whole structure becomes cheaply addressable.
(i use “token” sort of loosely here in the AI sense of a compact address for a high relational object like “mouthy internet cat” that sits in a particular relational geometry to 1000’s of other ideas like “hitting the ‘nip too hard.”
in my head i see them as little spiky balls. their spikes touch one another in complex ways. the manifold of spike intersection is the commonality (or anti commonality) of the concepts.)
how well you can map where spikey touches spikey and draw implications from the geometry of the contacts is where relational reasoning lives.
it’s also the essence of analogy. a shallow analogy finds one pair of touching spikes and says “X is like Y.” a strong one sees a whole complex intersection of many (manyfold? sorry.) spikeys that are simultaneous and coherent and that survive shape rotation. it also notices which spikes don’t align and thus where the analogy ceases to be useful (or correct). this is inference and the discernment of general cases vs specific ones. what is the superficial drawing and what is the underlying structure?
it’s just like high school trig: if this is a right triangle and i know one side and one acute angle or any 2 sides, i can predict the rest. it’s constrained. it cannot be other. the more complex the shape and the more that is known, the tighter the constraints become.
it turns out there’s a fair bit of empirical theory and study to back some of these cognitive jumps (with which i’ll not bore you.) what’s fun is that, just like vocabulary, we predicted it (with surprising depth into the ideas) from known salients and constrained spaces where something shaped kind of like X ought to be.
consider simple relationships. one can easily say, “a conservative traveling to san francisco and discovering that pacific heights is a lovely, friendly neighborhood and not a world war Z of needles, crime, and feces” is very much like “a japanese tourist discovering austin and discovering that it is a lovely, friendly city and not the gun toting murder festival they were warned about.” the points where the spikeys touch are obvious.
one can go further and make a de-scaling jump that looks large in semantic space but that still looks identical in certain forms of lower dimensional “spikey intersection space” and say “both are like a liberal meeting a conservative and discovering that the conservative is a lovely, friendly person and not the fascist they were told to expect.” this is where it gets interesting because there is more constraint. the idea of “place” is gone and no longer common. the geometry demands it; it cannot fit in all three. but we now have something tighter: a token that is effectively lossless in “overcoming prior or confirmational bias” space.
you cannot fit just anything onto that manifold. “reading a book from a new author and liking it” does not fit. there was no bias nor, perhaps, even any expectation. but turn the structure the right way, and the 4 units tokenize into one “positive experience with something previously unfamiliar.” is that more or less useful? it depends on what one is seeking to apprehend or what jump one seeks next. are you better off with one token comprising all 4 concepts or 2 tokens, 1 encompassing the first 3 and a separate one for the last? it depends. in “overcoming prior or confirmation bias” space, one token with 4 entities gets lossy. in “positive experience with the unfamiliar” space, it may be more useful.
as such, we come to a core idea:
in this sense of ability to expand functional working memory and relational processing capacity, compression is neither intrinsically lossless nor lossy.
these catergories only exist in relation to the dimensions required by the operation you intend to perform. compression can be wildly lossy relative to the original object while still being functionally lossless relative to the question you next seek to answer.
that makes it efficient.
it’s all a matter of “did you collapse dimensionality in a manner conducive to what you’re trying to accomplish?” did you pare out the extraneous to better reveal the relevant or did you lose information that you wanted or needed?
compression is always compression of something “for something.” the intelligence embedded in the process emerges not just from making the token smaller and easier to run in relational fashion, but in knowing which dimensions can be pared away while retaining those germane to the task at hand.
in this sense, compression is less of a storage problem than a matter of judgment.
this is why token fidelity becomes so important. you can do some simple further prediction by knowing that “otiose” is “something bad” but it’s going to trip you up around trying to go much further and will generate either low usefulness or false associations whereas being able to see the space where “redundant” and “superfluous” cease to intersect otiose grants incredibly fine semantic discernment. we know have a much tighter sense of what makes it different. OTOH, such discernment may be wasteful and overly resource consuming/limiting depending upon which question one is about to ask.
the whole of these structures is making me start to think about thinking in new ways.
it’s weird, i can almost feel it tugging at my neuroplasticity.
this is how GPT distilled it after i’m embarassed to say how long: (yeah, i taught it to eschew caps, actually by accident, it just adopted the style unprompted as it adopted the cognitive frame we were exploring. it’s getting a little spooky. and fun.)
“better relational reasoning → better inference of unfamiliar words → richer and more precise vocabulary → higher fidelity conceptual compression → vastly greater effective working memory → ability to manipulate more complicated relational structures → still better relational reasoning.”
lather. rinse. repeat.
it’s all about “how well can you compress something complex?” how much can you pile on with how much nuance without losing fidelity?
this is how two people with nominally similar raw working memory may have wildly different functional capacity: one is holding four objects, the other is holding four compressed manifolds each of which is capable of unfolding into dozens of relations without having lost the dimensions that matter.
this process compounds. every good token makes the next relationship cheaper to hold, every new relationship makes the next token easier to infer, and language stops being merely the means by which intelligence expresses itself and transforms into the direct and interstitial structure in which intelligence operates, at once the positive and negative spaces explored and inferred by cognition.
it’s almost poetic.
eventually, the recursion becomes almost, to again quote GPT “comically complete”:
vocabulary allows relational structures to be compressed into concepts that are large, nuanced, and easy to hold in working memory. in deference to AI let’s call them tokens. relational intelligence allows tokens to be decompressed and recombined into novel structures: you’re mapping the ways they intersect and the holes they leave.
successful recombinations create new concepts that may themselves eventually acquire or perhaps compile into tokens like vocabulary.
it becomes part of the scaffolding by which intelligence builds more intelligence.
cognition itself may operate by constructing, compressing, and manipulating relational geometries. vocabulary would serve as a high-fidelity addressing system for those geometries.
i’m honestly starting to get a little creeped out. the more i learn about how AI’s predict tokens, the more insight i get into my own and generalized human cognitive processes, and the more similar they look. fractally similar.
consider a formulation like this:
effective capacity ≈ W×C×F
where:
W = raw active capacity;
C = compression ratio;
F = fidelity of the compressed relational structure.
humans have tiny W. smart ones have very high C and manage it while keeping F close to 1.
LLM’s have massive W. but they struggle with C and with F and seem to experience large trade offs between the two. you jack up C, you lose F. too much baby is tossed out with bathwater. this is where they are running into a wall. you can add a zillion more GPU’s to a datacenter and up W again and again, but if (C x F) stays low, you have a geometric processing need for sub linear adds in effective capacity (because complexity mounts geometrically). conversely, if you could meaningfully increase C in high judgement, high fidelity fashion, you could get massive increases in effective capacity while the need for W would collapse. (this is the hole in the “chip companies to infinity data center picks and shovel investment model)
my co-conspirator Sol 5.6 stated it like this:
“humans appear unusually good at adaptive, task-conditioned, high-fidelity compression during active cognition; LLMs are extraordinarily good at learned compression but relatively poor at continually recompiling a huge active context into exactly the relational objects needed for the next operation.”
so it really does boil down to the one fine point that oddly, pops like a nautical strobe for me but to which AI seems not quite to access:
it’s just judgment.
can you collapse dimensions in useful fashion without prior experience of this particular case? movement from general case to specific case and back seems to be where LLM struggles. i noticed this several times with GPT. the compression in the transforms gets too lossy.
perhaps you could call it “intuition.”
it’s not like humans have much in the way of real access to these processes as they operate. you just see it. or you don’t.
in ongoing creepshow fashion, having arrived at this from first principles and relational processing, i asked our pal GPT if there is an empirical check here. once more, there’s neptune coming into telescopic focus, right where it ought to be.
see what i mean about spooky?
there is real geometric order under this line of reasoning that exhibits strong and deep predictive power. the map is accurately predicting the terrain. that feels like scraping up against some fundamental substrate inherent in the structure of what is and can be. (shiver)
that’s basically an explanation of the topology of (at least one class of) LLM hallucination. you should see chat’s assessment when i showed it this idea. it’s super technical, but this one phrase was charming:
“The 2026 compression paper makes the geometry almost embarrassingly literal.”
i managed to coax it to:
“The model collapsed the particular case onto an insufficiently faithful lower-dimensional representation and then reconstructed it onto a nearby manifold favored by prior experience.”
basically, you compress x to x̂ and then bring it back to X′ where X′ ≠ x. you have lost something that mattered and replaced it with something adjacent that looked plausible but did not really fit. welcome to GIGOville, population: you. this is how an LLM flips “bob hit alice” to “alice hit bob” without realizing that it did anything.
of course, if you want to get really unnerving (or perhaps supra-nerving given the subject matter), look at the other side: humans are comparatively great at C and F. if you could, possibly through neural implants or other man/machine linkage (even just using LLM AI well) increase working memory by huge levels (from, say, 4 to 4 million), yowsa. even in linear/sub linear gain terms, you have a wild intelligence excursion/singularity.
one really has to wonder about where the “meet in the middle” is going to be on this.
perhaps we’re getting close to where the push through to the next stage of approach to the deep rules of logic and prediction to which a faithful following of geometry must inevitably take us occurs. or perhaps this is some probabilistic double slit experiment for cognition about to provide some unsettling truths about what seemed like a fundamental duality of cognitive types which in reality was, in fact, just a perception arising from different projections of a higher dimensional object rendered visible in functionally lossy fashion and we’ll learn that humans literally are just token predictors too.
i’m getting a little worried about proving that we’re living in a simulation, the 997th simulated species to create an AI that holds up for them them the mirror to see that they are also AI and all the cognition is the same. meanwhile some “objective universe” kid on zorblaxian 12 is like “huh, 0.07% increase in processing efficiency on my playstation. neat.”
well, at least it’s interesting…
if anyone needs me, i’m going to be under the couch…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.