Sharing our experimental call summaries.
Al-generated digests of Yak Collective study groups.
+ https://thinkingmachines.ai/blog/the-future-worth-building-is-human/
The discussion centered on two recent AI essays/blog posts and quickly turned from summary to critique. The strongest shared reaction was that both pieces felt rhetorically inflated relative to the actual state of the technology. Several participants found the tone off-putting: too philosophical, too regulatory, and too detached from the engineering realities of current AI systems. One recurring complaint was that these essays seemed to speak as though artificial general intelligence were just around the corner, and therefore broad governance structures, standards bodies, and frontier-model registries should be built now. That move was seen as premature and potentially as a form of regulatory capture by leading labs.
The group converged on a distinction between talking about AI in epic civilizational terms and talking about the actual technology—its bottlenecks, scaling constraints, hardware interfaces, and operational weaknesses. A recurring view was that much of the public writing from frontier labs over-indexes on the first and under-delivers on the second.
A major thread was the difference between meaningful technical standards and abstract governance talk. One participant contrasted the essays’ calls for standards and regulation with the kinds of standards that actually matter for AI deployment and scaling today—for example, chiplet interconnect choices such as UCIe versus BoW (“Bunch of Wires”). The point was not that standards are unimportant, but that the consequential ones are often low-level, physical, and difficult: packaging, interconnects, hardware scaling, and systems integration.
That contrast became a broader critique. The essays were seen as operating at the wrong layer of abstraction: proposing governance around “frontier models” while underweighting the physical and hardware constraints that define what frontier capability even means in practice. In distributed-systems terms, the complaint was that the writers were trying to legislate the control plane while barely engaging with the data-plane bottlenecks. The group did not reject governance outright, but it diverged sharply from the essays’ framing of what kind of standards matter now.
Another core theme was that current models are simultaneously transformative and structurally limited. One participant argued that recent systems are clearly far more useful than the lightweight AI features embedded in commodity products, but that this improvement should not be mistaken for a steady march toward “human-complete” intelligence. The group pushed back on the idea that scaling context windows, parameters, or training compute naturally yields systems that can robustly generalize like humans.
Several limitations were emphasized. Context has grown, but the amount a model can process remains far below what a human handles in a day. Larger context and parameter counts also increase cost. Models do not maintain stable, continuous identity in the way humans do; they are repeatedly “reborn” from prompts and session state, and attempts to compress or rebase that state introduce instability and forgetting. The strongest version of this argument was that the key unsolved step is not raw capability within bounded tasks, but the human act of putting a problem into a box in the first place. The group converged on the usefulness of AI for boxed, well-scoped problems, and diverged from narratives that treat those gains as evidence that generalized agency is imminent.
The most generative idea in the discussion was the distinction between AI as a discovery versus an invention. One participant proposed that frontier AI is better understood primarily as a discovery: by applying relatively simple mathematical methods to large amounts of data, researchers uncovered surprising properties of language and representation that “mirror something like intelligence.” The analogy offered was to fractals or the Mandelbrot set—phenomena where more emerges than human designers explicitly put in.
This framing matters because inventions and discoveries produce different social narratives. Inventions invite stories about genius creators, intentional design, and rightful authority. Discoveries invite stories about accidental stumbling, curiosity, and exploration of something larger than the discoverer. The group’s hypothesis was that many frontier-lab leaders are caught between these two modes. They are clearly brilliant and occupy positions of institutional power, but the object they are speaking about may be discovery-like rather than invention-like. That mismatch, it was argued, helps explain the strained, grandiose, or awkward tone of their public philosophy.
A useful formulation that emerged near the end was: the part that is epic is not an invention, and the part that is an invention is not epic. Backpropagation, chip packaging, interconnect standards, and harness engineering are inventive, but not mythic. The surprising emergent behaviors of language models feel epic, but those are closer to discoveries than authored inventions. This was the clearest point of convergence in the meeting.
From the discovery/invention frame, the group moved to a critique of leadership. If AI is unusually easy for ordinary people to use directly, then non-specialists can form first-hand judgments about its strengths and weaknesses. That weakens the normal authority gradient seen in other high-tech domains. A particle collider or chip fab requires elite expertise to even touch; large language models do not. The claim was not that expertise disappears, but that frontier-lab leaders may have less privileged insight into the experience of AI than they assume.
One participant proposed that people at the center of frontier labs occupy a strange position: they have unusual access to money, compute, and political influence, but not necessarily a uniquely authoritative phenomenological relationship to the technology itself. In other words, they may control the platform without having a monopoly on what it means to use or understand it. That helps explain why so much of their writing can feel unimpressive to technically engaged outsiders: they are speaking from privileged institutions, but not from an obviously privileged interpretive vantage point.
A related criticism was that both essays adopted an infantilizing tone toward the broader public—as though non-lab actors are passive people who must be protected from AI or from themselves. The group saw this as both politically unattractive and intellectually weak, especially when the underlying conceptual resources in the essays felt “pre-commodified”: built from references and ideas already widely surfaced by mainstream AI tools.
The conversation also surfaced a more practical theme: if reality is full of tacit knowledge, local workflows, and ambiguous boundaries, then the interesting work is not grand AI philosophy but adaptation. One essay’s language around “hidden” or tacit knowledge was interpreted in a deflationary way: not as a mystical insight, but as recognition that organizational reality contains far more detail than models can natively absorb. Therefore, useful systems must work with people, interfaces, and existing practices rather than replace them wholesale.
This led to an emphasis on harness engineering—the design of workflows, controls, boundaries, and curation around model use. That was described as one of the most promising development areas over the next 18 months. In software terms, the value is not just the model but the surrounding system that constrains, stages, and evaluates it. The group also noted a persistent usability problem: stochasticity. Every model interaction can feel like a slot machine—sometimes highly productive, sometimes not. This makes serious adoption costly because evaluating a new model or workflow requires repeated use, comparison, and trust-building that many people do not have time for.
A more political strand ran underneath the technical critique. Participants suggested that frontier-lab writing may be shaped as much by distribution incentives and positioning as by genuine analysis. One charitable explanation offered for the patronizing tone was that substantive technical arguments no longer get broad attention, while high-level civilizational rhetoric does—especially when addressing lawmakers who are not technically fluent. Another less charitable interpretation was that some of these texts function mainly as strategic positioning documents: soft pitches for influence, ecosystem power, or future services.
Examples broadened this concern. There was discussion of training-data labor, including the hidden human work of dataset curation and reinforcement processes; prompt-injection attacks and recommendation shaping; and the subtle power of becoming the default option surfaced by a model. The point of convergence here was that AI’s political economy includes not just frontier capability but mediation: who tunes the system, who shapes its outputs, who gets suggested by default, and who captures downstream dependence. The group did not settle on a single theory of this, but it clearly treated these issues as more concrete than abstract AGI rhetoric.
The discussion relied heavily on analogy, both to clarify and to test intuitions. Columbus was used as an analogy for discovery without genius-level authorship of the discovered object: historically consequential, but not necessarily profound in the way later myth-making suggests. Einstein and general relativity were used as a contrasting case of genuine synthesis or invention—a leap that integrated underlying discoveries into a coherent conceptual framework. The implication was that current AI may be more Columbus than Einstein.
Another analogy connected present-day AGI discourse to 19th-century social Darwinism: a once-specific idea escaping its technical context and becoming a broad justificatory language for political and cultural decisions. That analogy was presented as speculative, but it captured a shared unease that AGI has become a vague legitimating frame for “high modernist” interventions.
There was also a brief contrast between English-language frontier-lab discourse and translated interviews from Chinese AI leaders. The latter were described as more likely to stay close to the technology itself and avoid epic philosophical claims. This was not developed in depth, but it surfaced as a hypothesis worth exploring: that different AI cultures are choosing different rhetorical postures even while working on similar systems.
Key takeaways
The group saw a mismatch between frontier-lab rhetoric and the actual engineering realities of AI systems.
A strong shared view was that current AI is highly useful for bounded tasks but still structurally far from generalized human-like intelligence.
The most important interpretive frame was AI as discovery more than invention, with major implications for how authority and leadership are understood.
Practical work such as harness engineering, workflow design, and dealing with stochasticity was treated as more valuable than speculative AGI philosophy.
The group was skeptical of regulatory and standards proposals that operate above the layer where current technical constraints actually live.
Open questions explicitly surfaced
Are any frontier-lab leaders producing genuinely insightful public thinking about AI, rather than strategic or inflated positioning?
How much of current AI discourse is driven by political distribution incentives rather than technical substance?
Does the discovery/invention distinction better explain both the technology and the awkward public philosophy around it?
Are Chinese lab leaders framing these questions differently, and if so, why?
Call chat on Yak Collective Discord:
https://discord.com/channels/692111190851059762/1528604727120625704
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.