When I turned 16, I couldn’t wait to get my driver’s licence. I studied the Ontario Ministry of Transportation handbook assiduously and took the written test as soon as I could, passing on the first try. But when my father then put me behind the wheel of our late 1970s Chevrolet Chevette with manual transmission, I soon realized that the real experience was far different from book learning. Undoubtedly, my hesitant attempts at simultaneously shifting gears, steering and accelerating or braking contributed to Dad’s grey hair, but with the help of a professional driving course facilitated by my high school, it didn’t take too long for me to drive both safely and confidently. Naturally, my turn to grey a little prematurely came when my kids each turned 16 and I had my own thrills in the passenger seat with an adolescent in control of the vehicle. But I’m proud to say they both learned quickly and thus far have unblemished driving records of their own.
Thank goodness I never tried to teach a computer how to drive, nor that my kids got their driving lessons from ChatGPT.
In the conclusion of my last post, I introduced Meta head of AI Yann LeCun’s comment that the average teenager can learn to drive in under 20 hours of instruction and practice, but after decades of research and billions of dollars spent on technology, AI has not yet achieved the same feat. LeCun is skeptical that any further progress can be had with current generative AI technologies. In the same post, I also mentioned the departure of Mira Murati from her position as Chief Technology Officer at OpenAI, and the founding of her new company, Thinking Machines Lab. In contrast to LeCun, Murati is optimistic about the future of generative AI but believes it needs more governance and oversight. An article last year in Wired magazine quotes her as saying “The technology is not intrinsically good or bad” but that it’s up to humans to guide AI towards good outcomes.
Let’s take a closer look at both LeCun and Murati’s positions on AI and the possibility of achieving Artificial General Intelligence.
From Yogi to YannYann LeCun is a co-winner, along with Geoffrey Hinton and Yoshua Bengio, of the 2018 Turing Prize—the most prestigious award in the field of computer science. The three made pioneering contributions to machine learning, neural networks, backpropagation and other fundamentals of what we now know as large language models (LLMS) and generative AI. LeCun developed the first machine learning systems for handwriting recognition in the 1980s, technology that is still used by banks and post offices today.
It’s interesting, then, that at a technology conference in 2024 he told students entering the industry to “Stop working on LLMs … there’s nothing you can bring to the table. You should work on next-gen AI systems that lift the limitations of LLMs.” But his perspective becomes clearer when we realize that he developed the concept of convolutional neural networks, which are the foundation of computer vision and image processing systems.
LeCun’s critique of generative AI had me thinking about the famous quote from Yogi Berra: “You can observe a lot by watching.” His point, as detailed in a recent Newsweek article, is this: as of 2025, the amount of data that has been fed to a typical large generative AI system is roughly 10^14 bytes. This, he estimates, would take the average human about 400,000 years to read.[1] By contrast, LeCun tells us that the bandwidth of the human optic nerve is about one megabyte per second. A four-year-old child has been awake for probably 16,000 hours in their life, and therefore can, in that amount of time, visually absorb 10^14 bytes of information. To paraphrase Yogi, clearly, one can observe a great deal more by watching than by reading.
So the problem, according to LeCun, is not the amount of data but how the data is received—and reading text is a pretty inefficient way of absorbing information. He goes on to comment that, although we have sophisticated LLMs that can read text and carry on conversations, we still don’t have robots to do the household chores, nor even a robot that can move as precisely in three-dimensional space as a cat can. As Newsweek puts it, “The strange paradox of LLMs is that they have mastered the higher-order skills of language without learning any of the foundational human abilities.” This is because language is considered to be, mathematically, low-dimensional. Reading text, speaking or producing tokens as LLMs do, is basically linear. We process one word, one sentence, one paragraph after another. But when we look around us, walk or pick up an object, we are intuitively experiencing our environment in all three dimensions at once. This is a higher-dimensional and instinctive way of interacting with the world, and we have not yet been able to teach the machines to do it. And that is why we still don’t have self-driving cars—we’re trying to apply a low-dimensional, linear solution to a high-dimensional interactive problem, just like my own early experiences in making the transition from reading the driving handbook to being in the driver’s seat.
Jump to JEPANaturally, Yann LeCun has something else up his sleeve. For the last few years, he’s been building on his background in computer vision and image processing to come up with a new approach to artificial intelligence. In a 2022 paper, he outlines his ideas, which he calls Joint Embedding Predictive Architecture (JEPA). He emphasizes that JEPA is completely different from generative AI, which essentially works according to the rules of linear algebra, trying to predict the value of variable y given a known value of variable x. JEPA, by contrast, works in what LeCun calls a representational space with latent variables—variables that do not appear directly in the equations but can be inferred from the context.
I won’t go into all the mathematics behind it, but these concepts are key to the application of machine learning to image processing. Give the ML algorithm an image formed from pixels, and nothing further to go on, and let it try to make sense of the image. The algorithm will use self-supervised learning, trying to identify patterns of similarly coloured pixels, including and excluding pixels on the boundaries, until it finally identifies what the picture is.
An analogy would be, say, completing a 1,000-piece jigsaw puzzle. Once you’ve built the frame, the rectangular outer edge of the puzzle, you have your representation space. Now you must fit in all the pieces inside the frame. The variables you know are the knobs and sockets of individual pieces, their vertical or horizontal orientations and their colours or patterns. The latent information you don’t know would be associated with the meaning of each piece—is a white piece with blue part of a cloud, or part of a wave on the water? Is a green piece tree leaves or grass? You start by grouping pieces of similar colours or patterns, and testing how they fit with each other. You get some small clusters of pieces, re-arrange your remaining pieces, and keep trying until eventually the whole puzzle comes together—as long as you haven’t lost any pieces along the way.
Recognizing an image or solving a puzzle is a single instance of the machine learning process. LeCun’s insight is to work with multiple images and even video content (V-JEPA) simultaneously. He describes techniques for training a JEPA system to learn how to ignore content of low-relevance and focus on what is most meaningful, thereby making predictions about probable outcomes.
He returns to the analogy of driving a car. The representation space is everything the driver experiences—what they see out the windshield, the feeling of the steering wheel in their hands and the pedals at their feet, the information displayed on the dashboard. A driver quickly learns to ignore irrelevant content like the scenery off in the distance and pay closer attention to oncoming traffic and road conditions. The speed of the car and the density of the surrounding traffic allow the driver to predict with some accuracy a latent variable like the probability of safely navigating an upcoming curve. Another latent variable might be the time of arrival at their destination, which can be predicted with a wider range of error based on speed, weather conditions or hazards like road construction or lane closures.
JEPA relies on additional concepts like statistical mechanics and calculating low-energy states in physical systems, mathematical constructions that are directly applicable to minimizing cost functions in optimization—and also used in machine learning to determine how close an algorithm is to achieving a meaningful result. This is probably due to LeCun’s close collaboration with Geoffrey Hinton, who pioneered the application of Boltzmann distributions to machine learning.
But LeCun differs sharply from Hinton in his assessment of how imminent Artificial General Intelligence, or AGI, might be. To Hinton’s assertion that there is a 20 per cent probability of rogue AI exterminating humanity within 30 years, he responds “that’s completely false.” He suggests, correctly, that it would require AI systems to not only be intelligent but also be able to operate in and control the physical environment—which, as he has demonstrated, is currently far from possible. Of course, if JEPA eventually comes to fruition, this problem will be solved, but LeCun further points out that we “give too much credit and power to pure intelligence.” Intelligence, he notes with only a hint of sarcasm, is clearly not a necessary attribute for human political leadership. In addition, we will keep building and improving the controls and guardrails around AI systems—humans have proven to be a pretty resilient species, after all.
Shifting gearsMira Murati shares Yann LeCun’s optimism that AI is controllable, but differs from his pessimism about generative AI. After all, as the former CTO of OpenAI working with Sam Altman, she did as much as anyone to bring the current generation of generative AI systems to the market. She disagrees with LeCun’s assessment that AGI is still far off, saying in a 2024 interview in Wired magazine that “There’s not a lot of evidence to the contrary. Whether we need new ideas to get to AGI-level systems, that’s uncertain. I’m quite optimistic that the progress will continue.”
Murati has the authority to back up her comments on AI, having worked at Tesla prior to OpenAI and subsequently having founded Thinking Machines Lab with an almost unheard-of initial $2B in financing from a consortium of Silicon Valley investors. Here’s an indication of the perceived threat she poses to the AI establishment: earlier in 2025, Yann LeCun’s boss and CEO of Meta, Mark Zuckerberg, decided to launch an all-out raid on Thinking Machines, offering nine- and 10-figure compensation packages to key employees if they would leave their position and join him at Meta. Not a single one took up Zuckerberg’s offer[2].
What’s interesting, though is that Murati has shrouded her work with Thinking Machines Lab in quite a bit of secrecy. The company was incorporated as a Public Benefit Corporation, which commits it and its board of directors to consider not just financial profitability but also the social and environmental good of its activities—often referred to as the “triple bottom line.”[3] Its web site is a single page with a high-level description of the company’s objectives and a link to its single product thus far—Tinker, a tool for optimizing the training and fine-tuning of most commercial LLMs.
So, what is Mira Murati’s vision? To figure it out, we have to look at the few interviews and articles available online, and try to read between the lines of her one-page web site. A recent profile in Fortune provides a few details. After a couple of years at Tesla working on their electric SUV, she joined OpenAI in mid-2018, quickly leading development teams for all the company’s products—ChatGPT, Dall-E for image generation, Codex for software development and Sora for video generation.
She has occasionally courted controversy—for example, in a 2024 speech at Dartmouth College she commented on AI’s impact on the arts by saying “Some creative jobs maybe will go away, but maybe they shouldn’t have been there in the first place.” She was criticized for being insensitive, but perhaps she was partly right. It might be too soon to tell the impact of AI-generated “actress” Tilly Norwood on the film and television industry, but, on the other hand, human visual artists like Refik Anadol[4] are incorporating AI into stunning compositions while remaining sensitive to the environmental impact of their work.
Despite this, Fortune notes that Murati has consistently advocated for the responsible development of AI with guardrails, and, disagreeing with many or most of her Silicon Valley colleagues, is even in favour of government regulation of the industry. Her launch of Thinking Machines Lab emphasized the themes of accessibility, customizability and human-alignment in AI systems but offers few technical details in how this will be achieved. I’m reminded of IBM’s commitment to responsible AI, while many other AI technology vendors have quietly disbanded their AI ethics teams.
Thinking Machines Lab’s Tinker allows AI professionals to train and fine-tune their models much faster and more efficiently. It’s an interesting choice of name, suggesting that all that current generative AI needs is some tinkering in order to move ahead in the direction of AGI. I don’t think that’s the case, and Fortune also notes that Tinker reflects Murati’s belief that the future of generative AI is not in ever-larger models, but in democratizing access to AI with more flexible and adaptable tools. This, again, would appear to be similar to IBM’s approach with WatsonX, building smaller fit-for-purpose models and offering tools like InstructLab to manage and curate the underlying data.
I’ve always been a fan of IBM’s approach to AI, and I’m happy to see it borne out and supported by such an industry heavyweight as Mira Murati. They may be approaching AI from somewhat different perspectives, but I believe their objectives are aligned in producing safe, reliable and responsible generative AI tools.
The big questionYann LeCun and Mira Murati’s roads diverge when it comes to the future of AI, and, to paraphrase Robert Frost, both roads are less-traveled. But will Murati’s approach deliver on the promise of near-term AGI? I don’t think so. The improvements she is working on are necessary and welcome, but fundamentally generative AI is still a closed system, needing something more. Will LeCun’s JEPA eventually get us there? I don’t know for sure; however, I would agree with his assessment that it’s still a long way off—and I’d like to see a working self-driving car first, to take us down that road.
But maybe if, how, or when we achieve AGI are the wrong questions to ask. Maybe the most important question of all is, why do we want to?
[1] By my own rough calculations, this assumes you could read one double-spaced page of typewritten text approximately every 30 seconds, continuously, without breaks for eating or sleeping.
[2] Except maybe one. The Wall Street Journal reported on October 11 that Andrew Tulloch, a co-founder of Thinking Machines Lab, accepted a six-year, $1.5 billion offer to join Meta. Such pay packages only illustrate the difficulty of recruiting talent and are not sustainable over the long term.
[3] OpenAI, similarly, maintains both non-profit and for-profit entities.
[4] I was lucky enough to see Anadol’s installation at the Guggenheim Museum in Bilbao, Spain last summer and was very impressed with his work.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.