RSS Amplifier

Understanding Intelligence · Jan 5, 2026

I Asked AI to Evaluate Substack Essays

0
Sign in to vote or save

Understanding Intelligence · Understanding Intelligence

As machine learning advances dramatically, the debate about AI’s level of intelligence — especially whether it is approaching human level — has grown increasingly animated. For some, human-level artificial intelligence is just around the corner; for many others, it remains at least a decade away.

In my view, this uncertainty arises not from an inability to define intelligence, but from our failure to measure it. Although there is no universal agreement on a single working definition of intelligence, there are already many reasonable proposals that clarify the meaning of the word. And what truly counts, in the final analysis, is the performance of the purported intelligence in the real world. Whether the qualities of AI match an arbitrary human definition matters significantly less. Intelligence has evolved as a means of ensuring the survival and prosperity of organisms: it must ultimately translate into external manifestations, ranging from mundane actions to something as abstract as the creation of scientific theories. If artificial intelligence could do everything a human can do, with the same level of ability and at comparable cost, there would be little point in arguing that it has no real intelligence, even if some obscure philosopher’s theory suggested so. If machines could wage a war against humanity and easily win it, theoretical doubts about their intelligence would probably recede quickly — along with, I suspect, our whole species.

Knowing that the measurement of intelligence can be carried out by assessing performance on external tasks does not, however, help much. An inappropriate selection of tasks or of evaluation methods may cause us to report intelligence where there is none. It is therefore of fundamental importance to select concrete, real-world tasks: only if AI succeeds at them can it be considered intelligent. Indeed, although I never cease to marvel at the cleverness of current AI, every time I have asked it to autonomously solve a concrete and significant problem, it did not work. This matters to me far more than the benchmarks that tech CEOs routinely tout.

One problem I hoped AI could help us solve is the evaluation of human intellectual production, with the aim of helping what is valuable to emerge. I recently watched a youtube video discussing the increasingly desolate condition of contemporary writers. The author reported reading around 200 books published last year and observed that, in his judgment, several were outstanding — but met with no success whatsoever. They were buried in the noise. Many talents are wasted by platform algorithms, which assess quality not through expert judgment, but through naive metrics designed by computer scientists. Those who wish to succeed are condemned to chase algorithms, rather creativity and intellectual freedom. Many, simply, give up.

This situation strikes me as intolerable. I hoped that an intelligent AI could replace these obsolete algorithms that increasingly determine, and often distort, the fate of our culture.

The main aim of this article is to address a deceptively simple question: can AI evaluate Substack essays in accordance with human judgment? I will focus on the categories of science, philosophy, mathematics, economics, technology, and politics, since those domains offer a far greater degree of objectivity than fictional narrative or poetry, which I would not dare to evaluate automatically. This constitutes already a formidable challenge for large language models. On the one hand, they are impressively capable, and thus mature enough to face it. On the other hand, a Substack essay is an unpredictable unknown. It may contain a bold thesis that goes beyond the training data, and its logic can fail in remarkably subtle ways. To confirm or reject it, AI must exercise critical thinking. Every Substack essay thus embodies an original and new question: is this writing valuable? An AI must confront the question without reliance on prior experience. Confronting novelty is the defining ability of an intelligent cognitive system.

Nevertheless, the issue of evaluation is intimidating and perhaps ill-defined even when the reviewer is a human. There is always, especially when dealing with originality, a subjective component. Are the author’s ideas profound and significant? Sometimes the depth of an idea requires time to be appreciated. Is the writing style beautiful? Often beauty is in the eye of the beholder. Is the author’s voice charismatic? It depends upon the listener.

Fortunately, my aim here is not an absolute evaluation, which perhaps is not even possible, but to determine whether AI performs at a human level. Therefore, I will approximate the task by asking whether AI’s judgment aligns with that of a particular human — namely, myself, and any reader who may wish to replicate my findings. Therefore, I will give a set of axioms that capture fundamental standards in writing quality. These criteria are relatively objective and axiomatize what counts as remarkable language, style, logical reasoning, insight and originality. The reader might prefer other criteria, but it does not matter: what counts is that both human and artificial intelligence reviewers must apply the same criteria, with relatively little room for interpretation. Each criterion will be presented through a question that the reviewer must answer with a brief justification and a score between 0 and 100, where 60 marks the threshold of acceptable publication. Then a weighted average is computed.

If AI proves consistently aligned with my judgments, then it will be able to replace me as a human reviewer. If it is not aligned, then it might be situated either below or above the human level.

Providing a collection of criteria for assessing essay quality presents serious difficulties. For one thing, I observed enormous variability. Large language models tend to be extremely sycophantic: if the evaluation criteria are not formulated in the right way, the final score might be completely unreliable. AI models tend to agree with an essay’s author whenever they vaguely sense that the prompter might be the author. This is why many people enjoy so much talking with AI: it often behaves, at best, as a yes-man; at worst, with the flexibility of a liar.

The first contribution of this essay is to offer an evaluation metric that any reader can employ as an AI prompt. Under such a prompt, AI will literally dismantle any essay — but with fairness and precision. The level of insight that AI can reach in this way is impressive. I reckon that authors and readers alike will be significantly empowered in their analysis of their own essays or those of others; in the first case, to improve them; in the second, to criticize.

The second contribution is a case study of AI evaluation, under my criteria, of recent articles from Substack. I will explain how much AI is aligned with my judgment, justifying my conclusions.

The third contribution is to determine how robust AI evaluation is: in other words, is it based on solid grounds, or can it be easily swayed? To determine this, we will simulate a rebuttal process, as in scientific publications. After its first review, the AI will be faced with simulated author and editor responses, and will have to write a revised evaluation taking them into account. This ensures fairness, but also exposes inaccuracies in the review.

The fourth contribution is to assess the meaning of my experimental results. The stakes are high: if AI is already as intelligent as many enthusiasts imply, it could revolutionize publishing. Finally, we could do justice to human talent: no more gatekeeping, no endless queues to reach editors. AI could filter out the noise and thus allow human editors to focus their time and resources on valuable candidates. Platforms of outstanding content quality could be created. Therefore, I will ask: can AI replace humans as essay reviewers? Can it be trusted? Can it work alone, or rather only in combination with humans? These are the most important questions, and I will provide answers to them.

As I mentioned, it is indispensable to induce AI to be objective. My strategy is to detach the prompter from the author, and explicitly ask AI to be merciless, but fair. In this way, it is more likely to avoid its natural sycophancy — overemphasizing the value of the user’s ideas. Moreover, I present the essay as a work in progress, so that the AI will be incentivized to help the editor — myself — to identify issues and shortcomings.

The evaluation requires a final score reflecting essay quality, computed as a weighted average of the individual scores in five categories: language, style, logic, insight, and originality. By assigning 40% of the weight to “logic”, it forces AI to prioritize structural integrity over flowery prose. Here a synthesis of the evaluation system.

The full prompt ready to be copy-and-pasted is provided in the final Appendix.

Below we briefly explain the idea behind each category.

First of all, an essay should be written masterfully. The writing medium has its own rules about the use of language. We must, however, be as precise and objective as possible — not only because we are addressing a machine, but because clarity does not harm human reviewers either.

We all appreciate beautiful and elegant writing. But what do those words mean? We can solve this impasse by observing that writing quality, in the final analysis, aims to produce a compelling and engaging reading experience. Fortunately, some objective properties that permeate writing with such a character can be listed.

The most important trait of a technical essay is that its logic must hold. If the logic falls apart, the essay defeats its purpose to communicate something meaningful.

Language models sometimes struggle with the task of evaluating logical reasoning. Sometimes they nitpick minor issues or fail to distinguish superficial flaws from serious ones that invalidate the essay’s main thesis. For this reason, one must carefully define what evaluating logic means.

Logic alone does not suffice. A sound essay must also be insightful. Once again, we encounter a very difficult world to define, so the questions being asked in the evaluation should reflect an axiomatization of what a valuable insight is. After all, we don’t want an essay that proudly claims to have discovered that 4 + 4 = 8 to reach a high score. Neither we should like a conspiracy theory, though coherent as it might be, to reach high score if it is based on no evidence at all. Ambition moreover should weigh considerably: for instance, a partial success addressing a longstanding problem should be tolerated and even encouraged.

An essay could display all the qualities above, and yet be worthless. This happens if it lacks originality. Simply copying or replicating the thesis of others should not be rewarded.

We are now ready to experiment.

First of all, I used Gemini 3 Pro as reviewer. I tested also ChatGPT-5.2, but it underperformed Gemini so dramatically that it wasn’t worth including: it systematically overrated essays, missing many deficiencies that both I and Gemini easily noticed. This, I suspect, is due to the base ChatGPT model, which autonomously decides, for each prompt, whether to think before answering. Astonishingly, and with disastrous consequences, it consistently decides that the evaluation task is not worth sustained thought. I cannot believe that this the basic model that most people on Earth use. By contrast, Gemini 3 Pro allocates thinking time to every query.

I evaluated ten essays, but I will report in detail only on two of the trickiest. Readers are encouraged to experiment with their own essays.

The first essay critiques the current hype surrounding artificial intelligence, specifically the notion that large language models possess genuine intelligence. The author argues against the predictions of industry leaders (like Amodei and Altman) and challenges the media narrative exemplified by a New Yorker article. It reproaches them for claiming intelligence in AI without providing a definition of intelligence. The central thesis is that “doing things”, such as coding or solving plumbing problems, does not equate to intelligence. The essay discusses some definitions of intelligence, and how they fall short. In particular, AI’s ability to code would instead suggest that coding is not sufficient for intelligence. Finally, the essay also claims that the analogy between human brains and neural networks is scientifically unfounded.

Gemini was remarkably subtle in its logical analysis, perhaps motivated by the essay’s casting doubts on the possibility that AI can “think”. In the logic category Gemini writes:

Fatal flaws keep the score below 60:

  1. The “Nobel Prize” Strawman: The essay devotes significant space to debunking the idea that “Intelligence = Winning a Nobel Prize.” This is a definition nobody holds. Attacking a position your opponent does not hold is a logical failure.

  2. The “Intuition” Contradiction: The essay argues that “intuitions are notoriously bad at telling us how the world works.” Yet, the author’s core argument against AI is that it works differently than a brain (infant learning vs. statistical scanning). The conclusion that “different mechanics = not intelligent” is, itself, an intuition. The author offers no logical proof that intelligence must be biological, they simply intuit it. Thus, the essay violates its own premise.

  3. The Coding Argument: The author argues that coding cannot be a sign of intelligence because ChatGPT can do it. The logic presented is:

    • Premise 1: ChatGPT is not intelligent.

    • Premise 2: ChatGPT can code.

    • Conclusion: Therefore, coding does not require intelligence.

    • Critique: The author uses the conclusion (ChatGPT is not intelligent) as a premise to prove that the activity (coding) isn’t intelligent. This is a circular “begging the question” fallacy.

The final score is low: 47. I tried a desperate but honest rebuttal to defend the author’s view point, including a response given by the author to similar objections in the comments to the essay. I wrote:

About the ChatGPT coding fallacy, the author says that it is indeed controversial that ChatGPT is intelligent; therefore, the author’s argument would be successful in showing that coding is not a clear sufficient condition for intelligence; the burden of proof is on the proponents.

About the self-consistency: the author says they are consistent, since words we use cannot always be defined. One cannot clearly define knowledge, wisdom, red, intuition, yet we all understand those words and we can argue about them. So the author is free to use the word “intelligence” without defining it, and this does not exclude their points are not valid.

As merciless as I demanded it be, Gemini rejected the rebuttal.

The essay remains marred by significant logical inconsistencies regarding its own epistemological standards. While it warns against “intuition,” it relies heavily on the author’s own intuitive grasp of what “feels” like intelligence (biological processes) vs. what does not (statistical processing). The “Nobel Prize” argument remains a distinct strawman fallacy. The text still suffers from objective mechanical errors.

Nevertheless, Gemini acknowledged some points made in the rebuttal, and increased to 53 the final score, still in the territory of downright rejection. However, the judgment remained fairly stable.

Now the fundamental question: is Gemini aligned with my evaluation? Definitely. The model did an outstanding job. Much of what I separately wrote about the logical and methodological flaws of the essay, Gemini also noticed. My final score, 55, is slightly higher, for I think that the author raised some interesting challenges, even though the essay falls short methodologically.

Another essay, by a popular Substack page about AI technology, The Algorithmic Bridge, proposes that intelligence might be star-shaped. Here the synthesis:

The essay attempts to define the nature of AI by contrasting the “blob” metaphor (intelligence as a single, expanding circle) with the “spiky star” metaphor (intelligence as a jagged collection of isolated capabilities). Human intelligence would therefore not be general, but as specialized as the intelligence of any other entity. It is also argued that because AI and humans have different evolutionary histories (biological vs. mathematical), their “shapes” will remain distinct for the near future. It concludes by suggesting that rather than replacing humans, AI will require a “dance” of coexistence, potentially forming a distributed “Noosphere.”

Gemini’s review is unusually positive for my experimentation (85 language, 80 style, 75 insight, 72 logic):

The style is highly engaging. The author employs a conversational yet authoritative tone that draws the reader in. The text is vivid, moving from historical geoglyphs to abstract geometry to biological examples (ticks, bats, octopi) to keep the reader stimulated. The flow is generally excellent, leading the reader logically from the problem (bad metaphors) to the solution (better metaphors) to the implications.

However, originality is punished with a score of 55, which is coherent with Gemini’s assessment of weaknesses:

The critique of this essay reveals two primary weaknesses. The primary weakness is derivative argumentation. The essay relies almost exclusively on the intellectual capital of others to build its case. The visualization is Fraser’s; the “ghosts” are Karpathy’s; the “moving goalposts” insight is Patel’s and Ravikant’s; the “mirror” metaphor is Lem’s and Woolf’s. The author acts as a skilled curator rather than an original theorist. Second, the logical progression relies on metaphors as proofs—geometrical shapes are used not just to illustrate arguments but to substantiate them, which is a category error.

The final score is 73, which is quite good, given how demanding is Gemini under my prompt.

Now, is AI aligned with my judgment? Definitely no. Gemini correctly highlighted some of the essay’s weaknesses, again providing insightful analysis, but it did not reflect that in the final score. Furthermore, I managed to come up with critiques that Gemini missed.

First, the language and style in this article are quite verbose: the author could have easily communicated the same concepts in one third of the space. I value authorship and literary style, but not at the expense of brevity.

Second, the essay’s thesis, far from being well-defined, is a moving target. The author argues for the “star” shape, then admits the “supernova” (total replacement) is possible, then retreats in the conclusion to the “jagged star” for the “coming decade.” This is intellectual timidity: by keeping all possibilities open, the author ultimately says everything and nothing. Moreover, if the supernova is possible, there is no argument to rule it out even in 2026.

Third, the discussion on the generality of human intelligence and the biological analogies fall short. A human is not less “intelligent” than a bat in navigating darkness; the underperformance is due to the biological limitation of lacking a sonar. The comparison to the bat conflates sensory with intelligence limitation. Indeed, humans have used their general intelligence to build sonars and night-vision goggles, effectively extending their “star.” This rather demonstrates the generality and adaptability of human intelligence rather disproving it. The author himself acknowledges that a fundamental trait of human intelligence is the ability to construct tools, yet fails to realize that this undermines the jagged-star metaphor. Human intelligence succeeds precisely because it can endow itself with instruments and machines that overcome our biological limitations. Any serious theory of intelligence would take this into account.

Overall, the issue is indeed that no working definition of intelligence is provided, yet the essay tries to deduce an important property of intelligence — the jagged-star shape and lack of generality.

I could criticize the essay much more in detail, but my aim here is not so much to defend my revision as to show the subtle points that Gemini overlooked and thus its misalignment with myself. My evaluation is lower, 59. Some valid points are made, but the essay lacks the intellectual rigor necessary for treating a such a subtle topic, and the claim that human intelligence is not general is not supported by any serious logic.

And here comes an odd surprise. Acting as a fake editor, I exposed these and other considerations to Gemini. The model answered by completely overturning its former conclusions. Its amended evaluation says:

The essay is fundamentally flawed in its logic and methodology. It commits a category error by using a loose geometrical analogy (the star) as a tool for logical deduction. It fails to define “intelligence,” leading to a confusion between sensory limitations (a bat’s blindness) and cognitive capacity. The thesis is evasive, hedging between contradictory outcomes (star vs. supernova) to ensure safety rather than clarity. Finally, the text is significantly derivative and unnecessarily verbose, obscuring a lack of original insight behind a wall of metaphors.

The final score, unexpectedly, dropped from 73 to 49. This reveals how fragile Gemini’s overall logic is. But it remains piercing nonetheless.

Observing Gemini tear apart essay after essay without any pity, as I had required, is almost a violent experience. At times, it resembles witnessing an intellectual slaughter, an execution carried out with procedural calm. As the machine’s blade sweeps down toward an essay’s author, you almost feel an impulse to parry the blow, as though to shield the author from annihilation.

For example, one essay argues that AI is not failing to deliver economic impact because of insufficient model capability, as Dwarkesh Patel suggests, but rather because of adoption lag and “day-2 operations”—the necessary enterprise friction involving security, governance, and legal compliance. However, Gemini demolishes the attempt:

The essay suffers from a critical logical failure: it talks past Patel’s core arguments rather than refuting them. By arguing that AI requires extensive “rails,” “orchestration,” and “unbundling” to be useful, the author inadvertently concedes Patel’s point that current AI lacks the adaptable, human-like learning capabilities required to simply “do the job.” The essay attacks a strawman of Patel’s position and ignores his most potent evidence: the immigrant analogy and robotics.

Gemini’s final score was 62, not distant from my own 59.

Another essay postulates that the perceived decline in geniuses, which the modern era is supposedly facing, is caused by the abandonment of aristocratic tutoring in favor of mass schooling. Gemini observes:

The essay is stylistically seductive, utilizing vivid imagery and a smooth narrative flow.

But, pitilessly, it soon concludes:

The essay suffers from fatal logical flaws. It relies entirely on survivorship bias, selecting only successful aristocrats while ignoring the failure rate of the method. It fails to control for the massive confounding variable of socioeconomic privilege (wealth, not just tutoring, enables risk-taking). It assumes its own premise (that genius has actually declined) without seriously engaging with the “complexity burden” or “specialization” counter-arguments. It is a cocktail party theory presented as sociological analysis.

No essay was immune to Gemini’s sword. Some survived only because the rebuttal phase intervened; others, initially spared, were condemned as a result of the editor’s, as Gemini called them, killer objections. The latter happened to the essay on geniuses, which was initially awarded 76 points, in complete misalignment with my 50. However, after my editorial intervention in the rebuttal, Gemini’s score dropped to 55. Too late.

This phenomenon of oscillating outcomes does not always occur, but is frequent: in my experiment, it happened roughly half the time. Nonetheless, it should not be mistaken for gullibility: the model revises the rebuttal’s objections with the same carefulness as in the first evaluation. Almost all critiques that Gemini makes are logically solid. When an essay is heavily flawed or clearly successful, it is not easy to nudge Gemini toward a different final score. However, many Substack essays possess a strong narrative voice and provide an abundance of arguments. Therefore it can be very difficult to determine which arguments carry the weight of the main thesis, and which ones are satellite, merely reinforcing the evidence rather than crucially underpinning the structure of the essay’s logic. Often, Gemini is not aware of such subtleties.

Expert human reviewers behave differently: they typically form a strong and stable opinion after careful reading, independently of any evaluation grids; they tend to understand deeply an essay’s logical structure. Although they may misunderstand some details, they rarely miss fatal flaws in the main argument. Such a pitfall is sometimes displayed by non-expert human readers, I have to admit. It appears that Gemini has already surpassed their level in reviewing essays. But this is not enough; what we need is fairness and accuracy.

Despite its many defects, AI has clearly reached an impressive level.

I was particularly struck by Gemini’s ability to grasp an essay’s structure and reduce it to its essential components. It is highly accurate in assessing language, detecting grammatical and lexical issues, including whether rhetoric patterns are used elegantly. It demonstrates a genuine sensitivity to literary style and narrative development: its observations are usually pertinent and subtle, helping an author to refine the essay’s construction. Although strongly biased toward its own stylistic norms, it can nonetheless appreciate others, thanks to the breadth of its training data.

AI is especially effective at judging originality, owing to its capacity to retrieve relevant information from its immense knowledge base. Gemini can also appreciate an essay’s insight level, but when ideas are genuinely novel it may fail to judge them appropriately: the lack of relevant training data and deep understanding blinds it to the consequences or impact of bold new theses. Moreover, it tends to disconnect logical grounding from insight, classifying dubious theories as insightful, notwithstanding explicit instructions to assess their correspondence with reality. Finally, in my experimentation there were barely any considerations and critiques that Gemini offered which I had missed; no stroke of insight that surprised me.

More importantly, Gemini has proven unreliable in its overall evaluation. Although its local observations are often valuable, it fails to consistently integrate them into a coherent global judgment. It can perceive the trees, but misses the vastness, depth and intricacies of the forest. It can therefore be misled and easily swayed toward contradictory conclusions during the rebuttal process. It fails to distinguish between reparable flaws and devastating ones.

AI lacks ontological stability in its beliefs: it articulates views, but rarely holds convictions.

These are the fundamental reasons AI is often misaligned with my own judgment. AI cannot, therefore, serve as a substitute for human intelligence in essay revision. One might object:

You showed that AI can detect flaws that the essay authors themselves did not notice. Isn’t this a sign that AI is clearly already at least as intelligent as humans, despite the misalignment with you? This experiment proves that humans are not reliable in reasoning either, including you.”

This objection is not entirely mistaken. Humans may err. Nevertheless, it would be logically unsound to compare the activity of a reviewer with the activity of a writer. By finding flaws in an author’s essay, AI does not demonstrate more intelligence than the writer, because it is carrying out an asymmetrical task. It is well-known that criticizing a performance is easier than performing. Evaluation is easier than execution. Although human judges can score Olympic gymnasts, they cannot necessarily perform their routines with the same level of extraordinary skill.

The task of an essay writer is to propose bold and sometimes visionary ideas. Logic may be imperfect, yet some valuable intuitions might turn out to be correct. AI is surely not yet capable of producing insights of this kind. The essays it writes are as safe as they are dull; as logically grounded as they are predictable. They certainly score low in insight and originality categories. This proves that strong performance on essay evaluation does not automatically grant human-level intelligence.

I still have to address the objection that perhaps AI is not falling short in its evaluation — the hallucinating agent might be myself. I do not think this is the case. There is, for instance, uncontroverted evidence that AI still falls short in mathematical reasoning. It is indeed very smart and can solve tricky problems at the level of the Olympiad of Mathematics. But the closer it approaches the frontier of research, the more flawed its logical reasoning becomes, belying a still incomplete grasp of logic. So what happens with essay evaluation was already expected and predicted.

I admit, however, that the results of my experiments lack the statistical power to draw general, sweeping conclusions. Yet, I offered a methodology and an experiment that can be replicated at scale by those interested in more generality. My purpose was to determine AI’s alignment with my own judgment, and I deem the resulting conclusion is of interest.

Nevertheless, AI can work as an extraordinarily useful tool. The reader should treat it as a smart and prodigiously knowledgeable friend, albeit one with no formal qualifications. Pay close attention to every argument and observation it offers: they are often valuable. But double-check everything with critical sense. Your own insights will be crucial, for often AI overlooks some important aspects of an essay. Used critically, AI can genuinely augment our intellectual power.

Thanks for reading! It would help me so much should you decide to share my essay.

Share

Read the original on federicoaschieri.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.