RSS Amplifier

Trillion Dollar Hashtag, by Antony Slumbers · Aug 18, 2026

Human Judgement Has A Timestamp

0
Sign in to vote or save

Antony Slumbers · Trillion Dollar Hashtag, by Antony Slumbers

I began with an idea for this newsletter. Working with AI, I then spent much of two days exploring two apparently more promising versions. Both turned into rabbit holes. In the end I had returned remarkably close to where I began.

Almost.

The original intuition had survived, but now I understood why. We had generated alternatives, tested different arguments, exposed weaknesses and discarded two plausible directions. The work was an extended exchange through which I discovered what the newsletter was really about.

Which, as it happens, is the subject of this newsletter.

EXECUTIVE SUMMARY

AI capabilities are jagged: a model can be exceptional at one apparently difficult task and hopeless at another that looks easy. This makes a fixed division of cognitive labour between people and machines unreliable. The effective boundary is set by the whole apparatus we build: model, harness, data, human and task. Human + AI work is a continuing process of discovering that boundary, allocating work, checking the results and changing the allocation as capabilities improve. For commercial real estate, this turns ‘human judgement’ from a permanent moat into a current operating decision. The more durable advantage belongs to organisations that can turn experience into evidence and evidence into faster learning.

TWO RABBIT HOLES

The first rabbit hole began with an excellent essay by Sangeet Paul Choudary on scarcity and strategy. If AI makes exploration dramatically cheaper, where does value move: to selecting from more possibilities, testing them against reality or framing the right question in the first place?

I thought this was the newsletter.

Several hours later it looked like a restatement of Choudary’s argument, combined with things I have written about before; constraints, moats and RIRA. Interesting. Just not interesting enough.

Rabbit hole number two was CRE’s belief in the invulnerability of human judgement. Pull the word ‘judgement’ apart and much of it turns out to be pattern recognition, evidence synthesis, prediction and accumulated exposure. All areas in which AI is advancing rapidly.

This was much better, though it still treated the boundary as a transfer of territory: AI advances, human judgement retreats.

For years I have argued that the future is less Centaur and more Cyborg: humans and machines intertwined across workflows, each changing what the other can do. The sharper question became: how much of what we assign to the human is permanently human? And how much is waiting for better models, data, tools or ways of checking the answer?

That brought me back to where I had begun: the moving boundary between us and our machines.

But with considerably more baggage.

THE JAGGED FRONTIER

AI would be much easier to understand if its capabilities advanced in a smooth line. First it masters easy tasks, then moderately difficult ones, before eventually reaching the genuinely hard stuff.

Instead the frontier is distinctly jagged.

In 2025, an advanced version of Google’s Gemini Deep Think model produced natural-language solutions to five of the six problems in the International Mathematical Olympiad within the competition time limit. The answers were officially graded at gold-medal standard.

That is an extraordinary level of mathematical reasoning.

Yet models capable of such feats can still misunderstand a straightforward instruction, miss an obvious fact in a document or confidently answer the wrong question. For a while, the internet’s favourite example was the inability of language models to count the number of r’s in ‘strawberry’. That particular example will date. The underlying absurdity persists.

The best explanation remains the ‘jagged technological frontier’, developed in research by Harvard Business School, Wharton, MIT and Boston Consulting Group. In an experiment involving 758 BCG consultants, access to GPT-4 made people faster and substantially improved the quality of their work on tasks that fell within the model’s frontier.

On a different type but equally realistic task, the results reversed. The model processed the obvious numerical evidence but missed a crucial clue in the interview material. Consultants using AI tended to trust its polished analysis and followed it towards the wrong answer. Those working without AI performed better. The research has now been published in Organization Science, and the HBS account of the findings remains one of the best introductions to AI at work.

Tasks that look equally difficult to us can sit on opposite sides of the frontier.

Worse, the frontier moves. A new model, a tool connection, access to proprietary data, a better workflow or simply a better brief can turn yesterday’s failure into today’s routine capability.

No wonder people are confused.

EXHILARATING. EXASPERATING.

The jagged frontier helps explain why people react so differently to AI.

For some, working with it is exhilarating. You can explore ideas, test alternatives, acquire capabilities you never possessed and wander productively into areas that would previously have been inaccessible. Every conversation can produce a surprise. Occasionally a very good one.

For others, exactly the same characteristics are exasperating. They ask the same question twice and receive different answers. The model is brilliant one moment and obtuse the next. It sounds equally confident whether it has understood the problem, misunderstood it or made something up.

We have spent decades learning that software should be predictable. Press this button and that happens. If it does something else, there is a bug.

AI breaks that mental model.

Our businesses already operate amid incomplete information, ambiguous signals and unknowable futures. Now we are introducing powerful tools whose own capabilities are uncertain. Before deciding whether to trust the answer, you must decide whether the AI is competent to answer that particular question.

You have to judge the judge.

Which is why working effectively with AI is becoming a genuine management skill. You need to brief it, provide context, set standards, challenge its reasoning, verify important outputs and learn where it performs well. Then a new model arrives and part of that knowledge needs recalibrating.

Some people enjoy this open-ended process. Others want the machine to work reliably and get out of their way. Both responses are perfectly understandable.

That uncertainty makes engagement essential.

We are all feeling for the stones.

CROSSING THE RIVER

Crossing the river by feeling the stones’ became the shorthand for China’s gradual and experimental approach to economic reform under Deng Xiaoping. You take a step, test your footing, learn from the result and then reach for the next stone.

It is a good description of working with AI.

There is no durable map showing which tasks a model can perform reliably, where it improves a competent professional, where the professional degrades a better machine answer, or where the combination produces something neither could achieve alone.

And your frontier will differ from mine. It depends on the model, the workflow, the quality of the data, the expertise of the person, the consequences of error and the method used to verify the result.

AI literacy increasingly consists of finding this out safely.

This requires calibrated delegation: knowing what to delegate, how much autonomy to allow, what good looks like, how the work will be checked and when it should return to a human.

Conventional software is configured. AI increasingly has to be briefed, supervised and reviewed. The operating skill resembles management more than traditional software use.

This goes far beyond ‘prompt engineering’. A good manager understands the capabilities of the person, provides the context needed to perform, sets boundaries, reviews what matters and adjusts as experience accumulates.

AI requires much the same discipline.

Except your new colleague may change significantly between Friday and Monday.

THE CYBORG’S MOVING BOUNDARY

The original jagged-frontier research identified two broad patterns of human + AI working.

The Centaurs divided the work. The human performed some parts and handed other parts to the machine. The Cyborgs moved backwards and forwards continuously, integrating AI into the way they thought and produced the answer.

My belief for some time has been that the future is predominantly Cyborg. As AI becomes woven into our research, analysis, communication and decision-making, it will become increasingly difficult to say where the human contribution ended and the machine contribution began.

But human + AI comes with no guarantee.

A 2024 meta-analysis in Nature Human Behaviour examined 106 experiments that compared humans, AI and human-AI combinations. On average, human + AI performed better than humans alone. Against the stronger of the human or AI working alone, however, the average combination underperformed.

On average matters.

But the date matters too.

The paper appeared in 2024, but the experiments it synthesised were published between January 2020 and June 2023. Its ‘decision’ category largely involved people choosing among predefined options. That is some distance from the open-ended, multi-stage judgement that CRE claims as its special preserve.

It also predates the modern reasoning-model era. OpenAI released o1 in September 2024, trained to spend more time and computing power working through a problem before answering. Later models were trained to decide when and how to search, run code, inspect files and combine tools.

The technical frontier has moved faster than the human-performance evidence. Whether these systems produce better collaboration remains an empirical question.

Harnesses move it again.

A harness is the surrounding system that turns a general model into something capable of performing a particular workflow: instructions, proprietary context, retrieval, tools, memory, task decomposition, checking loops, permissions and human approval points. As Anthropic puts it, when we evaluate an agent, we are evaluating the model and its harness together.

Give the same model the leases, comparable evidence, a cash-flow engine, a second pass that challenges its assumptions and a rule requiring human approval before release, and you have created a materially different system.

The relevant unit is the model + harness + data + human + task.

The meta-analysis is most useful for the variation it exposes. Human + AI sometimes produced genuine gains and sometimes made the result worse. When the human alone was stronger, the combination was more likely to help; when the AI was stronger, human intervention frequently damaged the answer. Stronger reasoning models could therefore make uncalibrated human involvement more dangerous.

The finding that creation tasks were more promising than decision tasks belongs to the systems tested at the time. The requirement to keep feeling for the stones persists.

And the academic benchmark of ‘use whichever is best’ can be rather moot in practice. Few companies have access to the best human in the world, and fewer can afford them. The economically relevant comparison is often your available people working with your available AI against the people and processes you have today.

More recent research complicates matters further. In The Cybernetic Teammate, a field experiment with 791 Procter & Gamble professionals, individuals using AI matched the performance of conventional two-person teams while working faster. AI also helped technical and commercial specialists produce more balanced answers, crossing professional boundaries that had previously required collaboration. The average-quality difference between AI-enabled teams and teams without AI was not statistically significant. What did move was the upper tail: AI-enabled teams were roughly three times as likely as the solo control to produce a top-decile solution.

This is powerful evidence for human + AI working. It is also evidence that AI can reproduce some of the benefits we previously attributed to human teamwork.

The combination that wins today is unlikely to stay fixed.

HUMAN JUDGEMENT HAS A TIMESTAMP

Commercial real estate has a ready answer to all of this: our moat is human judgement and relationships.

Perhaps. But ‘judgement’ is currently doing far too much work.

CRE bundles a great many different cognitive moves into that one flattering word:

  • remembering similar transactions;

  • selecting and adjusting comparables;

  • recognising patterns;

  • applying rules of thumb;

  • reconciling conflicting evidence;

  • predicting how a market or counterparty will behave;

  • generating alternatives;

  • assessing risk;

  • choosing objectives;

  • negotiating agreement;

  • taking responsibility for a decision.

These are several capabilities, and each will cross the human-AI boundary at a different time.

A more useful decomposition is:

Layer / The question being answered / Likely direction

Repeated inference / What is likely to happen? What does the evidence suggest / Increasingly AI-absorbed

Framing / What problem are we solving? Which variables and discontinuities matter / Human + AI, with a moving boundary

Preference / What should we optimise? Whose interests count? Which trade-offs are acceptable? / Institutionally human, heavily AI-mediated

Commitment / Will we allocate capital, sign the opinion and accept the consequences / Durably human because authority must sit somewhere

Some of this judgement will migrate into the model. Much of it may migrate into the harness around it. A valuer’s comparability rules, an investment committee’s mandate, downside thresholds and approval rights can become retrieval rules, tests, permissions and escalation gates.

Human contribution then moves upstream: designing and governing the decision system, deciding where autonomy ends and accepting responsibility for the commitments it makes.

The sharpest distinction is between prediction, preference and commitment. Framing shapes all three.

AI can increasingly tell us what is likely to happen, generate alternatives, test assumptions and identify evidence that contradicts a recommendation. Someone still has to determine what the organisation wants, which trade-offs it accepts and whether it is prepared to bear the consequences.

Even here, continued human involvement says little about human intellectual superiority. Institutions require named people to authorise acquisitions, sign valuations, approve safety decisions and carry fiduciary responsibility. That is a choice about where society wants authority and liability to sit.

The human may remain because someone has to hold the pen.

THE UNAUDITED ADVANTAGE

CRE’s diffuse feedback loops may have protected mediocre judgement from scrutiny.

A property is bought. Its performance takes years to emerge. The market moves, the business plan changes, the asset manager leaves and the economic cycle overwhelms the original assumptions. Success is attributed to judgement. Failure is blamed on the market.

Often nobody can demonstrate whether the original judgement was well calibrated.

‘Human judgement’ can therefore describe a decision process whose quality has never been measured because the feedback is slow and hopelessly confounded. That is an unaudited incumbent advantage. Before calling it a moat, we should audit it.

Consider three areas.

Valuation

Comparable selection, adjustment analysis, lease-event modelling, scenario weighting and consistency checking are all exposed. The valuer’s role may concentrate around choosing the appropriate basis, recognising genuine structural discontinuities, communicating uncertainty and taking professional responsibility for the opinion.

For how long each of those remains primarily human is an empirical question.

Investment committees

Much of an investment committee’s supposed secret sauce consists of interrogating assumptions, finding inconsistencies, comparing an opportunity with previous cases, constructing downside scenarios, testing the recommendation against the mandate and identifying missing diligence.

AI can already participate in each of these, although reliability varies by task and setup. The committee’s role may become narrower and more consequential: determining the institution’s appetite for an irreversible exposure and legitimising the commitment.

Leasing and relationships

Remembering histories, assessing negotiation positions, anticipating objections, preparing communications and selecting incentives all look increasingly addressable by AI.

The word ‘relationship’ needs decomposing as well:

  1. Discovery: knowing who owns something and who makes the decision.

  2. Privileged access and context: hearing early, being taken seriously and receiving candid information about the history, personalities, intentions and unwritten context.

  3. Credible commitment: being trusted to honour promises, exercise discretion fairly, resolve problems and remain accountable over time.

AI and platforms attack discovery aggressively and make relationship intelligence easier to preserve and share. But knowing who matters is different from being someone they trust, call early or speak candidly to. AI can improve how you use that access through better preparation and organisational memory; it cannot manufacture another person’s willingness to disclose, reciprocate or rely on you.

Credible commitment looks more durable. Buildings, leases and developments require prolonged cooperation because no contract can specify every future contingency. In most current arrangements, people still need to trust that someone will behave reasonably when things go wrong.

Some trust may also migrate from ‘I know this person’ towards ‘I trust this system’s evidence, controls and audit trail’. The part of relationship capital sustained by market opacity is exposed. Permission, reciprocity and earned confidence are different.

My belief that #HumanIsTheNewLuxury survives this argument. Luxury value has always been separable from technical performance. A mechanical watch keeps worse time than a cheap quartz one; its value lies elsewhere. Human attention, empathy, recognition and meaning can become more valuable even as machines outperform us cognitively.

Human value and human cognitive superiority are different propositions.

RIRA HAS A TIMESTAMP

This matters directly for my RIRA framework and the CRE Automation Matrix.

Quadrant D, hard-to-verify cognition, is currently human-led with AI acting as challenger. But it must never become a protected human reservation.

Hard to verify today does not mean uniquely human forever.

Tasks can migrate as organisations improve their data, record outcomes, build simulations and shorten feedback loops. Models may also improve at judgement without the underlying task ever becoming objectively verifiable, simply by learning across more imperfect cases than any individual could encounter.

Every classification in the Automation Matrix therefore needs a date at the top.

It also needs to record the system being tested: model, harness, data, human and task. Change any one of them and the classification may change with it.

And the RIRA false-negative challenge should become tougher. For every task claimed to require human judgement, ask:

  1. What precisely is the human contributing?

  2. What evidence shows that humans perform it consistently well?

  3. Which future capability, data source or verification method would change the classification?

‘I know it when I see it’ should not pass.

RIRA itself is a way of feeling for the stones. Release the inherited constraints. Imagine what becomes possible. Redesign the work around the best available combination of human and machine. Actualise it, observe what happens and feed the result into the next redesign.

The map is provisional. The learning is cumulative.

THE LEARNING ARCHITECTURE

Choudary’s three possible destinations for value are judgement, experimentation and framing. His resolution is that none provides an enduring advantage on its own. The scarce capability is the architecture connecting them.

An AI-native organisation needs to:

  1. frame the right search space;

  2. generate many possibilities within it;

  3. eliminate weak options cheaply;

  4. test promising ones against reality;

  5. identify where irreversible commitment is justified;

  6. record the result and improve the next round of search.

For CRE, this means capturing why decisions were made, connecting them to eventual outcomes, measuring calibration, preserving organisational memory and continually reallocating work as the frontier moves.

The harness is where this learning architecture becomes operational. Evidence becomes retrieval, accumulated experience becomes context, outcomes become evaluations, and the organisation’s thresholds and approval rights become checks within the workflow. Each round can improve the next one.

My hypothesis is that the more defensible advantage lies in the organisation’s ability to turn experience into evidence, evidence into better decisions and decisions into faster learning, while retaining legitimacy when real people and irreversible capital are affected.

That is a much more demanding proposition than ‘our people have judgement’.

It is also much more defensible.

BACK WHERE I STARTED

Which brings me back to this newsletter.

It started in one place, explored two alternatives and returned close to its origin. The process produced more than words: it turned an intuition into a position I can defend.

That is Cyborg work.

No one can give you a permanent map of the human-AI frontier. But you can build the skills and the organisation required to keep finding it.

We are all feeling for the stones.

The difference will be how quickly we learn from each step.

And that leads to next week’s question: what does a commercial real estate company designed to learn this way actually look like?

No posts

Read the original on spaceasaservice.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.