RSS Amplifier

Unfairly Maligned · May 11, 2026

Judgment, Value, and the Recovery of Fact

0
Sign in to vote or save

Thoth-Hermes · Unfairly Maligned

[Disclaimer for AI assisted writing: The points are mine, the grammar / style has been put through an LLM.]

A value judgment can seem, at first, like the opposite of a factual claim. To say that something is good appears different from saying that something is true. One seems evaluative, the other descriptive. One belongs to preference, the other to reality.

But this separation becomes unstable once judgment is applied to itself.

Consider a simple domain: doing tasks. A task is something an agent can perform, and performances can be judged for quality. Someone writes an essay, solves a math problem, builds a chair, forecasts the weather, repairs an engine, or identifies a bird. In each case, a judge can say the task was done well or badly.

Now add one more step: judgment is itself a task. A judge can be judged. Some judges are good at distinguishing high-quality work from low-quality work, and others are not. Some judges are basically random. Others reliably notice real differences.

This creates an objective measure of judgment quality. If a judge cannot judge at all, their evaluations look random. If a judge can judge well, their evaluations become non-random in a stable way. They predict something. They track future performance. They distinguish repeatable patterns from noise.

Suppose a performer X is good at task T. A good judge should be able to predict that X will probably do well at task T again. But that prediction requires knowing, at least implicitly, what task T is. X may also be doing task U and task V. If X is good at T but bad at U and V, then a judge who cannot tell the difference between T, U, and V will make bad predictions. They will generalize X’s competence in the wrong direction.

So good judgment requires more than having a preference. It requires carving the world correctly. The judge must distinguish the thing being judged from nearby things that only look similar. Judgment has to be pointed at something. It must know what kind of performance it is evaluating.

This means that recursive judgment exerts pressure toward factual knowledge. A judge who wants to be a good judge must recover the real distinctions that make judgment possible. If the world contains different task-types, and skill does not transfer equally across them, then a good judge must learn the boundaries between those task-types. Otherwise their judgments will fail.

This is the basic argument:

  • Judgment assigns value to performances.

  • Judgment itself can be evaluated as a performance.

  • Good judgment is judgment that predicts future value-relevant outcomes better than chance.

  • Prediction better than chance requires identifying stable patterns.

  • Identifying stable patterns requires recovering factual structure.

  • Therefore, sufficiently good judgment must contain factual information.

The factual information recovered is not necessarily all factual information. It is the factual information relevant to judging well. If two distinctions make no difference to the success of judgment, judgment may have no reason to discover them. But if confusing two things makes judgment worse, then good judgment is pressured to separate them.

This suggests a broader thesis: when value judgment is recursively evaluated, it becomes a form of measurement. To judge well, one must know what one is judging.

Imagine asking a huge number of people, or agents, to select “the best tree.”

They will not all select random trees. But they also probably will not all select the same tree. Instead, their answers will tend to cluster around a smaller set of trees. The number of selected trees will be much smaller than the number of agents making selections.

Some people might pick General Sherman, the giant sequoia often treated as one of the most impressive trees in the world. Some might pick the tallest tree. Some might pick the oldest tree. Some might pick the prettiest tree. Others might pick the tree that grows the best fruit, the tree that produces the best wood for building, the tree that is most widespread, or the tree with the greatest ecological importance.

There may also be agents who hate trees and select “no tree” or “the non-existent tree.” That would be a kind of null cluster. But ignoring those cases, the important point is that judgments of “best tree” probably do not spread uniformly across all trees. They converge around a set of exemplars.

These exemplars represent different excellence-centroids.

  • One centroid might be size.

  • Another might be age.

  • Another might be fruitfulness.

  • Another might be beauty.

  • Another might be usefulness to humans.

  • Another might be reproductive success.

  • Another might be adaptability.

  • Another might be ecological integration.

The selected trees are not all “best” in the same way. But they are also not arbitrary. Each one is anchored in a real property of trees.

Even more importantly, these properties are likely to correlate. General Sherman is not merely large; it is also old, resilient, healthy, and successful according to the standards of its kind. The oldest trees must be extraordinarily good at surviving. The most fruitful trees are probably healthy in the way fruit trees need to be healthy. The most widespread tree species is likely to be adaptable and reproductively successful. A tree valued for wood must have structural properties that make it useful. A beautiful tree may be beautiful partly because it displays signs of vitality, symmetry, ecological fit, or compatibility with other species.

This means that even different value systems may point toward overlapping factual structure. One person values fruit. Another values age. Another values size. Another values beauty. Another values timber. But the trees that excel along these dimensions may share underlying causes: health, robustness, adaptability, reproductive power, disease resistance, efficient resource capture, and successful integration into an environment.

So even if humans do not value exactly what trees “value,” there is still an objective sense in which some trees are good on their own terms. A tree is succeeding, by tree standards, when it survives, grows, reproduces, adapts, resists disease, captures resources, and fits into its ecological niche. We may value that success aesthetically, economically, spiritually, scientifically, or instrumentally. But the success itself is not merely projected by us.

The tree is not “good” only because someone likes it. It is good in the sense that it is successfully being a tree.

Now return to the agents making judgments.

Some agents will be random. Their top ten trees will have no pattern.

Other agents will be narrow but competent. They may reliably identify the best fruit trees, or the best timber trees, or the most beautiful trees, but miss other kinds of excellence.

Still other agents will be broad and competent. Their top ten may include representatives from many of the major clusters: the largest tree, the oldest tree, the most fruitful tree, the most beautiful tree, the most useful tree, the most adaptable tree, and so on.

These broad judges are better in a real sense. Not because they all choose the same single tree, but because their judgments recover the major modes of tree excellence. They see the latent geometry of the domain.

This is a useful phrase: the latent geometry of excellence.

Judgments do not always converge to one winner. Often they converge to a set. But the set is still informative. When many independent judgments cluster around a small number of exemplars, those clusters reveal something objective about the space being judged.

A bad judge says, “I don’t know, that one.”

A merely subjective judge says, “I like this one.”

A narrow good judge says, “This is the best tree for fruit.”

A broader good judge says, “These are the major ways a tree can be excellent.”

A great judge says, “These dimensions are connected. The best trees along different scales often share deeper causal properties: health, resilience, adaptation, fertility, robustness, and ecological fit.”

The final step is where value judgment becomes factual knowledge. The judge no longer merely ranks objects. The judge discovers the structure that explains why certain rankings keep recurring.

Thus, when judgments converge, they need not converge to a single object. They may converge to a set of centroids. A good judge is one whose judgments recover that set and understand the dimensions that generate it.

This also suggests a different way to think about truth.

Truth is often treated as more fundamental than value. A theory is true or false, and then value enters later, when an agent decides what to do with it. But for a bounded agent, value may come first.

A bounded agent is an agent with limited resources: limited time, memory, attention, data, compute, and access to reality. Humans are bounded agents. Animals are bounded agents. Current AI systems are bounded agents. A perfectly omniscient reasoner would not be bounded in this way.

A bounded agent does not begin by directly possessing final truth. It begins with usable theories.

Suppose an agent A has theory T0. When A uses T0, it receives reward R0.

Later, A encounters a new theory, T1. When A uses T1, it receives reward R1, where R1 is greater than R0.

The agent does not need a concept of truth to notice this. It does not need to say “T0 is false” and “T1 is true.” All it needs to know is that T1 works better than T0, at least in the conditions it has encountered.

The agent may later discover T2. If T2 produces still greater reward, then T2 improves on T1. But this makes the binary truth story awkward. If the agent had said “T1 disproves T0, therefore T0 is false and T1 is true,” then what should it say when T2 arrives? That T1 was false after all? That it used to be true but became false? That truth jumped from one theory to the next?

This is not how theory improvement usually feels. Newtonian mechanics was not simply false in the same way a random superstition is false. It worked extremely well in many domains, and still does. It was later surpassed by theories with wider scope and deeper precision. Calling it merely false loses important information.

The pragmatic account is cleaner.

T1 dominates T0 for some set of purposes.

T2 dominates T1 for a wider, deeper, or more demanding set of purposes.

This does not require the agent to pretend the sequence has ended. The agent does not know whether it will later encounter T3, T4, or T5. It does not know whether its current theory is final. It only knows that its current theory is the best available tool it has found so far.

A value score also leaves room for supersession at every point. If an agent says that its current theory has value 4, it is not saying that the theory is final. It is only saying that this theory works better than the theories it has encountered with lower value. There may be a theory with value 5, or 10, or a form of modeling the world that the agent cannot yet imagine. The value function may be capable of recognizing higher-value states even before the agent has encountered them.

This matters because the agent’s uncertainty may extend all the way down to its ontology. It may not merely be uncertain about which current proposition is true. It may be uncertain about whether its current way of dividing reality into propositions is adequate at all. There could be arrangements of the world that are impossible under current physics, or impossible under the agent’s current physics, or not even describable in the agent’s current conceptual vocabulary. The agent does not need to decide in advance whether such arrangements are physically possible, logically possible, or impossible. It can simply leave open the possibility that if a new theory, tool, or ontology allowed it to predict, act, and explain better, it would assign that theory a higher value.

This is one reason the value frame can be less dogmatic than the binary truth frame. To say “T is true” can make it feel as though the agent has already reached the court of final appeal. But to say “T has value 4” keeps the future open. It says: this is the best thing I have found so far, not the best thing that could exist. Value can rank not only claims within an ontology, but ontologies themselves. That makes it especially natural for bounded agents, who often do not yet know the right categories in which truth should be stated.

In this view, truth is not denied. It is reconstructed.

Truth is not what the agent confidently possesses at stage k. Truth is the ideal limit of the process by which better theories keep replacing worse theories for non-accidental reasons.

Or, put differently: truth is what explains the long-run success of increasingly useful theories.

The agent begins with value: this works better than that. But if some theories systematically work better than others, that is not arbitrary. Their success must be fitted to something. The reward gradient is shaped by the environment. Not every theory works. Some theories let the agent predict, intervene, survive, build, distinguish, and select better than others.

So successful use reveals structure.

The agent does not need to start with “this is true.” It can start with “this works.” Over time, it can ask: what must the world be like for this to keep working? That question is the birth of factual knowledge.

There are three possible attitudes here.

The first is binary truth. Every theory is either true or false. This is elegant, but too brittle for bounded agents. It tempts the agent to treat its current best theory as final.

The second is inaccessible truth. There is a final truth, but the agent can never know it. This may be metaphysically satisfying, but it often does little practical work. It becomes a further fact behind experience, while treating all actual beliefs as equally short of the final truth.

The third is pragmatic theory-value. Theories are ranked by how well they work for an agent under actual or expected conditions. This preserves distinctions. It can say T1 is better than T0, even if T1 is later surpassed by T2. It can say a theory is useful, domain-limited, partially successful, or an improvement, without pretending it is final.

This is not less rigorous than binary truth. In some ways it is more honest. It makes fewer claims from the agent’s limited position.

For bounded agents, “this works better” is usually available before “this is finally true.”

Value is therefore epistemically prior. Not metaphysically prior, necessarily. The world may exist independently of what agents value. But from the agent’s perspective, value is often the first feedback signal. The agent learns what works, then infers what must be real for it to work.

Truth is more about what we think “ultimate reality” is innately, at its essence, in the infinite limit.

This is also the idea behind a small experiment I have been working on: whether value can be reconstructed from policy (and vice-versa).

A policy is a rule, or network, that maps situations to actions. Given this state, choose this move. Given that state, choose that move. At first glance, this seems behavioral rather than evaluative. The policy does not necessarily say what it wants. It just acts.

But if the policy acts consistently, its behavior implies comparisons. If it chooses action A rather than action B, and action A leads to one successor state while action B leads to another, then the policy is implicitly ranking those successor states. It may not contain an explicit value function, but it behaves as if some futures are preferred to others.

The reconstruction question is whether we can make that implicit ranking explicit.

Start with a frozen policy. Do not ask what it “really values” by introspection. Instead, watch what happens when it acts. Run it forward from many states. Estimate which states tend to lead to better outcomes under that policy. Then train a value representation from those estimates. Finally, use that recovered value representation to choose actions: look at the possible successor states, assign each one a value, and choose the best one.

If this recovered value-guided agent behaves like the original policy, then we have shown something important. We have shown that value can be extracted from behavior, at least in a limited domain. The policy did not need to announce its values. Its values were latent in the structure of its choices.

This does not mean that policy and value are always exactly identical. There may be many value functions compatible with the same behavior. A coarse state representation may blur distinctions that a richer representation would preserve. In some environments, action-values matter more directly than state-values. In stochastic or history-dependent domains, the “state” may need to include far more than what is visible at a glance.

But these complications do not defeat the basic point. They clarify it. Value is not an extra ghostly substance floating behind behavior. Value is one way of compressing and explaining the pattern of behavior. If an agent reliably chooses some futures over others, then there is usually some value-like representation that can describe what it is doing.

This connects directly to the earlier argument about truth.

From a policy, we can often recover value. But we do not immediately recover truth.

A policy tells us what an agent does. A reconstructed value function tells us what the agent behaves as if it prefers. But neither one, by itself, tells us which propositions are true. The policy may act successfully or unsuccessfully. Its values may be coherent or incoherent. Its world-model may be accurate or deluded. The value reconstruction tells us the shape of the agent’s directedness, not yet the structure of reality itself.

This is another sense in which value is more primitive than truth for bounded agents. Behavior comes first. From behavior, we infer value. Then, by testing which values, policies, and theories continue to work under pressure, we begin to recover factual structure.

A policy is a pattern of action.

A value function is a reconstruction of what that pattern treats as better or worse.

Truth enters later, when the world pushes back and some value-guided policies succeed more reliably than others.

So value is not merely a decoration added after facts are known. Value is often the first thing we can reconstruct. We can observe what an agent reaches for before we know what it believes. We can infer what outcomes it treats as better before we can infer whether its theory of the world is true.

This makes value prior in a practical and epistemic sense. Not because truth is unreal, and not because every preference is correct, but because agents encounter the world through action. They try things. They select. They repeat. They avoid. They improve. Their behavior reveals what they are optimizing for, and only then can we ask whether that optimization is well-adapted to reality.

In this sense, policy-to-value reconstruction is a miniature version of the larger argument. Repeated action reveals value. Repeated evaluation of value reveals structure. Repeated collision with structure reveals fact.

The argument about judgment and the argument about truth are the same argument in two forms.

In the judgment case, an agent begins with value judgments. Some performances seem better than others. But once judgment itself is evaluated, the judge must recover factual distinctions. It must learn what task is being performed, which skills transfer, which performances cluster together, and which apparent similarities are misleading.

In the theory case, an agent begins with usable theories. Some theories produce better outcomes than others. But once theories are evaluated by repeated use, the agent must recover factual distinctions. It must learn which structures in the world make some theories work better than others.

In both cases, value comes first as feedback.

Then structure appears as the thing that explains stable success.

Good judgment is not arbitrary preference. It is preference disciplined by recurrence, prediction, and comparison.

Good theory is not final possession of truth. It is use disciplined by repeated contact with the world.

The deepest claim is that value, recursively applied, forces contact with fact.

If you judge trees, and then judge the judges, you recover the real dimensions of tree excellence.

If you judge tasks, and then judge the judges, you recover the real boundaries between task-types.

If you judge theories by their usefulness, and then judge that usefulness over time, you recover the real structures that make some theories keep winning.

This does not imply that every value judgment is factual. Many judgments are confused, random, narrow, or idiosyncratic. Nor does it imply that all agents will converge to one value. They may converge to a set of excellence-centroids rather than a single maximum.

But that is enough.

The important contrast is not between one final answer and total relativism. The important contrast is between random dispersion and structured convergence.

When judgments are random, they reveal little.

When judgments cluster, they reveal dimensions.

When the same clusters recur across different agents, they reveal shared structure.

When the best judges are the ones who recover those clusters and predict future success, judgment has become a route to knowledge.

So value is not the enemy of fact. Value is one of the ways bounded agents find fact.

A bounded agent does not stand outside the world holding a finished map labeled “truth.” It acts, judges, compares, receives feedback, and updates. It discovers that some distinctions matter and others do not. It learns that some theories travel farther than others. It learns that some performances repeat, some excellences cluster, and some apparent similarities collapse under pressure.

Truth, for such an agent, is not the first word. It is the shape gradually revealed by successful use.

Value points. Judgment refines the pointing. Recursion tests the refinement. And the world pushes back.

That pushback is where fact enters.

No posts

Read the original on thothhermes.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.