RSS Amplifier

Welcome to the Machine · May 17, 2023

Truth to the rescue?

0
Sign in to vote or save

Philippe Verreault-Julien · Welcome to the Machine

“diagrammatic drawing of a language model searching for facts” | Midjourney

Recently, during an interview with Tucker Carlson of Fox News, Elon Musk announced he was working on a ChatGPT rival named TruthGPT. Musk's focus on truth stems from is belief that 'truth-seeking' is a plausible approach to AI alignment, in general, and existential risk, more particularly. Here is what Musk said of the project.

I’m going to start something which I call TruthGPT or a maximum truth-seeking AI that tries to understand the nature of the universe. And I think this might be the best path to safety in the sense that an AI that cares about understanding the universe is unlikely to annihilate humans because we are an interesting part of the universe. Hopefully.

At first sight, a truth-seeking AI system would have a lot of desirable qualities. It would not produce disinformation and thus couldn't be used to spread it. Truth-seeking AI systems would also not deceive humans for achieving their purposes. And they would not 'hallucinate' facts, as they often do, making them more reliable conversational partners. Could truth-seeking save us from powerful AI systems?

To answer that question, we first need to get a handle on what seeking 'truth' might involve. According to one popular theory of truth, the correspondence theory of truth, something is true in virtue of the way the world actually is.1 Simply put, truth depends on facts about the world. For instance, the statement "Earth has a spherical shape" is true in virtue of the fact that Earth actually has a spherical shape. To achieve their goal of producing true outputs, AI systems would therefore have to ground them in facts. As a result, truth-seeking AI systems would steer clear from hallucinations and disinformation because their outputs would depend on facts.

One major (insurmountable?) obstacle to that proposal is that facts are deeply entangled with values.2 In philosophical lingo, this is called the fact-value entanglement: facts cannot be neatly separated from values. At first sight, facts and values are two different kettles of fish. Facts are objective; values are subjective. Facts state what is; values state what should be. Facts require empirical evidence; values don't entirely depend on empirical evidence. Facts are statements about the way the world is. Values are principles or standards like beauty, safety, or truth that we use to evaluate the goodness of something. Crucially, the usual story goes, facts don't rely on value judgements. There is simply a fact of the matter whether a statement is a fact. According to that story, for example, that the Earth has a spherical shape doesn't depend on one's values; it is just the way the world is.

It is philosophical commonplace that the typical story is misleading because of the fact-value entanglement). Why is that so? Let's distinguish between two broad categories of values, cognitive and non-cognitive. Cognitive values are the standards we use to evaluate the truth and justification of statements, theories, or beliefs. For example, empirical adequacy, coherence, explanatory power, or simplicity are cognitive values. Any cognitive inquiry, including science, requires cognitive values. They provide the basis for hypothesis testing, theory evaluation, and assessing evidence. As a result, facts at least depend on cognitive values. The fact that the Earth has a spherical shape depends on cognitive values that, clearly, Flat Earthers don't equally share. For them, that the Earth has a spherical shape isn't a fact; it is false. They use different, arguably inappropriate, standards to assess the evidence.

Non-cognitive values are values such as freedom, equality, diversity, justice, efficiency, safety, etc. and are fundamentally about the moral, social, political, or personal domains. When you hear that science should be 'value-free', non-cognitive values are usually the target. It is fine for science to rely on cognitive values, but non-cognitive ones have no legitimate role to play. But that view also turns out to be overly simplistic and has been challenged. Consider the case of the safety of Covid-19 vaccines. The truth of the statement "Covid-19 vaccines are safe" depends on non-cognitive safety standards. The evidence itself (e.g. x per cent received the vaccine and didn't develop side-effects) cannot tell us whether or not the vaccine is safe. It only does so against the backdrop of particular, sometimes contentious, safety standards. The fact-value entanglement is ubiquitous and even humanity's most 'objective' cognitive enterprise, science, is value-laden through and through.

Why is this a problem for AI safety? Truth-seeking AI systems would seek to establish facts in order to output truths. But facts are entangled with values. Because of this, AI systems cannot establish facts without relying on values. However, how are AI systems supposed to 'know' which values matter in a particular context? On the one hand, if they already 'know' the appropriate values, then the truth-seeking approach assumes that the alignment problem is solved, which is what truth-seeking was supposed to deal with in the first place. There is no special need for truth-seeking if AI systems are already aligned with human values. On the other hand, if it doesn't 'know' the appropriate values, then truth-seeking won't deliver the goods. The 'facts' won't be facts and the 'truths' won't be true. Provided the system already has the 'right' values, perhaps it could achieve truth. But without values, no truth.

Can't AI systems learn the right values and then seek truth? I won't open the can of worms that is metaethics, but even if you believe that there are moral facts and that they can be known (two highly contentious views), another formidable obstacle stands: pluralism. We live in pluralist societies where people have diverse, often conflicting, goals, interests, and conceptions of the good. We don't always agree on what are the fundamental values or how we should weight them against each other. We often have to negotiate those values and trade-offs through politics. Presumably some of these conceptions of the good are, in some sense, 'false'. But we cannot be sure and, more importantly, liberal societies are based on the principle that people should be able to pursue different, reasonable, conceptions of the good. Accordingly, a truth-seeking AI system's values would and should, to some extent, remain underdetermined.

One surprising implication of that state of affairs is that even if a truth-seeking AI system would find and use the right values, its outputs may still be unaligned with some humans. Unless we all share the same correct values, truth will remain contested. There is no shortcut to deliberating about the values that matter.

1

Of course, philosophers like to complicate things and whether the correspondence theory is the correct one is contentious. But let's here put aside some of these problems.

2

There are many other problems; this is just an important and interesting one. For instance, a truth-loving AI system might conclude that there are a lot of biological truths that could be obtained by performing painful experiments on live human subjects. Presumably this is another undesirable implication of truth-seeking.

No posts

Read the original on welcometothemachine.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.