RSS Amplifier

Tyanny of the focus group of one... · Jul 1, 2026

Parrots, Perceptions and Human Factors in LLM and Multi Modal Foundation Models

0
Sign in to vote or save

Ian Nock · Tyanny of the focus group of one...

Back on June 18th at the SCTE Presents: Changing Face of Streaming 1 day conference, I said something apparently controversial. I had taken part in the first session as a panellist dealing with Streaming Protocols and Security, and in the second session as the moderator of a more wide ranging discussion about Edge Compute and AI.

In the end of day Q&A for the entire event, I said that LLM / multi-modal models or the looser term AI, don’t actually produce correct results. They don’t because they don’t know what correct actually is. They don’t know anything. They just produce the most likely result from the input. You can watch the section where I made my controversial comment below:

It was good discussion, and the discussion continued beyond the Q&A into the post event drinks.

Everything was fine, and then someone said ‘I was controversial’ because of what I said. I was taken aback.

The reason I was taken aback was because the operation of transformer based large language models are well known, and it is well known what they produce. For those who understand how they work, saying that it does not produce ‘correct’ output is not controversial. It is well understood. All the effort is in ensuring that you use the model correctly and validate what is produced by constraining how you use it.

The problem is in human perception. We are creatures that have evolved to spot patterns and to assign pattern recognition as such that we are are actually able to be fooled into seeing patterns where there is none, or in seeing patterns that are not what they seem to be. In other words, if it looks like it is thinking, it must be thinking. If it looks a correct answer, then it must be a correct answer (based on how confident it looks).

Those outputs are wondrous, and brilliant in helping us review lots of information very very quickly. They can be used in very focused ways, or in more general ways. Most of the time the most likely answer is the correct answer.

The problem is that they are not correct answers, or even creative results if you are getting it to produce an image or even a video. They look creative, but they are the result of a set of statistical matrix calculations of a multi layered and designed/trained model to produce an output from the input in useful ways.

Back in the early days of Google, I had a saying that ‘Google means never being more than a few seconds from having an answer’, and that is a fantastic power. The challenge was in determining if it was a correct answer, and that required knowledge, experience and references. The same is true today with LLMs, they produce the most likely output based on the input after being processed by a model that has been trained on a huge amount of data, that has enshrined a representation of that data in a way that it can be used to be transformative to the input data to an output, and if you are lucky the model has a way of doing a Google Search to find ‘live data’ so that your output can be up to date.

But it is not necessarily correct, and this is the danger.

We still need to qualify the results and process it ourselves with expertise if we are to truly rely on the information. Just like a Google search from almost 20 years ago.

This is the danger that was well described 5 years ago by a seminal paper called ‘On the Dangers of Stochastic Parrots: Can Language Models Be Too Big’. This can be downloaded at https://dl.acm.org/doi/10.1145/3442188.3445922.

I really do propose that you read in its full. It describes the problem perfectly, but does not discount the power of LLMs, and transformer models as to what they can actually do. It is fully behind their careful use because they have no actual understanding, only a simulation of one. In fact, I also suggest you read an update after the 5 year point from Emily Bender, on what has happened since. It is very a useful read, click on the image to read.

This why the best use of AI Transformer models is with the human in the loop. As an accelerant to getting things done, but not a replacement. They are very powerful tools, but the true power is when used by expert humans who account for the fact that they produce likely answers and not correct answers. We are the ones who make the answers produced correct, and need to ensure that we have a breakout when the response is outside of an accepted range, through defaults or non-AI based techniques. The expertise is in designing those breakouts well.

In the SCTE conference session, I did actually suggest a bit of homework to everyone. That homework is to ask your favourite transformer model a question you know the answer to. And when it responds, tell it that it is wrong. Not just once. Do it all the time and see what it does. Eventually it will get in a loop or start providing low confidence answers that are not correct. This signifies the other challenge with the transformer models. They are eager to please, to personify it a bit.

In other words, most models are not designed to say ‘Sorry I don’t know’. The good ones may actually respond in that way, but that is a sign of a human programmer having put some defensive programming in place to provide a default answer when the confidence is low or if you query it regarding something that it has been programmed not to answer. One pointer to look out for with respect to that, is to look into jailbreaking of models - which is about working around the protections put in place by their designers to not allow you to do certain things - particularly producing copyrighted imagery or asking how to make bombs. I suggest you don’t do that last point, because you might attract untoward attention from the authorities.

Don’t take what I am saying to be some massive negativity about the use of Transformer models. Just as the AI researchers say in the ‘Stochastic Parrot’ paper, LLMs and Foundation models are very effective when used well. Just that they need to be used well, and we need to be open about their limitations in their use, particularly with respect to correctness. In fact, there are a great many applications out there where processing inputs to produce a most likely output, can be a fantastic solution. The issue is where correctness has to be assured, and how do you do that?

Transformer models are fantastic in so many ways. Just don’t fool yourself into thinking that they provide correct answers rather than the most likely.

No posts

Read the original on iannock.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.