It seems like every week I’m presented (at times against my will) with an exciting new use case for generative AI or a major update to its capabilities: circle an image to search, generate a video from mesh in minutes, test drive a new code buddy, and so on. Yet sometimes it isn’t the output that deserves our focus, but the process itself. Yes, the watch tells you the time, but how much more can we learn by taking apart the intricate system beneath the dial?
Most of us can imagine the nearly artistic way the mechanism of a classical watch is arranged, but what exactly lies beneath the dial of AI? To explore this, I will borrow a term from Douglas Hofstadter’s fantastic Gödel, Escher, Bach. We are going to “step out of the system” and observe from a higher vantage point.
This is not a discussion of AI’s outputs such as the poem, the image, or the decision. It is about the underlying machinery that produces them. By shifting our focus from what AI creates to how it learns to create, we step into a rich but less-traveled domain.
The reason this topic first intrigued me, and then kept tugging at me like a relentless squirrel, is our implementation of Multi-Agent Reinforcement Learning (RL) in several of our digital twins. Notice how we can’t take more than a few steps without pausing to clarify terms? No worries. A quick introduction to Reinforcement Learning may be the last one before we are free to push forward.
In a nutshell, Reinforcement Learning is an AI method that trains a virtual agent to discover solutions to problems. The agent is given a virtual environment that models a real-world situation, the ability to perform certain actions, one or more goals, and also elements that penalize failure. As the agent acts in this virtual environment, it progresses either towards or away from its goals. These actions may also change the virtual environment, thus altering the next actions selected by the agent. These sets of actions may be repeated many times, either by changing some of the parameters or by relying on randomness to discover new solutions.
In one of our projects we trained virtual agents in the digital twin of a concert hall to simulate reactions to fire. They began as blank slates, learning to navigate the hallways, avoiding the fire and leaving the concert hall quickly.
The first set of runs had our agents wobbling around and bumping into everything possible, if they were even able to move far enough to hit something. Over countless, but rapid, sessions they went through variations of ineptness: awkward, eerie, but mostly embarrassing, until they learned to run out of the virtual burning room with reasonable mobility and a sense of urgency.
As impressed as we were with the autonomy achieved by the agents, our goal was to add value to the concert hall through the digital twin. We therefore expanded these iterations by introducing multiple agent types, some trying to escape and others trying to put out the fire. We also altered the environment by placing fire extinguishers in different positions. What we change in the virtual must remain viable in the real world, whether that means adding fire escape signage, repositioning extinguishers, or training personnel through worst-case scenarios in the digital twin using XR technologies.
It isn’t hard to extend this to many different domains. Imagine modeling an active volcano, along with the surrounding topography and the effects of heat and lava, and letting a variety of virtual agents, whether they represent robots or humans, discover the safest way to approach this hazard.
How about keeping it simpler and closer to daily life? Let’s say you are considering opening a combination pastry shop and café. Why not create the setting in the virtual, letting agents, both customers and staff, try out various combinations of layouts, services, and more? Let your imagination run wild while the agents try to keep things grounded.
Just letting our minds wander through such examples should reveal that the more we know about the internal workings of the agents and the environment, the greater our capacity to alter both the agents and the virtual setting to expose deeper insights. This, I believe, would make us more comfortable with giving external entities the capacity to act, at least in some sense of the word, as we are learning and growing alongside them.
But I would like to draw your attention to a critical point and steer us toward the central premise of this section: in reinforcement learning you aren’t always pushing the agents to make the best move at each step. This may seem counterintuitive. Why shouldn’t we memorize some of the safer, faster routes and simply keep moving in that direction?
Because just like in chess, where a gambit is a tactical losing move that opens the way for a more powerful endgame, a succession of the locally best steps may not accumulate into the best overall path. In metaheuristics this is sometimes called a local peak. Kilimanjaro and Fuji are certainly high, but they are not Everest.
We can make agents in RL seek new and potentially better solutions by increasing their tendency toward exploration in lieu of exploitation. Finding the right balance between exploration, trying new actions, and exploitation, sticking to series of actions that are known to be beneficial, depends on the goal and the context within which the agent is operating. But the crucial lesson we gain by looking beneath the surface of this AI method is the power that comes with being able to adjust the balance between exploration and exploitation in our own lives.
Let me unpack this a bit, as it requires a deeper look at how the tension between the status quo and the search for novelty operates in humans at a subsurface level.
Although my focus in this writing is on the exploration tendency, it is not my intention to cast the exploitation element in a bad light. Keeping the status quo surely has its advantages. We are more efficient when working with the familiar, we are calmer, and we are more likely to form stable goals when interacting with what already works. It is no wonder that humans place a high priority on preserving the status quo and have a bias toward risk aversion.
Yet I feel that the exploitation mode, especially beyond childhood, becomes too dominant, keeping us from reaping the benefits of adventure and discovery far too early. This may be partly a result of our modern lifestyles, but also of over-protective patterns laid down in the subconscious through the messages we received in childhood. I would argue that the majority of adults live clearly on the side of safety, inadvertently missing out on the lifelong benefits of exploration.
It is through exploration that we keep a keen mind in old age, it is how we turn a couple of trite ingredients into a world-class soufflé, it is how we ride the flow state outlined by Mihaly Csikszentmihalyi, it is how we live and relive the hero’s journey described by Joseph Campbell, which is as old as human history. It is how we keep ourselves in the running to make the next breakthrough in science, art, or the social fabric.
In fact our bodies are happy to reward us for taking new paths even when we come back empty handed. Experiments have shown that talking to a stranger on a bus for at least ten minutes increases happiness, even if the conversation is not particularly pleasant. Studies also show that couples who share novel activities have a more positive view of their relationships. In a way our underlying wiring gives us gifts even if we do not return with the flame like Prometheus.
Unfortunately it is also our subconscious, and the various hormones and chemicals at its disposal, that may persuade us to stay dormant and refuse the call of adventure when we should accept it. Just how does this occur? Is there a switch in our minds that gives us the green light to explore and grow, or a red light that tells us to stay on the couch? In a way, yes.
One of the most basic instincts shared by humans and other complex creatures is the fight-or-flight mode. It is a way for our bodies to allocate resources so that we are ready to take a challenge head-on or to flee in a hurry. The underlying mechanism involves the sympathetic nervous system and some familiar components such as adrenaline and cortisol, the longer-term stress hormone. This mode is activated rapidly in response to immediate risk or when entering an environment with too many unknowns.
It is this mode that gives us butterflies before a public speech, the rush of adrenaline when inviting someone to dance, or the tension of wondering whether your quarterly report will net you a promotion. It is natural and elemental. But it becomes destructive when we are not able to regulate ourselves after the perceived risk has passed. Stress is the modern killer, but it is not an enemy. It is a friend that has overstayed its welcome.
Our subconscious learns to regulate itself at an early age and carries this ability into adulthood, but only if it is convinced that we grew up in a safe environment. In other words, we build a healthy pattern of engaging the fight-or-flight mode while exploring and taking risks as toddlers, and then returning to a calm state once the mini-adventure is complete. As adults we may assume this happens automatically, as long as a child is not left in the wild or raised in a war zone. But that is not the case.
Our subconscious determines the safety of our environment, in fact for the rest of our lives, based on the responsiveness of our primary caregiver (most likely our mothers). It isn’t enough to be fed and cuddled from time to time. Children seek a highly responsive interaction from an adult who knows when to step in and when to hold back, who can provide the required eye contact or a caress at the right time, telling the body that fight-or-flight is over. It is now time to reallocate those resources for growth and repair.
Numerous studies in Attachment Theory suggest that for roughly 65% of people this regulation occurs at a young age, resulting in a healthier lifelong habit of connecting with others, in essence integrating into a social structure that supports us as we face risks and venture out to explore. Some 21% learn early on that they must go it alone, at a high cost to their long-term health. And finally 14% form brittle connections to others, at times causing more harm than good.
Similar outcomes have also been observed in animals. The work by Meaney on maternal licking and grooming in rats shows that high levels of this behavior result in offspring that are more courageous in exploring their environments as adults, and therefore more likely to find food, mates, and new opportunities. Those that were neglected early in life, on the other hand, have a much harder time recovering from the anxious state of fight-or-flight and over time learn not to explore or put themselves at risk.
I’m not trying to make the case that 65% of us are natural explorers and the rest are doomed to the couch. Clearly there are other factors that go into the equation. But it is important to realize that if we wish to inject a greater element of growth, discovery, and adventure into any stage of our lives, we need to look beyond curiosity and courage and also work on the social support structures that tell our bodies to go sailing, as a safe harbor is at hand.
In fact the 80-year Harvard study, which I will cover in later sections, concludes that a feeling of deep support from our social circles is the strongest predictor of a long and healthy life. It is a result well worth revisiting, and one that connects naturally to the role of the digital twin. In the meantime, let us build for ourselves and our loved ones robust harbors so that we never cease to seek adventure, no matter where we are in life.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.