I like to picture everyone going through life with an invisible toolbox at their hips, filled with tools acquired over a lifetime and ready to be used in times of crisis or during those rarer moments when we find the time and energy to restructure our lives. These tools might take the form of a philosophical school such as Stoicism, an understanding of everyday human fallacies, or the ability to cook a killer soufflé.
As intriguing as it is to witness people, at practically any age, discover new life tools and wield them with glee, I believe this reaches another level with firms. A software startup adopting Agile methods is a transformation worth watching, especially that moment when hardcore coders start believing in the power of a wall full of post-its. Now imagine riding along as a growing company re-evaluates its core principles, reshapes its value propositions, or discovers an entirely new category of funding options. And that’s before even opening the vast warehouse of tools made available by the recent growth of AI platforms.
Do we unleash generative AI on our legacy knowledge base and see what new doors it may open? Shall we dive into reinforcement learning agents and let them run wild in search of new solutions? Or do we seek immediate impact through machine-learning tools that uncover bottlenecks and hidden patterns in our processes? It all feels a bit like a teenager inheriting an exotic car gallery.
I would recommend taking any one of these AI tools for a ride, as one of them may take your organization to thrilling new places. But a quick question before you rush off: do you have the fuel to run them? I do not mean manpower or computing resources, which are also essential, but the substance that actually powers AI systems: well-defined and accessible data.
You may have noticed that I did not simply call this fuel “data.” That distinction is vital. Over the years, I have seen how often it is misunderstood across organizational hierarchies, with quite dire consequences.
To illustrate the power of mature data, and the risks that come with relying on poorly structured data, I’ll walk through a few stages of data maturity, peppered with some personal experiences.
I’ll set our starting point as data captured on paper. This has been a staple of civilization for so long that those of us not born with tablets in our hands can still find our way to a room full of documents using only our noses. Clearly engraved in my mind are the particular shade of yellow, the dourly artistic way the edges of paper crinkle and bend, and an odd type of dust that tickles the scholarly tendencies.
Please let there be no misunderstanding: I do not miss my runs through the literal hallways of bureaucracy, in search of yet another signature or stamp to complete an application. This was long before those stacks could be endlessly replicated, safely archived, or sent across the world with the click of a button. Those binders full of paper were just as fragile as they were vital.
But I truly felt how fragile the whole affair could be when, some ten years back, I walked through the national archives of land ownership for a nation I will not name. The rooms looked pristine and modern, yet the drawers were filled with rolled, yellowish paper deeds that were, at the time, the only existing copies. I sincerely hope much has changed since then.
It felt so odd to realize that I could have reached in and literally torn apart the only record of land ownership that had passed through so many generations, carrying values that are hard to fathom. Or that an unlucky fire could have broken out and left an entire nation facing a major crisis.
My initial plan was to visit a variety of data-related examples I’ve witnessed over the years, ranging from the very personal to the national. But as I sifted through these experiences, I realized that one of our ongoing projects already brings many of them together. It even has the potential to lead the industry in showing how well-defined data makes AI integration faster and more robust, while also delivering deeper insights.
This unifying example comes from a manufacturing facility with legacy systems accumulated over decades, coupled with modern production lines that create rich and useful cases for our current topic. Although the use of paper was predictably limited, it still appeared in the form of material lists and assembly layouts for operators. These printouts, while not the most environmentally friendly approach, nevertheless made some sense given the limited space available on the assembly line.
The operators seemed at ease with the paper outputs, which made me curious whether they also felt the limitations. Paper does not allow for quick propagation of design changes through a central system, and it lacks interactivity such as zooming in or accessing detailed component information. Still, I would not consider this a clear case where an immediate overhaul is necessary. I cannot say the same for the next level of data maturity in use at the facility.
Just to be clear, this facility boasted a great staff across the board and a genuine willingness to invest in modern technology. It operated using a combination of mature off-the-shelf solutions, such as their ERP and product management systems, as well as custom-developed manufacturing and performance-tracking software. Even so, some gaps remained. One such gap existed between the production planning and assembly departments. A quick and versatile solution, albeit with downsides I will cover next, was to share daily plans using Excel files.
Closing the gap using Excel files is quick, as it does not require additional software development or deployment. It is also versatile, since production schedules can be updated simply by typing in new values. However, this approach comes with downsides that push it into the “overhaul immediately” category for me. Most importantly, it does not support intelligent business logic that can detect errors, optimize processes, or record progress automatically. This creates a clear set of risks and inefficiencies.
There is also the issue of traceability. In other words, this setup does not allow other people in the organization, or external stakeholders, to see the current state of operations. This means mistakes will likely be realized too late, if at all. The lack of broader access also limits historical analysis. In the modern production landscape, there is little room for firms that cannot examine their past performance in detail or explore solutions using AI and code-based analysis.
Not surprisingly, the first application developed by my team was designed to close this gap using well-defined data, error checking, extended access, and more. That same application can also link specific product groups to the digital twin, opening the door to deeper AI methods. But I am getting ahead of myself. Let’s first look at some examples of more mature data use at the facility.
The performance-tracking system used across different stages of production boasted a well-defined database structure that supported the various internal applications in operation. However, to my knowledge, it did not offer API (Application Programming Interface) access for external system integration, which meant we could not rely on the real-time data it captured.
In contrast, the facility’s ERP (Enterprise Resource Planning) system offered the same advantages along with API support. This allowed my team and the facility’s IT experts to develop the new application in just a few weeks. To me, this clearly illustrates why organizations should strive to push their data structures to higher maturity levels so they can respond rapidly to new challenges. Waiting until a crisis strikes to overhaul legacy systems is a gold-gilded, priority-mailed invitation to disaster.
There is yet another step up the ladder, one that we are ascending right now, and that is semantic knowledge graphs, also referred to as ontologies in computer science. At this level, we aim to automate integration across systems and generate new knowledge through code. I look forward to covering this in more detail, but let’s wrap up this section first with another dimension of data that has already paid dividends in this project: fidelity, or the level of detail needed to engage successfully with AI.
The assembly lines that were initially linked by Excel files, and are now digitally bridged, were also something of a black box in terms of performance data. Completed products were recorded in the ERP in batches of several hundred at a time, often at irregular intervals. In essence, anyone trying to analyze the process would see something like, “600 units of product X were completed sometime this morning.” This level of information does not provide deep insights.
Were there stoppages during those hours? Were some assembly lines more efficient than others? How did performance compare with past iterations? Were there steps along the way that caused slowdowns? All impactful questions, yet the answers remained out of reach.
In this case, we developed a relatively simple IoT device, a load cell that detected whenever a product base was added to the assembly line, signaling the start of an assembly process. You may remember a similar setup from the coffee-tracking example in my introduction to IoT. By adding this simple sensor to the line, we now knew the exact pace at which individual items were assembled.
By feeding this high-fidelity data into the AI service, developed in collaboration with a partner firm, we were able to extract detailed breakdowns of the assembly process. Even this simple step revealed how assembly activity was distributed across the week, with Wednesdays predictably leading, as well as noticeable slowdowns about an hour after shift changes, possibly when a buffer of prepared items had run out.
One of the most striking findings was the difference in average assembly time for the same product across different lines and shifts. This observation alone exposed a rich opportunity to transfer best practices between teams, resulting in an immediate return on investment for the project.
At this point, I hope you share my enthusiasm. Getting the best out of technologies such as AI and the digital twin does require a mature data infrastructure. But regardless of the end goal, investing in data maturity is always a smart move. The benefits go far beyond AI readiness, helping organizations adapt more easily to change, reduce human error, and gain clearer insight into both past performance and future opportunities.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.