Acknowledging Wiebke Hutiri & Ricardo Segundo for their review of this post.
The most memorable, high-paced, and high-tech job I ever had was as a Data Scientist at a large financial company — a truly humbling experience that helped me grasp just how risky data initiatives can be from a Return on Investment (ROI) perspective, and just how essential sound Data & AI strategies are in order to be on the “enormous upside potential” track and avoid the “financial sinkhole” track .
While the company as a whole was struggling, our department grew from nothing to a staggering 200 FTEs in a matter of years—“feels like a Ponzi scheme” we often joked. Those 200 weren’t just anyone—most were PhDs in STEM fields or Palantir software engineers. It’s certainly a humbling experience to work alongside A-teamers and boy were some of those skills remarkable — not good for my imposter syndrome, but as they say “if you’re not confused you’re not learning”. Needless to say, the budget allocated must have been astronomical.
I was in a team of about 50 FTEs, comprised of platform managers, data engineers, data scientists, and product managers. Details can’t be shared, but the data lineage of our software looked something like this:
We built it all on Palantir’s fancy Foundry platform — a company whose data software is miles ahead of anything I have ever used — and it worked damn well. The sophistication and deep attention to software engineering best practices were awe-inspiring. I recall a series of meetings where the lead engineer led heated “column nomenclature” discussions—yes, indeed—discussions on how the columns in the gold tables should best be named for ease-of-understanding for end consumers.
It was those moments that a deep appreciation for data system design took root - those atomic every-day decisions that, if disregarded, stack on top of each other, resulting in data systems to creep into the realms of code spaghettification leading to costly technical debt.
The company needed our software—we strongly believed it had potential for value-add across the organisation for a number of different use cases. So we reached out to business stakeholders to begin to train them around what we’d done, and ideate with them on how our novel tech solution could help them. At this stage the generic technical core of our software had been built, but the final business applications for specific departments needed to be worked out — and we understood very well that the business stakeholders needed to be part of this journey.
We asked them to formalise their requirements, because they were our stakeholders, so we needed formal requirements to know precisely what we had to deliver to them.
But in practice it was more like…
This was the same story, over and over again. Perhaps it sounds familiar? I have seen this movie play out in a number of different companies. Ultimately it’s a complex change management challenge—one requiring cultural shifts at every level. This starts with senior managers aligning around the idea that we’re all on the same team (no them vs us), and not delivering for siloed priorities.
Easier said than done. Cross-team communication is time consuming and draining—we speak different languages. Workplace dynamics also mirror the Prisoner’s Dilemma, driving temptation for short-term personal gains and fear of exploitation—under tight deadlines, cooperation falls aside as teams focus narrowly on their own deliverables. The result? Suboptimal outcomes for the company.
Breaking this cycle demands intentional strategies—like transparent incentives or trust-building rituals—to override the “defect” me-first instinct.
The problem to solve, that connects us all, can be easily framed:
That’s the data-value gap that needs to be bridged. More realistically, a time dimension should be added. Then it looks like this:
Because of:
the realisation by executives that data is an asset that has the potential to drive a competitive edge
the expansion of IoT devices that capture streams of data
the nature of Deep Learning architectures that enable their predictive performance to increase as a function of the size of the datasets they are trained on
the fact that the cost of data storage has decreased exponentially from the 1950s to now
companies now store more and more data.
While your data is the asset itself that can give you a long-term competitive edge, the hard truth is that storing data does nothing to generate value. You have to do something useful with that data — think personalised marketing; customer retention modelling; dynamic price optimisation; predictive maintenance; or supply-chain optimisation.
Additionally, with recent breakthroughs in generative AI — all triggered by a landmark research paper published by Google in 2017 that has an astonishing 168’121 citations as of this writing — text-based and agentic AI systems have clear potential to automate mundane tasks, as well as be a guide in business critical decisions. How? Because they can understand your internal structured and text data (see RAG). The list of potential ideas is enormous.
As with my personal example above, bridging the gap is a formidable undertaking—there are countless roadblocks along the way that it often feels like all the planets need to align to make it happen. According to a 2024 study by the RAND corporation, 80% of AI projects fail. These low success rates are indeed noteworthy given the competitive advantage that such advanced techniques can provide if you’re part of the 20%.
In this Substack publication, I’ll be writing about varies strategies to bridge the data-value gap, generally focusing on research-backed methods to maximise for ROI in Data & AI.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.