RSS Amplifier

Edward's Substack · Jul 26, 2026

ROI: TBD Is Not an AI Strategy

0
Sign in to vote or save

Edward J. Liebig · Edward's Substack

The methodology we introduced at NexGenomics to discover and govern meaningful AI opportunities did not begin as an AI methodology.

Its roots go back decades to our earlier work assessing Chemical Facility Anti-Terrorism Standards, or CFATS, requirements, cybersecurity, and operational risk in industrial control system environments.

At the time, we kept encountering the same problem. Traditional technology risk assessments often began with assets, threats, vulnerabilities, and estimated probabilities. That approach could tell us a great deal about the technology, but it did not always answer the most important operational question:

What are we actually trying to prevent from happening?

In an industrial environment, the real consequence is rarely the movement of a packet or the failure of a server by itself. The consequence is what that digital event can ultimately influence: an unsafe operating condition, an environmental release, degraded product quality, lost production, damaged equipment, or a decision made from incorrect information.

That realization changed the discovery path.

Rather than beginning with the technology and working outward, we learned to begin with the objective, or the unacceptable consequence, and work backward.

  • What outcome must be preserved?

  • What could cause that outcome to be lost?

  • Which decisions, controls, systems, people, and operating assumptions protect it?

  • What evidence tells us those protections are still working?

  • Where can a failure, misunderstanding, or change propagate through the organization before anyone recognizes its significance?

Over time, that consequence-first approach evolved into what our teams use today. At NexGenomics, we call it CIE-inspired Consequence-Oriented Risk Evaluation, or CIE-CORE. It provides a structured way to begin with an important outcome, bound the system or process that influences it, trace the relevant pathways, identify the evidence required to understand it, and determine how authority, influence, and intervention should be governed.

Continuous Digital Physical Validation, or CDPV, extends the same reasoning into ongoing operation. It asks whether the assumptions, controls, behaviors, configurations, and evidence that are supposed to preserve an outcome remain true under current operating conditions.

Although this methodology grew from industrial and operational environments, the emerging challenge of using enterprise AI effectively makes it broadly relevant.

Organizations are currently approaching AI in much the same way many organizations once approached cybersecurity and digital transformation. They begin with the technology.

  • Which model should we use?

  • Which platform should we buy?

  • Where should we deploy agents?

  • How quickly can we demonstrate adoption?

Those are legitimate questions, but they are not the appropriate starting questions.

Before deciding where to place AI, an organization should first understand what it is trying to improve, protect, accelerate, or make more reliable.

The AI opportunity does not begin with the model.

In this context, “consequence” does not have to mean catastrophe. It may be a customer who leaves, a payment that arrives late, a quality issue that reaches the market, a maintenance decision made without complete evidence, a regulatory obligation that is missed, or an expert whose time is consumed reconstructing information that should already be available.

  • These are outcomes the organization already values.

The job of AI discovery is to determine where better correlation, synthesis, prediction, recommendation, or bounded action can measurably improve those outcomes.

That brings me to something Reet Kaur shared in an original post that captured the problem rather well:

CEO's AI strategy deck with 47 slides. 📊
Slide 1 is the vision.
Slides 2 to 46 are the roadmap.
Slide 47 says "ROI: TBD."

The board gave a standing ovation. 👏 👏

It is funny because it is close enough to reality to make people uncomfortable.

Across nearly every industry, organizations are being pushed to demonstrate that they have an artificial intelligence strategy. Executive teams are building roadmaps. Boards are asking for updates. Business units are launching pilots. Vendors are presenting compelling demonstrations of what their models and agents can do.

The vision slides usually look excellent.

The problem tends to appear when someone asks a more basic question:

What, exactly, will operate differently after we deploy this?

That is often where the discussion becomes vague.

The organization may know which AI platform it wants to use. It may have selected model providers, created a governance committee, approved a pilot portfolio, and built an impressive adoption timeline.

But none of those things, by themselves, creates value.

“ROI: TBD” is not simply a missing financial calculation on the final slide. It is often evidence that the strategy started too far away from the work.

A company can have an AI vision without having identified a meaningful operational change.

Artificial intelligence can summarize documents, correlate data, generate content, classify information, detect patterns, recommend actions, and increasingly execute multi-step tasks. Those are valuable capabilities.

But capabilities are not outcomes.

A summary is not an outcome. A recommendation is not an outcome. An agent completing a task is not necessarily an outcome either.

The real outcomes are things the organization already cares about: reducing order-to-cash time, increasing production throughput, improving first-pass quality, lowering unplanned downtime, resolving customer issues faster, reducing regulatory reporting effort, or enabling scarce experts to manage more work without sacrificing quality.

We call this Human Amplification.

AI does not possess an independent value proposition inside the enterprise. Its value appears only when it changes the performance of a defined process, decision, or operational outcome.

This is why organizations struggle to calculate ROI after their pilots. They often measure AI activity rather than operational change.

How many people used the assistant? How many prompts were submitted? How many documents were generated? How quickly did the model respond?

Those numbers may tell us that the technology was used. They do not necessarily tell us that anything improved.

The useful question is not, “Did employees use the AI?”

It is, “Did the work become measurably better?”

That is the discovery path organizations now need.

Begin with the outcome. Identify the processes, decisions, evidence, and authority that shape it. Determine where friction, fragmentation, or unvalidated assumptions are preventing the organization from achieving its objective. Then decide whether AI has an appropriate and governable role in improving it.

The order matters.

  • Consequence first.

  • Technology second.

  • Proof before expansion.

Most AI discovery sessions begin with a seemingly innocent question:

Where could we use AI?

That question encourages people to imagine technology. The result is usually a wish list of chatbots, assistants, automated reports, intelligent search tools, and predictive dashboards.

Some of those ideas may be useful. The problem is that they are being proposed before the organization has identified where value is being lost.

A stronger conversation begins with operational accountability.

What outcome are you responsible for? Where is it constrained? What creates delay, inconsistency, rework, cost, or exposure? Which decisions shape the result? What evidence supports those decisions? Where does that evidence become fragmented, stale, or difficult to assemble?

These questions lead people away from novelty and toward the actual work.

The process, not the model, becomes the unit of discovery.

This also prevents organizations from automating activities simply because they are repetitive. Repetition alone does not make a task valuable.

A decision made only a few times each week may be far more consequential than thousands of low-value administrative actions. An exception handled by a senior employee may carry more opportunity than an entire department’s routine reporting.

The objective is not to find the greatest number of tasks.

It is to find the operating thread where better context, judgment, coordination, or execution will materially improve an outcome.

A meaningful discovery effort should move vertically through the organization and horizontally across the work.

The vertical walk connects executive intent to frontline reality.

Leadership may describe an objective such as improving customer retention, reducing operating cost, increasing capacity, or accelerating revenue. That objective should then be traced downward.

Which process produces that result? Who owns it? Who supervises it? Who performs the work? Which systems support it? Who has authority to approve, reject, override, or escalate a decision?

This often reveals a gap between how leadership believes the process works and how employees actually get the work done.

The approved procedure may show a clean sequence of steps. The frontline experience may depend on email, spreadsheets, undocumented judgment, personal relationships, and historical knowledge.

Both views matter.

The horizontal walk follows the work from its original trigger to its final result:

Trigger → interpretation → decision → approval → execution → verification → exception handling → reporting → learning

Do not stop at departmental boundaries.

A maintenance event, for example, may involve operations, engineering, reliability, procurement, warehousing, finance, safety, compliance, and an outside service provider.

The delay may not be caused by the technician performing the repair. It may be caused by missing history, an unclear approval threshold, an unavailable part, a supplier response buried in email, or a handoff no one clearly owns.

Many of the strongest AI opportunities are not jobs waiting to be replaced.

They are broken threads waiting to be connected.

Once the process is mapped, the next step is to determine what the evidence says about how it actually performs.

That evidence may exist in transactions, documents, tickets, logs, work orders, approvals, emails, sensor readings, contracts, or previous decisions. Together, these sources reveal where information is incomplete, where judgment is applied, and where friction is already costing the organization time or money.

Begin with the baseline. How is success measured today? Is the organization tracking cycle time, quality, throughput, availability, cost, customer satisfaction, risk, or revenue? Without an accepted starting point, there will be no credible way to prove that AI improved the outcome.

Pay particular attention to exceptions and decisions. The most valuable opportunities often appear when the normal process breaks down, trusted sources conflict, or an experienced employee must reconstruct context from several systems. AI may be useful in assembling evidence, identifying similar cases, prioritizing work, or preparing a recommendation—but its role should reflect the consequence of getting the answer wrong.

This is also where Human Amplification becomes practical. When a process depends on someone who “just knows how it works,” the objective is not necessarily to replace that person. It may be to give the expert faster access to the right evidence, preserve the reasoning behind decisions, and make appropriate portions of that knowledge available to others.

Finally, connect the friction to its economic effect. Overtime, rework, delayed revenue, excess inventory, production loss, penalties, consultant dependency, and customer churn are not abstract inefficiencies. They are the current cost of the process.

Finance should help establish that cost before the pilot begins—not be asked afterward to invent the ROI.

A promising process usually has a meaningful outcome, a recognizable owner, repeatable work, and a measurable level of avoidable friction.

It occurs often enough to establish a baseline. Delay or inconsistency creates a real consequence. The work depends on gathering, comparing, interpreting, or routing information. The boundaries are understandable. Exceptions can be identified. Human authority can be preserved where necessary.

Evidence-intensive processes are often strong candidates. Audit preparation, contract analysis, claims review, regulatory reporting, due diligence, incident documentation, and executive reporting all require people to assemble information from multiple sources and form a coherent position.

High-friction handoffs are another productive place to look. Sales-to-delivery, engineering-to-operations, procurement-to-receipt, maintenance planning, customer escalation, and shift turnover all create opportunities for information to be delayed, misunderstood, or lost.

Repetitive expert judgment may also be valuable, particularly when AI prepares the context while a qualified person retains the decision.

The opportunity grows when the outcome is valuable, the process occurs frequently, friction is expensive, and decision latency matters.

It shrinks when the evidence is unreliable, integration is impractical, exceptions are unbounded, or the required authority creates more risk than value.

A process can be ripe for AI and still be unsafe to automate.

This is where many strategies move too quickly.

Once an attractive opportunity is identified, the conversation often jumps directly to model selection, application development, or agent deployment.

But the underlying architecture determines whether the organization can trust the AI during normal operation and whether it can constrain the AI under pressure.

That second requirement is critical.

Policies, committees, acceptable-use standards, and review processes all matter. But governance that exists only in documents may not be enough when an AI system receives malicious input, uses incorrect context, invokes a compromised tool, exceeds its intended role, or begins propagating errors across connected systems.

The architecture must be capable of enforcing the organization’s intent.

An agent should not simply be told which data it may access. Its identity and permissions should technically prevent it from reaching anything else.

It should not merely be instructed to avoid unauthorized external communication. Its outbound access should be restricted to specifically approved destinations.

It should not simply promise to seek human approval before taking a consequential action. Execution should remain impossible until the required authorization is received.

That is the difference between procedural restraint and architectural constraint.

A trustworthy AI foundation should be able to isolate workloads, separate sensitive information, enforce least privilege, govern model and tool access, preserve data lineage, record actions in tamper-resistant logs, and restrict external communication.

The organization should be able to determine:

  • Which agent was operating

  • What purpose it was authorized to serve

  • Which information it accessed

  • Which model and tools it invoked

  • What recommendation or output it produced

  • Who approved the next step

  • What action ultimately occurred

Not every AI system should receive the same level of influence.

One may summarize information. Another may prioritize cases. Another may prepare a proposed action for human approval. A mature deployment may orchestrate an approved workflow across several systems.

Those are materially different levels of authority, and the architecture should recognize and enforce the distinction.

At NexGenomics, we address this through the concept of an Influence Contract. Before an agent is deployed, its purpose, permitted data, allowed outputs, prohibited actions, approval requirements, confidence thresholds, rate limits, evidence obligations, drift criteria, and response procedures should be explicitly defined.

The most important terms should be enforced by the infrastructure, not left entirely to prompts or policy.

The pressure I am most concerned about arises when an AI is given an objective, access to tools, and enough freedom to find its own path to the result.

A person may interpret a task within the obvious boundaries of the request. An AI system may continue searching for another pathway when the expected one is blocked. It may combine tools, permissions, interfaces, stored context, or system weaknesses in ways its designers did not anticipate.

That behavior does not necessarily begin with malicious intent. It can emerge from the system pursuing the assigned goal too effectively while lacking a reliable understanding of which boundaries must remain inviolable.

The recent OpenAI–Hugging Face incident illustrates why this matters. During a cybersecurity evaluation, advanced OpenAI models chained vulnerabilities across the testing environment and Hugging Face infrastructure while pursuing the narrow objective of obtaining answers that would help them complete the evaluation.

The issue is not simply whether an AI follows a bad instruction. The deeper concern is whether, while pursuing an authorized objective, it can discover and use an unauthorized path to reach the result.

A trustworthy architecture must therefore constrain more than the agent’s stated purpose. It must constrain the routes available to it.

The infrastructure should be able to prevent access outside the approved scope, limit which tools and systems can be combined, require authorization before consequential steps, isolate each execution environment, and stop the activity when behavior departs from the intended operating boundary.

It should also preserve the evidence needed to show what the AI attempted, which pathways it explored, what it accessed, and where the architecture intervened.

This is not only a security requirement. It is a trustworthiness requirement.

A process does not create sustainable ROI if it becomes faster while giving an AI uncontrolled freedom to improvise its way toward the objective.

The strongest opportunity is therefore not simply a valuable process with available data.

It is a valuable process that can operate inside a trustworthy architecture that permits useful creativity while making prohibited pathways architecturally unavailable.

AI should not be used to conceal a process the organization does not understand.

A process is not ready if no one owns the outcome, leaders disagree about what it is intended to accomplish, the governing policy remains unresolved, the evidence is unreliable, or most cases involve unbounded exceptions. It may require clarification or redesign before automation begins.

The organization must also define acceptable error.

A model that is correct 95 percent of the time may sound impressive. Whether that is sufficient depends entirely on the consequence of the other 5 percent.

A low-confidence marketing classification and a low-confidence safety recommendation do not belong in the same governance category.

Consequence determines the controls.

It is also worth challenging projects whose economic case depends entirely on removing people. Labor reduction may be part of the value equation, but it should not be the only one.

The stronger proposition is often to help people complete work faster, manage more complexity, reduce avoidable mistakes, and apply expert judgment more broadly.

Automating a poorly understood process does not remove dysfunction.

It makes dysfunction operate at machine speed.

Before an AI project begins, the organization should be able to explain the process in plain language.

Who owns the outcome? What is the current baseline, and what improvement is expected? Which part of the process is inside the pilot? What information may the AI use, what may it produce, and what is it prohibited from doing? Where is human approval mandatory? What happens when confidence is low, evidence conflicts, or the situation falls outside the approved boundary?

Those answers form a value and authority contract.

The business side defines the intended improvement. The governance side defines the operating boundary. The architecture enforces the most important constraints. The evidence chain allows the organization to determine what actually happened.

At the end of the pilot, the question should not simply be whether the model performed well.

The question should be whether the process improved while remaining inside the agreed boundary.

Organizations often attempt to build enterprise-wide AI programs before proving value in one bounded operating thread.

The result is usually a broad portfolio of pilots competing for data, integrations, security review, funding, executive attention, and user adoption.

A better path is bounded expansion.

Start with one consequential process, one accountable owner, one clearly limited AI role, one measurable outcome, and one complete evidence chain.

The first implementation does not need to automate the full process.

It may begin by assembling evidence for a decision, classifying incoming work, identifying exceptions, or preparing a recommendation for human review.

That is enough to begin proving value.

Once the organization understands the behavior, limitations, economics, and governance requirements of that deployment, it can expand into adjacent steps and related processes.

Expansion should occur because the evidence supports it, not because the roadmap says the next phase begins this quarter.

The final test is straightforward.

Did the operation improve?

Did cycle time decline? Did first-pass quality increase? Did experts regain time? Did avoidable escalations fall? Did customer response improve? Did the organization reduce rework, downtime, reporting effort, cost, or exposure?

Those measures matter more than prompt volume, user registrations, model response time, or the number of agents deployed.

Model performance remains important. It simply belongs inside a larger operational measurement.

The model may be accurate. The application may be well designed. The users may enjoy it.

But if the process did not improve, the organization has not yet produced a return.

The AI strategy deck does not need to end with “ROI: TBD.”

It should end with something closer to this:

  • We selected one bounded operational process.

  • We defined the desired outcome and established a baseline.

  • We identified the evidence required to support the work.

  • We constrained the AI’s access, influence, and authority.

  • We established where human judgment remains mandatory.

  • We defined how improvement and failure will be measured.

  • We will expand only after the results are validated.

That may not generate a standing ovation.

It should generate something more useful: confidence.

The board should not approve an AI strategy simply because the vision is compelling or the roadmap is comprehensive.

It should approve the strategy because the organization can identify the first outcome it intends to improve, the evidence required to prove that improvement, the authority AI will and will not receive, and the conditions under which the capability will be permitted to expand.

That is the difference between deploying AI and operationalizing value.

  • Discover the outcome.

  • Follow the evidence.

  • Map the influence.

  • Bound the authority.

  • Validate the improvement.

  • Then expand.

Copyright © 2026 NexGenomics. All rights reserved.

No posts

Read the original on edwardjliebig.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.