If you’ve spent any time in energy technology commercialization, you’ve encountered the alphabet soup: TRL, ARL, CRI, MRL, SRL, IRL. Each of these metrics was built to answer a version of the same question: how ready is this thing? And each, in its own way, falls short of giving you a complete answer.
And in most cases, that’s ok.
Metrics and yardsticks are highly important in this space. They enable quantitative assessments. They allow comparisons across projects. They give program managers a shared vocabulary for articulating goals and tracking progress. But they are seldom sufficient on their own. In our experience designing and running DOE programs, the most valuable function of a readiness metric is the conversation that objective setting and the scoring process forces you to have. The number itself is secondary.
From the perspective of a program designer at DOE (or any mission-aligned funder), readiness metrics serve several concrete purposes, and their value begins well before a program launches. During program scoping, structured conversations around readiness across multiple dimensions can reveal the true maturity and challenges of the industry a program intends to serve. That understanding shapes better objectives, more realistic eligibility criteria, and program structures that reflect the actual state of the field rather than assumptions about it.
Once a program is in motion, readiness metrics help target specific challenges: if a program is aimed at bridging the gap between bench-scale validation and pilot demonstration, saying “we’re looking for TRL 4-6 technologies” communicates that instantly. They help manage applicant burden: a well-articulated readiness threshold tells prospective applicants whether a program is meant for them, reducing wasted effort on both sides. And they provide a basis for tracking progress, which feeds into program evaluation and helps the agency learn from outcomes to improve future funding actions.
These are real, practical benefits. The challenge arises when the metrics are asked to do more than they can.
Without a consistent, third-party assessment and validation capability, groups often overstate or misstate their readiness levels. This is especially true with TRL, where the incentive to claim a higher number is strong and the verification infrastructure is thin. A team might genuinely believe their technology is at TRL 6 when a more rigorous assessment would place it at TRL 4. That gap is rarely dishonest. TRL definitions leave enough ambiguity that reasonable people can disagree, particularly as system complexity increases. But it means that readiness scores, taken at face value across a portfolio, can be unreliable as absolute measures. They work much better as comparatives, useful when you’re evaluating closely related technologies against each other, less useful as standalone verdicts.
The bottom line: Readiness metrics are important, particularly to articulate goals and objectives for projects and programs. But use them mostly to drive deep conversations and relationship building, rather than getting hung up on exact definitions and the precise placement of projects on each scale.
TRL is the original and most widely used readiness framework. Developed by NASA in the 1970s to assess technological maturity for complex systems (think Apollo-era hardware), it tracks a technology’s progression from basic scientific discovery (TRL 1) through laboratory validation, prototype demonstration, and ultimately to a system proven in an operational environment (TRL 9). DOE - like many other agencies and research organizations - adopted and adapted the framework for its own technology portfolio. Today TRL is embedded in nearly every funding opportunity announcement across the agency.
TRL’s strength is its simplicity and universality. For standalone technologies with clear performance benchmarks (a new photovoltaic cell chemistry, a novel membrane for carbon capture) TRL provides a useful shorthand for where the technology sits on the development continuum. It’s particularly well-suited for applied R&D, where the primary question is whether the technology works and at what scale it has been validated.
But TRL has real limitations. As system complexity increases, it becomes harder to assign a single TRL score that means anything. A long-duration energy storage system might have a core storage medium at TRL 7 while its power conversion system is at TRL 4, and its integration into a specific grid application is at TRL 2. What TRL is the project? The answer depends on what you’re scoping, and reasonable people will scope it differently. As we explored in our post on the necessity of a systems approach to risk, system-level complexity demands a different analytical lens than any single readiness metric can provide.
More fundamentally, TRL measures technical readiness, hence its name. It tells you nothing about whether the technology can be manufactured at scale, whether there’s a market for it, whether the regulatory pathway is clear, whether the supply chain exists, or whether communities will accept it. A technology can be at TRL 9 and still be decades away from commercial deployment if these other barriers remain unaddressed. For anyone who has worked in energy commercialization, this is familiar territory. TRL is often most useful and complete early in the commercialization spectrum, and becomes progressively less complete as a metric as the solution or technology matures.
TRL is also, among all the readiness metrics, the most susceptible to wishful thinking. Because it’s the most commonly used and often the most consequential for funding decisions, there’s a natural pressure to adjust to align with funding requirements. Without independent validation, self-reported TRL scores should be taken as conversation starters, not conclusions.
The ARL framework was developed by DOE’s Office of Technology Commercialization (formerly the Office of Technology Transitions) and the Office of Clean Energy Demonstrations in partnership with industry stakeholders to address exactly what TRL leaves out. Where TRL asks “does the technology work?”, ARL asks “what else needs to be true for this technology to actually get adopted?”
ARL assesses 17 different dimensions of adoption risk across four core risk areas, then translates that assessment into a 1-to-9 readiness score. The four risk areas cover:
Value Proposition—Assesses the ability for a new technology to meet the functionality required by the market at a price point that customers are willing to pay, to meet the market demand (a broadened definition of “product-market fit”).
Market Acceptance—Captures the target market(s) demand characteristics and risks posed by existing players — including competitors, customers, and other value chain players.
Resource Maturity—Determines risks standing in the way of inputs that are needed to produce the technology solution.
License to Operate—Identifies the societal (national, state, and local), non-economic risks that can hinder the deployment of a technology.
The breadth is the point. ARL’s value is in forcing a comprehensive conversation about all the reasons a technology or solution may fail beyond core technical performance. A team might be laser-focused on hitting a performance target in the lab, but the ARL process pushes them to confront questions like: Is there actually demand for this? Can you get the raw materials or components? Who’s going to build the factory? Will communities where it will be deployed accept it? Is the regulatory pathway understood? These are the questions that determine whether a technology makes it to market or dies in the valley of death despite having a high TRL score.
At the program level, ARL is particularly useful for scoping programs and articulating their challenges and objectives. When OCED designed its demonstration programs, walking through the ARL dimensions with industry stakeholders before programs launched helped the team understand the real maturity and barriers facing potential applicants. That process shaped which adoption risks each program should target, which were outside its scope, and how program structures could be tailored to the actual state of the industries involved. At the portfolio level, it enables comparative analysis across subsectors and investment areas. Because OCED’s portfolio spans everything from carbon capture to industrial decarbonization to small modular reactors, a consistent framework for comparing non-technical risks across these very different technology areas is essential for strategic decision-making.
The trade-off is complexity. With 17 dimensions, an ARL assessment is inherently more involved than a single TRL number. It’s harder to communicate succinctly, which can limit its utility as a quick screening tool. But for anyone making or informing a significant investment decision, the additional nuance is worth the effort. The Commercial Adoption Readiness Assessment Tool (CARAT) that operationalizes the ARL framework is designed to be deliberately simple. Its power is in quickly identifying where critical barriers exist, and the scoring precision matters less than the conversation it generates.
The CRI was developed by the Australian Renewable Energy Agency (ARENA) in 2014 and takes a different but complementary approach. Where TRL tracks technical development and ARL spans the full range of adoption barriers, CRI focuses specifically on commercial and financial maturity: essentially, the financeability of the technology.
CRI uses a six-level scale, from a hypothetical commercial proposition (CRI 1) to a bankable asset class comparable to mature energy technologies (CRI 6). It evaluates eight indicators: regulatory environment, stakeholder acceptance, technical performance data availability, financial cost and revenue propositions, industry supply chain and skills maturity, market opportunities, and company capabilities.
CRI’s design reflects its origin: it was built to help a national renewable energy agency decide where to deploy limited public funding to greatest effect. It serves as a close partner to TRL. In ARENA’s framework, CRI explicitly begins where TRL starts to plateau (around TRL 7) and extends through the commercialization journey that TRL doesn’t cover. If TRL asks “does it work?” and ARL asks “what could prevent adoption?”, CRI asks “can this be financed?”
That’s a valuable question. In energy infrastructure, the ability to attract private capital at reasonable terms is often the binding constraint. A technology that works, has a market, and has community support can still fail if lenders and investors can’t underwrite the risk. CRI provides a structured way to assess and communicate that financial readiness.
However, like TRL, CRI is still fundamentally focused on the core technology and its immediate commercial proposition. It’s less useful for discussing system-level effects: the interactions between a new technology and the broader grid, market structures, policy environment, and infrastructure that determine whether deployment actually scales. And as a framework developed for the Australian renewable energy context, some of its indicator definitions need adaptation when applied to other sectors or geographies.
There’s no shortage of additional readiness frameworks. Manufacturing Readiness Levels (MRL), originally developed by the Department of Defense, assess manufacturing maturity from concept through full-rate production. Supply Chain Readiness Levels focus on the maturity and resilience of input supply chains. System Readiness Levels and Integration Readiness Levels attempt to address the system-of-systems complexity that TRL struggles with, evaluating how well individual components work together rather than assessing them in isolation.
Each of these was developed to solve a real problem, and each has a legitimate use case. But proliferating metrics can create its own problems. If every dimension of readiness has its own scale, you can end up with a dashboard so complex that it obscures rather than illuminates. The goal should be clarity, not comprehensiveness for its own sake.
Based on our experience, a few principles emerge for how to use readiness metrics effectively. (We explore related themes in our posts on the necessity of a systems approach to risk and the wicked problem of program evaluation.)
Use multiple metrics in tandem. No single framework captures the full picture. TRL paired with ARL gives you both the technical and adoption perspectives. Adding CRI brings in the financial lens. The combination is more powerful than any one alone.
Seek consistent, third-party validation where possible. Self-reported readiness (i.e. from a potential recipient or a researcher) scores are useful but come with their own set of biases based on internal perception and, on occasion, the specific funding being sought. Where you’re comparing closely related technologies or making investment decisions, having a consistent assessor apply the framework improves reliability significantly. This is particularly important for TRL, where overstatement is common.
Match the metric to the decision. If you’re designing an applied R&D program, TRL is probably your primary lens. If you’re designing a demonstration program that needs to address commercialization barriers, ARL provides the right framing. If you’re evaluating the bankability of a near-commercial technology, CRI may be your best option. Using the wrong metric for the decision at hand creates confusion and can lead to poorly targeted programs. It takes work to set the targets within a solicitation, and it takes work from every applicant to assess their solution against those metrics. Poorly set targets and poorly chosen metrics causes undue burden, distorts the ground truth of the companies and technologies, and can complicate decision-making and funding landscapes inside and outside of the federal ecosystem.
Treat values as comparative, not absolute. A TRL score or ARL score in isolation tells you less than you think, but consistent evaluation across a set of related technologies and projects, by the same assessor, using the same methodology: that can be useful. Readiness metrics are strongest when used to compare, weakest when used to make binary go/no-go decisions.This is particularly a challenge for managers or organizational leadership. There is an inherent desire to boil large complex programs or portfolios of projects down to simple metrics. While the desire to distill the complex to the simple is only human, there is the risk that rigidly adhering to specific numbers can lead to misunderstandings or a false sense of certainty. Pushing for specificity on these values when it is not imperative to moving work forward, can lead to busy work for staff and disenchantment with using the metrics at all leading to a loss of all value.
And perhaps most critically:
Use them to facilitate and sustain conversations. This applies before programs launch, not just during selection or oversight. Sitting down with industry stakeholders and walking through readiness dimensions during program design surfaces assumptions about industry maturity, identifies barriers that the program should (or should not) try to address, and builds shared understanding that makes the eventual program more grounded in reality. During selection and project management, the same process surfaces blind spots and builds shared understanding across disciplines that might not otherwise be in the same room.
Ensuring energy sector competitiveness, affordability and sustainability depends on moving technologies from lab to market faster and more reliably than we ever have before. Readiness metrics are one of the tools we have for navigating that journey. They’re most useful when we’re honest about what they can and can’t tell us, and when we treat them as the start of a deeper conversation rather than a final answer.
Innovation Waypoints is brought to you by Waypoint Strategy Group.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.