It feels like an old New York City apartment that they just keep painting over and over and over again. And you’re sorta like, what color was this supposed to be? We should strip the paint and we should take a look and say, what are we trying to achieve?
-Sharon Yeshaya, Morgan Stanley CFO at the Federal Reserve’s 2025 Capital Conference
Policy sludge comes in different forms.
A recent paper I co-authored distinguishes between: (A) “horizontal sludge,” where rules overlap and conflict, (B) “vertical sludge,” where implementation of a statute from rule writing to compliance balloons a bank’s obligations, and (C) “cumulative sludge,” which builds up over time like barnacles on a ship or layers of paint in a New York City apartment.
This post focuses on cumulative sludge.
In particular I analyzed the duties of large bank boards of directors and the entire body of supervisory guidance from the U.S. federal banking agencies since the 1990s. The findings of each point in different directions and highlight the importance of doing deep analytical work, not just relying on priors, when approaching policy sludge clean up. Unfortunately, such work has been rare in the past because of the staff effort required.
As a result there’s been a chronic deficit of policy sludge analysis. Without objective, systematic analysis, it is difficult for agencies to prioritize which sludge problems to tackle and how to tackle them.
AI changes that calculus. It is now possible to do in days or weeks what took entire teams months or quarters. With the right tools and approach, regulatory agencies can now systematically analyze policy sludge and generate streamlining roadmaps without incurring debilitating opportunity costs.
This post provides a glimpse into how that can be done by working through two practical examples of cumulative policy sludge: bank board duties and supervisory guidance.
Nearly every statutory obligation and supervisory expectation for large U.S. bank boards of directors appears prudent and reasonable on its own. Over time, though, these accrete. The graph below tracks the dramatic growth in board duties over the last 40 years. Most were added after severe corporate governance weaknesses were revealed in the 2008 global financial crisis.
Inventorying all board obligations and supervisory expectations is not easy, even with AI, as they are scattered across a wide range of statutes, regs, and guidance documents. Every time a provision states, “boards should…” or “…such-and-such report shall go to the board”, it adds to the list of board duties.
When I asked earlier AI models to inventory large bank board duties, they tended to look up just a handful of regs — e.g., related to capital, liquidity and resolution planning — and miss everything outside of that (which is quite a lot).
The most recent frontier models are substantially better, though. And several AIs specializing in law and regulation have emerged.
I began this round of analysis by using a preview version of Caffrey, a specialized financial regulatory AI platform. Caffrey leverages frontier LLMs like ChatGPT and Claude, but augments them with a large annotated database of all U.S. financial regulations, enforcement actions, guidance documents, and other sources. Caffrey also uses a proprietary method involving pre-research, citation checks, and other techniques, to convert user prompts into legally-precise, regulation-aware instructions and to strengthen the final output.
For instance, this was my original prompt to Caffrey:
I need help identifying "policy sludge" -- i.e., duplicative, overlapping, and unduly burdensome requirements that have accumulated over time. In banking, a common complaint relates to the cumulative requirements for boards of directors, which are sprinkled throughout many regs and statutes. Build a comprehensive inventory of all obligations of bank boards of directors, so that I can begin to tackle this problem.
Caffrey grilled me with a series of pointed questions, then converted my three sentence request into the following prompt:
Compile a comprehensive, citation-level inventory of all current, in-force obligations imposed on boards of directors of Category I global systemically important bank holding companies and their insured depository institution subsidiaries, drawing from all applicable federal regulations and agency guidance issued by the Federal Reserve, the OCC, the FDIC, and the CFPB, as well as relevant interagency standards. The inventory should capture obligations at every level of board structure, including obligations of the full board, obligations specific to individual board committees such as audit, risk, compensation, and nominating or governance committees, and any obligations that distinguish between inside and outside directors or between independent and non-independent directors. The research should address obligations arising across all major regulatory domains applicable to Category I GSIBs, including capital planning and stress testing, liquidity risk management, resolution planning, recovery planning, operational resilience, model risk governance, third-party risk management, incentive compensation, cybersecurity governance, BSA and AML program oversight, consumer compliance oversight, and safety and soundness standards generally. For each obligation identified, the inventory should specify the precise regulatory source, the nature of the board-level duty, whether the obligation falls on the full board or a specific committee, and any distinction between inside and outside or independent directors. The intended audience is bank directors and their legal counsel, as well as policymakers engaged in regulatory reform efforts aimed at identifying and reducing duplicative, overlapping, and unduly burdensome board-level requirements. The output should be granular and citation-specific, suitable for use as a working policy reform document, and is expected to be extensive, approximating 40 to 60 pages in depth.
(The augmented prompt reminded me why my career as a lawyer was quite short!)
I separately fed my original prompt and the Caffrey-augmented prompt into GPT 5.5, Opus 4.8, and open weight models Kimi 2.5 and GLM-5.2.
To systemically assess the responses, I built three LLM-as-judge evaluators, using Codex (GPT5.5) Claude (Opus 4.8), and Gemma 4B (an open weight model from Google). Each evaluator assessed each report for its: (1) coverage and comprehensiveness, (2) source and citation integrity, (3) substantive accuracy and currency, (4) sludge analysis quality, (5) structure, granularity, and fitness for purpose, and (6) calibration and transparency, i.e., honesty about limits. In addition, each evaluator conducted sample testing for critical weaknesses, such as fabrications, misattribution, and false exhaustiveness.
Overall, Codex was the harshest grader, providing the largest differences in scores between models and reports. (Notably, Codex graded GPT 5.5 highest, while Claude graded Opus 4.8 highest — a bias that has been noted by AI researchers.)
Interestingly, Caffrey’s expanded prompt nearly doubled the scores of GPT5.5 and Opus 4.8 versus their responses to the original prompt (GPT5.5-orig and Opus4.8-orig, respectively). In other words, Caffrey’s harness, which converted my loosely worded request into a legally-precise and significantly more comprehensive prompt, enabled the frontier models to double their measured performance.
While the evaluator scores are interesting, they are just rough indicators. What matters to policymakers is the quality and credibility of the content.
The reports generated by each of the models had different strengths and weaknesses. I asked Opus 4.8 to compare, contrast, and synthesize the best from each model’s outputs. After several turns, it synthesized everything into a unified 67-page report and accompanying spreadsheet, both of which are highly detailed and nuanced. For instance, below is an excerpt from the unified report summarizing the analysis of board obligations:
The report also analyzed and ranked areas of high overlap and high burden:
In short, the report seemingly included everything that a reform-minded policymaker could want. BUT, while being a solid first draft, in practice a report like this would need serious senior level policy and legal review before being ready for formal consideration and decisioning by an agency.
The key takeaway here is that generating this draft took just several days of working with AI, rather than several months of manual work by an entire team (as was the case when the Fed reviewed board expectations roughly a decade ago). The staffing costs of doing such analysis have fallen dramatically, opening up a range of possibilities for improved policymaking going forward.
Imagine an annual streamlining process where high value sludge targets are identified and cleanup roadmaps are adopted and implemented. With today’s AI models this is now in the realm of the feasible.
Everything related to this project, from the model reports and the evaluator tool to the evals and final synthesis, is available here. I encourage readers — especially banking attorneys and policy analysts — to take a look.
Another area that has been cited as accumulating policy sludge over time is bank supervisory guidance.
Last year the Reserve Bank of India (RBI) actively addressed this by compressing 9,445 circulars into 244 Master Directions, marking one of the most ambitious policy sludge clean up efforts by a financial regulator. It reportedly took 40+ staff more than a year to accomplish.
I wondered if it was possible to take what RBI had done manually and equip an LLM to do it agentically for others. Specifically, I asked Opus 4.8 to reverse engineer RBI’s streamlining and develop a SKILL1 for LLMs that other agencies could use to achieve RBI’s results.
The initial analysis unpacked how RBI executed the streamlining, noting that it took four years of preparation and consisted of 5 design principles (e.g., as-is consolidation, full traceability, etc.) and 8 phases. That informed the creation of a SKILL.md file, which readers can download and use with any LLM.
I used the SKILL to analyze 30+ years of supervisory guidance issued by the Federal Reserve, FDIC, OCC, and UK PRA and FCA.
I happened to do this during the brief window when Anthropic’s Fable 5 was available. It was a beast, happily crunching through materials using the Wayback Machine to identify and reconstruct long forgotten guidance releases and terminations over the years, as depicted visually here:
This table summarizes its findings for the FBAs:
Notably, the Fable 5 analysis shows that the stock of in-force guidance at each agency has remained relatively constant over the years, indicating that the problem is not accretion. Rather, the problem is “navigability and verifiability,” i.e., guidance letters issued without an overarching framework and retracted without any formal signal.
The complete sludge analysis is here. The SKILL.md file and associated materials are available here for policy innovators to experiment with and test out.
Cumulative policy sludge hinders effective policymaking, but comes in different sub-flavors. This post honed in on two examples, using AI to do deep research to analyze and identify the policy sludge profiles of large bank board duties and of supervisory guidance. The divergent findings highlight the importance of conducting rigorous analysis, and the demonstrated efficiency enabled by AI should give policymakers hope of what’s now possible when it comes to policy streamlining and sludge cleanup.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.