RSS Amplifier

Mike Hsu · Jul 8, 2026

Cleaning up

0
Sign in to vote or save

Mike Hsu · Mike Hsu

Visual summary of streamlining 227 unorganized board duties (current) to 26 consolidated duties (proposed). BoardDistill graphic created with Claude Design

My last post identified the cumulative sludge of 227 unorganized board of director duties1 that has accumulated at GSIBs over the years. AI helped compress the time and effort required to robustly identify the scope and nature of the problem. Notwithstanding, to clean up sludge efficiently, credibly, and neutrally — to resolve policy sludge — we must do more than merely classify it. We need a theory and process for practically streamlining a messy corpus.

AI can surely help, but how exactly?

At the recent Open Source in Finance Forum (OSFF) conference in London, I attended a presentation, “Distilling Intent from Code” by Henry Garner, CTO of Juxt. He shared a riveting story about finding a bug in the Apollo 11 code, situating it in the broader challenge of keeping legacy code current and in Peter Naur’s “Programming as theory building.” The accumulation-of-code-over-time problem sounded eerily familiar. I scribbled in my notebook, “‘legacy code = policy sludge?”.

We specify, because we want to communicate clearly. And we keep specifying, because we are always in the process of finding out what we mean.

- Henry Garner, “On co-specs and tri-checks

Garner’s firm has developed a solution to the legacy code problem: distill from a codebase what the overall intent is and use that to contextualize the existing piecemeal specs to figure out what the canonical overarching spec should be, and then utilize that to elicit fully aligned code and to weed out excessive legacy code.

This approach allows one to step back and ask, “What are we actually doing here (distillation)? What do we want the codebase as a whole to achieve (canonical overarching spec)? Let’s make sure it does what we want (elicitation) and let’s also make sure we keep it that way (weeding),” without throwing everything away and starting from scratch.

The tool for doing this is called Allium and it is open source.

When I described this to my wife, who does criminal justice reform work, she recalled a project she had worked on for the DC government almost a decade ago. I looked up the details and was stunned.

In 2016 the DC Council formed an independent commission to overhaul the city’s criminal code. As part of the streamlining project, the commission engaged the public through multiple hearings, feedback sessions, and revision drafts; it didn’t delegate the project to technocrats to quietly execute. It operated transparently, encouraging community input, and systematically recording and sharing its work.

The commission documented everything — from every oversight hearing to every draft legal analysis to every advisory group report. It methodically inventoried each criminal provision that had accumulated on DC’s books since 1901 and recorded the dispositioning of every single one into a proposed revised criminal code, including the rationale for every change, e.g., for merging, splitting, rewriting, or retiring each provision.2 The project took over five years and the commission made all of its work public, posting over 70 documents.

In short, DC decided to tackle over 100 years of accumulated criminal code sludge, ran a highly transparent and democratic process to clean it up, and meticulously mapped its thought process — provision by provision — from a sludge-filled corpus to a clean, revised code.

I wondered if Allium could be operationally adapted to diagnosing and addressing policy sludge, not just legacy code — and if the DC Criminal Code Reform Commission’s work could be used to test and train it.

Using Opus 4.8 and GPT5.5, I developed a v0.1 legal and regulatory adaptation prototype of Allium, called Lex-Allium. Instead of a (legacy) codebase, it works with (legacy) statutes and regulations. Instead of lines of code, it analyzes legal and regulatory provisions. While the objects are different, the concepts and techniques of distilling, eliciting, and weeding are the same — they are just pointed at policymaking instead of coding.

I then used Fable 5 (which had just been made accessible again) to test and sharpen Lex-Allium for statute distillation using the DC code revision as a backtest. I wanted to know: could Lex-Allium be utilized to distill the (pre-reform) DC criminal code in a manner that would enable it to replicate what the DC reform commission did?

After analyzing the public reports released by the commission and doing deep research, Fable 5 distilled 14 types of changes that the commission made to transform the accumulated criminal code provisions into the revised criminal code. These included merging overlapping offenses or making an unstated requirement explicit. It called these 14 moves the “transformation calculus”.

Screenshot from an explainer of StatuteDistill, created with Claude Design

This was promising. But criminal law is quite different from civil and administrative law. How generalizable is this technique? After further analysis and testing by Fable 5, it looks like most of the transformation calculus is translatable to other fields, with a few adjustments on the margin depending on the domain.

So, we now had a generalizable technique for distilling a legacy body of statutes — StatuteDistillsimilar to what Allium enables for legacy codebases.

What if I aimed this at the accumulated sludge of large bank board duties? Would that enable us to consolidate the mishmash of duties accumulated over the years into a streamlined, coherent set in a neutral, non-deregulatory manner?

Preliminary analysis indicated that StatuteDistill could be applied to the inventory of 227 board duties that had been identified. But some notable differences were flagged.

First, the DC code revision was anchored to the Model Penal Code. Bank board duties have no Model Penal Code equivalent. There is no consensus reference model of what board governance should look like. So the anchor had to come from elsewhere. Two substitutes emerged: a statutory floor set by the duties mandated by Congress (e.g., the Dodd-Frank risk committee, audit committee independence under SOX, the BSA program statutes); and a clear hierarchy (i.e., statute over regulation over enforceable guideline over guidance over exam manual) in place of the case law analysis for criminal law.

Second, the DC Council was a single legislature that could enact whatever the commission proposed. Board duties span five agencies. Every proposed change therefore carries two extra fields: who has the authority to act, and what action is required (e.g., a guidance action, an interagency statement, a rulemaking, or a statute). This matters, as we’ll see below.

Third, criminal provisions are all binding. Board duties come in two flavors — binding rules and non-binding guidance. Inattentiveness to this distinction could lead to substantive changes in duties and pollute the streamlining. So “binding-force conservation” is a hard rule.

With those and other adaptations, I used Claude Code and Fable 5 to build BoardDistill — the StatuteDistill engine wrapped in a banking/finreg lexicon (named Lex-Balanus after barnacles by Fable 5).

The build started with guardrails, which are all deterministic and were tested (84 tests). Then came citation verification, which is hardest for duties buried in guidance docs and exam manuals. Every claim from BoardDistill is anchored to verified text where possible. Across the full run, it produced zero hallucinated citations.

The banking corpus also sharpened the transformation calculus, growing it from 14 moves to 16, as summarized in the screenshot below.

Most importantly, every substantive judgment — every merge, every retirement, every “parallel but justified” call — is gated and requires human approval, which is logged, creating a clear audit trail. The AI system processes and proposes, while humans decide.

In the end, the process converted 227 overlapping board duties across five agencies and five instrument types into 26 consolidated duties. The streamlining was done neutrally, with every statutorily mandated duty preserved. Readers can access a beautiful interactive visual here:

Interactive visual of the clean up

This spreadsheet contains the full crosswalk from today’s 227 board duties to the proposed consolidated set of 26:

The screenshot of the 26 consolidated duties above shows, for instance, that the same meta-duty regarding enterprise strategy and risk governance and risk appetite approval shows up in 26 separate places in the corpus, scattered across multiple sources. Conforming the mechanics of that and putting it in one place, logically situated and easily looked up, is one example of the power of sludge clean up.

Importantly, every statutorily mandated duty survives. For instance, each “parallel but justified” pairing, e.g., FDICIA versus SOX internal controls, is preserved rather than merged. Of the roughly 190 duties that disappear, nearly all are restatements or holding-company/bank mirror pairs. All policy dials — quantitative thresholds, cadences, the reconciled independence definition — are left unchanged and flagged for principals to decide.

Let’s assume for a moment that there’s broad consensus on the cleaned up end-state and that the agencies agree in principle to streamlining the 227 board duties into the 26.

Fable 5 estimated it would take at least 2.5 years to work through the bulk of the changes procedurally and likely 4 years to complete (with BHC/IDI rationalization and incentive comp being the wildcards).

When I first saw this, my heart sank. Having been part of many interagency policymaking projects, I had to admit that this was likely an under-estimate of the time required. Four years?

Maybe not.

As this project was unfolding, the Supreme Court decided Trump v. Slaughter, overruling the 90-year-old Humphrey’s Executor precedent and confirming the President’s authority to remove heads of independent agencies at will. (The companion Cook decision preserved a carve-out for the Federal Reserve.) The administrative law and democratic governance impacts of the Slaughter decision are hard to over-state and beyond the scope of this post.

But Slaughter does create an alternative implementation timeline.

Under the post-Slaughter framework, an Executive Order task force could direct an omnibus rulemaking, forcing coordination of key decisions on a 120-day calendar, and effectively compressing the 2.5 year timeframe to 1 year and the 4 year completion timeline to 2 years.

Much of the critical commentary on Slaughter has focused on the destruction of institutional norms, clipping of Congressional power, and heightened strongman risks. These need to be taken seriously and rigorously debated. From the perspective of policy sludge clean up, though, Slaughter offers a path to implementation that would have been highly challenging to pull off otherwise.

For policymakers and those interested in using AI to address long-standing challenges, this project3 holds several lessons.

AI capabilities can compound exponentially, but they need building blocks. I could not have skipped straight to building BoardDistill. The duty inventory came from the horserace, the verification harness from the LLM-as-judge work, and the consolidation method from reverse engineering RBI’s cleanup from my last post, as well as the calculus from the DC backtest this time. A seemingly random scattershot of earlier projects became critical building blocks for this one. (Anthropic’s development of Claude Code followed a similar path, with many components having been prototyped beforehand and serving as raw ingredients for later assembly.) Teams and leaders should not expect every prototype to generate immediate ROI. There is significant potential value in simply building things, even when the end state and utility aren’t clear, as many of the pieces that get built may end up supporting more ambitious projects in the future.

Cross-pollination and luck reinforce each other. The two key unlocks for this project came from outside banking: an open source legacy code tool presented at a conference, and a decade-old criminal justice reform project that my wife happened to work on. Neither is remotely close to bank regulation. Both were essential. I felt lucky to have had these grab my attention in close enough proximity for me to connect the dots. Upon reflection, it reminded me of a study by Richard Wiseman, where he asked self-identified “lucky” and “unlucky” people to count the number of photographs in a newspaper. The unlucky participants averaged 2 minutes for the task. The lucky participants took just seconds. Why? Because on the second page, there was a large ad declaring, “Stop counting - there are 43 photographs in this newspaper.” Lucky people seem to be predisposed to having a broader aperture, while unlucky people are heads down in tasks. Especially during times of great change, regulators would benefit from broadening their apertures a bit and looking for inspiration from other fields.

The value of human work will increasingly concentrate in ideation and review. AI powered the middle part of this project — the drafting, the cross-referencing, the verification. My value was concentrated at the two bookends: (A) deciding what questions to ask, what paths to explore, what challenges to tackle, and (B) reviewing the AI’s work and dispositioning each substantive judgment. Shifting my attention, time, and energy to these two tasks, especially review and dispositioning, took a lot of mental effort. Without AI, the mere process of doing a project provides a continuous stream of data points on the quality and direction of the work. With AI, those datapoints are calculated quickly and buried in the model’s traces. The risk of cognitive outsourcing is real and one must be disciplined in reorienting their time and attention to where it matters and makes the most difference.

Being much more ambitious: frontier AI raises the bar on what’s possible. A year ago, the most ambitious version of this work would have been a credible, hallucination-free analytical report. Today, one person working with AI can produce, in days, an end-to-end reform package — inventory, verification, a consolidated proposal, an A-to-B crosswalk, and a dual-track implementation roadmap. Regulators and policymakers need to look well beyond co-pilots and incremental workflow assistants, and push themselves to articulate extremely ambitious projects and objectives. I think most would be surprised to discover how many of those would be within striking distance of today’s frontier AI systems.

1

This post uses the term “duties” in the colloquial sense, roughly synonymous with “responsibilities,” “obligations,” or “tasks,” and not in the legal sense (“duty of loyalty" or “duty of care”).

2

The commission’s documentation, including all reports and advisory-group memos, remains publicly available at ccrc.dc.gov. The ending is a cautionary tale, though: the DC Council enacted the Revised Criminal Code Act in 2022, but Congress passed a joint resolution of disapproval, which President Biden signed in March 2023, nullifying the revision.

3

All of the code and artifacts for this project are available here.

No posts

Read the original on mhsu2112.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.