RSS Amplifier

Caleb’s Newsletter · Mar 15, 2026

AI Regulation, "Tempest in a Teapot?"

0
Sign in to vote or save

Caleb Gibbons, CFA, FRM · Caleb’s Newsletter

Note: Picture from the interweb, coffee pot, not teapot but the wave, “The Great Wave off Kanagawa” by Katsushika Hokusai (1831) was a great visual tempest for this post

No Nova Scotia statute or regulator (CUDIC, utility board, securities commission, insurance council) has issued AI/model risk guidelines analogous to OSFI’s E‑23. The province’s only related policy is its Digital Code of Practice (internal government policy) covering responsible use of AI in public services. Privacy protections in Nova Scotia (FOIPOP and PIIDPA) apply only to government-held data; provincially regulated financial firms must follow federal privacy law (PIPEDA/CPPA). Without a provincial AI framework, Nova Scotia institutions should voluntarily align with OSFI’s E‑23 principles.

Next Steps: Provincially regulated FIs in Nova Scotia should proactively adopt E‑23-style model risk governance. This includes creating a model inventory, defining risk categories, implementing lifecycle controls (design, validation, monitoring), and assigning clear board/senior management accountability, even absent a legal mandate. They should also follow federal privacy rules (CPPA/PIPEDA) for data used in AI. Engaging with industry groups and regulators to develop sectoral guidance is advisable. Until Nova Scotia or federal legislation specifically addresses AI risk (CPPA’s algorithmic fairness rules are forthcoming), the prudent approach is to treat OSFI E‑23 as best practice for model governance.

There are a few federally regulated (OSFI) credit unions in Canada, Coast Capital the largest with almost C$40bln in assets, but as a percentage <7% of Canada’s credit unions operate under a federal charter (50 of 750 million by assets). Provincially regulated credit unions in Nova Scotia offer deposit protection of $250,000 versus the $100,000 offered to OSFI regulated financial institutions covered under the Canadian Deposit Insurance Corporation (CDIC). Especially in Atlantic, the federal (OSFI) route (converting from provincial to federal regulation) has been deemed to be of limited strategic benefit versus the cost ($/years) and complexity. UNI Financial Cooperation is the largest and only federally regulated credit union in Atlantic Canada, based in New Brunswick. With $5.3bln in assets UNI is the largest Acadian francophone financial institution in Canada.

Why this matters. Tempest in a teapot is an idiom that describes a small issue being blown out of proportion and made to seem more significant than it is.

OSFI Guideline E‑23 (effective May 1, 2027) sets a principles-based, risk‑focused framework for all federally regulated financial institutions (FRFIs) to manage all models, including traditional actuarial models and AI/ML systems. It requires boards and senior management to oversee model risk, maintain a centralized inventory with risk ratings, and enforce robust lifecycle controls from design through decommissioning. Key requirements include formal governance and accountability (with clear roles and board reporting), rigorous data governance and lineage documentation, measures for explainability and bias/fairness in AI models, and ongoing monitoring for model drift and retraining needs. Third‑party (vendor) models must be validated and monitored to the same standard. Institutions must document every model and have escalation procedures for incidents and failures. An 18‑month transition period ends May 2027.

Compliance will require many FRFIs to upgrade resources and tools: for example, investing in model inventory platforms and hiring or outsourcing dedicated validation teams. Smaller institutions will need to scale their practices and may adopt technology solutions or third‑party validators for complex AI models. Internationally, E‑23’s principles echo global standards (e.g. U.S. Fed/OCC model‑risk guidance, UK AI‑governance consultations, and the EU AI Act’s high‑risk AI controls), but OSFI’s guideline explicitly names AI/ML and ties model risk to board responsibility.

Officer’s and Director’s insurance rates are not likely going down soon.

  • Applies to all FRFIs: Banks, insurers, trust companies, foreign branches, etc. (pension plans are excluded). All models – defined as any tool applying quantitative or AI/ML methods to process input data into outputs – fall in scope.

  • AI/ML included: “Emerging technologies such as AI and machine learning” are explicitly covered. Even simple AI applications (e.g. generative AI) must be assessed for model risk in context.

  • Risk‑based, principles approach: The guideline is not prescriptive. It encourages innovation but requires that higher‑risk models get stricter governance. Institutions determine model risk ratings based on complexity, materiality, autonomy, and data quality.

  • Board & Senior Management: Overall accountability stays with the institution. Senior management must establish policies, assign roles (owner, developer, reviewer, approver), and allocate skilled personnel, especially for novel AI techniques. Model risk (including from AI) must be reported clearly to the board. The board oversees model risk within the institution’s risk appetite.

  • MRM Framework: Institutions must build a formal Model Risk Management (MRM) framework aligned with enterprise risk policies. This includes: a model inventory, risk-rating methodology, lifecycle processes (development, validation, deployment, monitoring, decommissioning), reporting channels, and periodic reviews. The framework must be periodically reviewed and updated (e.g. as new AI tech emerges).

  • Resources & Team: MRM needs a multidisciplinary team (data scientists, risk experts, legal/ethics, etc.). Resources must be “commensurate with risks” – for example, firms using many advanced AI models need more sophisticated controls (OSFI notes “extensive use of advanced AI/ML techniques should have correspondingly mature governance”).

  • Accountability Clauses: Institutions cannot outsource responsibility. Even if models or components come from vendors, the institution is fully accountable for outcomes. (OSFI Guideline B‑10 on third‑party risk also applies.)

  • Comprehensive inventory: All models with non-negligible inherent risk must be logged in a central inventory. (Trivially low-risk tools need not be managed under MRM.) The inventory is an enterprise-level record.

  • Inventory content: At minimum, record each model’s ID, name, description/use case, owner/developer, origin (in-house or vendor), risk rating, and key dates (deployment, last review). For high‑risk models, also track version, dependencies, data sources, approved uses, limitations, monitoring status, and next review date. The inventory must be accurate, up-to-date, and controlled.

  • Risk Rating Methodology: Institutions must define a model risk-rating scheme. Ratings consider quantitative factors (e.g. portfolio size, potential financial impact) and qualitative factors (e.g. complexity, autonomy, customer or regulatory impact). Models are tiered (e.g. high/medium/low risk) so that the most critical or complex models get enhanced scrutiny. Ratings are reviewed regularly and updated after significant changes or performance issues. Firms may carve out a “negligible risk” category for trivial models, but this exemption must be carefully controlled and approved.

OSFI requires governance throughout the model lifecycle. Key lifecycle stages and controls include:

  • Model Design & Development: Models (including AI/ML) must have a clear business purpose and justification. Development standards include: comprehensive documentation (model logic, assumptions, limitations); defined performance criteria; traceable data sourcing; and version control. For AI models especially, data must be representative and pre-processed carefully to avoid embedding bias. Explainability requirements (e.g. transparency of inputs-output mapping) should be set based on the model’s use and impact. Any use of expert judgment must be documented. (Guideline encourages “expert judgment” checks during development.)

  • Independent Validation/Review: All higher-risk models require independent validation by parties separate from the developers. Validators assess conceptual soundness, data quality, methodology, and performance. The scope and frequency of validation are driven by risk rating. Events triggering a new review include: initial model build, major modifications (including self-learning changes), material performance deterioration, or new data sources. The validation must explicitly check explainability and fairness (e.g. “review novel algorithms for AI/ML” and confirm outputs are explainable). Findings are documented, and any weaknesses require remedial action or additional controls.

  • Model Approval: Before use, each model (and significant change) must be formally approved by a governance committee or senior function. Approval confirms the model is fit for purpose and that risks are acceptable or mitigated. A model can be approved even with known limitations, provided compensating controls are in place. OSFI expects approvals at key stages, e.g. post-validation and pre-deployment, and after periodic reviews.

  • Deployment and Change Control: Deploy models in a controlled environment. This includes thorough testing in a production-like setting, ensuring consistency between development and production data. Formal change-control procedures must be in place: clear roles, change-approval paths, and rollback plans. Related risks (cyber, tech infra) must be assessed before launch. For AI models, particular care is needed if models have many dependencies or use third-party components. Deployment plans should include communicating model outputs to stakeholders and updating documentation.

  • Ongoing Monitoring: Defined monitoring standards (frequency, metrics, thresholds) are required for all models. Monitoring checks both performance metrics (accuracy, etc.) and qualitative factors (ensuring the model remains used as intended). AI models need additional drift detection – tracking changes in input data or environment that degrade model accuracy. Institutions must set thresholds for revalidation or retirement and have contingency plans if a model fails (e.g. fallback models, ceasing use). Any significant issues (breaches, drift, unavailability) must trigger prompt escalation via pre-defined procedures.

  • Model Decommissioning: When a model is retired, a formal decommission plan is required. Stakeholders are notified, and the model (with documentation) is archived for a defined period as a fallback. Downstream impacts must be checked to prevent residual risks. If decommissioning a third-party model, review any additional contractual or technical steps needed.

  • Validation by Risk Tier: Models rated as higher risk demand more intensive, independent validation. Validators must be separate from developers; they can be internal specialists or external experts. The validation scope should test model logic, data, code and outcomes. E‑23 stresses “effective challenge”: reviews should uncover hidden flaws and ensure models meet objectives.

  • AI/ML Considerations: For “self‑learning” or continuously retrained models, OSFI requires criteria to determine when changes are material enough to trigger revalidation. Incremental updates may require periodic evaluation of performance drift and fairness. Institutions should predefine what constitutes a significant model change (e.g. architecture updates, new training data batches) to avoid unmonitored evolution.

  • Regulatory Alignment: OSFI’s use of “validation” is interchangeable with “review” in the guideline. The frequency of validation is risk-based (OSFI explicitly disavows mandatory annual validations). In fact, US regulators (OCC/Fed) similarly emphasize tailoring validation frequency to complexity and risk.

  • Data Quality Standards: OSFI mandates robust data governance for model inputs. Data must be accurate (free of material errors or biases) and representative of the model’s target population. Source data must comply with all legal/regulatory privacy and ethics rules. Data used for model training/validation should be traceable (lineage documented) so that transformations and sources are auditable.

  • Bias and Fairness: Institutions must explicitly consider and mitigate data bias. E‑23 highlights that biased training data can lead to “unfair model outputs” and reputational damage. Controls should include regular quality checks (outliers, missing values) and documented data cleaning procedures. The impact of synthetic or proxy data should be documented and monitored.

  • Explainability Support: Well-governed data also aids explainability. E‑23 notes that clear data lineage and documentation help end-users understand model outputs. Institutions should integrate model data governance with the enterprise’s broader data strategy (e.g. data dictionaries, master data management).

  • Proportional Explainability: E‑23 requires that each model’s level of explainability be assessed based on its use, autonomy, and impact. Models used in customer-facing or high-impact contexts need higher transparency (e.g. clear rationales for outputs). Firms must define what “explainable” means for each model and verify it during validation. During independent review, validators must check that outputs are sufficiently understandable to stakeholders.

  • Bias Testing: Beyond data bias controls (above), institutions should test models for biased or unfair outcomes. This might include checking that protected groups are not adversely affected by predictive models. Any bias concerns must be addressed by adjusting data, algorithms, or adding safeguards. OSFI’s guideline urges firms to “give consideration to the potential for unwanted bias” translating into unfair outputs.

  • Documentation: All explainability and bias analyses should be documented. If an AI model is essentially a “black box,” firms may need alternative controls (e.g. human overrides, limited scope) and must justify their adequacy. Notably, Basel (BCBS) has long warned that black-box models reduce confidence in results. OSFI’s text aligns with this view.

  • Drift Monitoring: For models that learn over time, banks must actively monitor for “model drift” (performance degradation due to changing data distributions). Monitoring should include statistical tests and triggers for retraining or validation whenever inputs or performance metrics change significantly.

  • Autonomous Changes: Self-learning AI models (e.g. with online learning) need guardrails. OSFI expects institutions to set internal thresholds indicating when model evolution constitutes a “material change” requiring re-approval. Real-time or frequent retraining frameworks should include audit trails and canary testing before full deployment.

  • Contingency Planning: Institutions must prepare for the possibility of rapid drift. The guideline explicitly requires contingency plans for a model’s unavailability or deterioration. This could mean parallel run strategies, fall-back models, or halting usage until a new model is validated. Documentation of drift events and responses (including revalidation results) is essential for auditability.

  • Third‑Party Risk (Guideline B-10): E‑23 mandates compliance with OSFI’s B‑10 guideline for all outsourced models. Vendor‑supplied models, even neutral network enabled Generative AI, must be subject to full model-risk management.

  • Due Diligence: Institutions should assess vendor models’ risks before use, including technology, data privacy, and concentration risk. Contracts should ensure access to model details (algorithms, training data) so the institution can validate and control them. OSFI emphasizes that banks are responsible for ensuring third‑party model use stays within risk appetite.

  • Validation & Monitoring: Third-party models must be validated and monitored just like in-house ones. OSFI specifically added that third-party models should receive validation and monitoring commensurate to their risk. For complex vendor models with proprietary tech, institutions need alternative controls (e.g. thorough end-to-end testing, output sensitivity analysis) since internal access to the model may be limited.

  • No “Grace Period”: There is no OSFI‑granted grace period for validating updates to vendor models after deployment. Firms may create limited-use exceptions but should aim to complete independent validation promptly.

  • Governance Structure: A governance framework must cover all third-party AI usage. OSFI expects each institution to decide how to govern its “black box” vendor models in line with its risk appetite. For layered cases (vendor uses another AI model), institutions should include feeder models in their review scope.

  • Risk Reporting: Regular reporting of model risk is required at all levels (team, management, board). This includes reports on inventory status, high-risk models, validation results, and monitoring outcomes. Board reporting must cover the organization’s overall model risk profile.

  • Escalation Protocols: OSFI requires clear escalation procedures for model issues. Any breach of performance thresholds or validation failures should be promptly escalated to appropriate management (and ultimately the board if systemic). Incident management should track issues, corrective actions, and communicate lessons learned.

  • Audit and Supervisory Oversight: Institutions should treat model governance audits like any risk audit. OSFI (and supervisors) will expect evidence of compliance: e.g., documentation of reviews, approvals, testing logs, incident logs. Firms should prepare to demonstrate their model governance processes during examinations.

  • Comprehensive Record keeping: Institutions must document every aspect of each model. Beyond the inventory, firms should retain: model development documentation, validation reports, code repositories, change logs, and user manuals.

  • Audit Trail: All model-related decisions (design choices, parameter settings, approval decisions) must be traceable. The guideline implies that model risk management activities, from development to retirement, should be audit-ready. For example, maintaining historical versions of models and data ensures reproducibility of past results.

  • Regulatory Audit: OSFI may conduct reviews of model governance during examinations. Having organized and accessible documentation will be critical to pass audits. E‑23 expects institutions to “provide evidence” of compliance, so clear records of policies, committees, and validation work are essential.

  • Publication and Effective Dates: OSFI published the final E-23 guideline on September 11, 2025. It takes effect on May 1, 2027, giving an 18‑month transition.

  • OSFI Guideline B-10 is currently in force and deals with risk from 3rd party technology providers (e.g. Microsoft 365) and AI model providers (Co-Pilot in the case of MSFT). OSFI clearly saw the import of AI specific guidance as we collectively increase our reliance on digital platforms and advanced analytics.

  • Basel Committee (BCBS): Basel’s model standards have long emphasized understanding and validating models. Notably, BCBS warned that a model “black box that is not well understood by bank personnel does not provide confidence”. OSFI’s focus on explainability and “effective challenge” aligns with this.

    Models should be “As simple as possible, but no more so.” Responsible AI must focus on the concepts of explainability and interpretability.

  • US Regulators (OCC/FRB): The US Fed/OCC guidance (SR 11-7) similarly defines models broadly and requires governance, validation, and risk-based policies. The recent OCC bulletin (2025) underlines that model-risk practices (including frequency of validation) should be scaled to the institution’s size, complexity and risk exposure, echoing E‑23’s proportionality principle.

  • UK (PRA/FCA): UK regulators advocate a technology-neutral, outcomes-based approach. Recent Bank of England/PRA discussion papers stress that oversight should focus on actual consumer/outcome risks (not on AI label) and be proportionate to each AI application’s impact. This resonates with E‑23’s risk-based, principle-driven stance, though OSFI sets specific lifecycle milestones.

  • EU AI Act: The upcoming EU AI Act classifies many financial AI systems as “high-risk,” imposing strict obligations (risk management, documentation, data governance, human oversight, etc.) on providers. E‑23’s requirements on documentation, bias mitigation, and oversight are broadly consistent with such high-risk controls. (Unlike the AI Act, OSFI applies to all models but implies higher scrutiny for critical ones.)

    The AI Act is a proposed European law on artificial intelligence (AI), the first law on AI by a major regulator anywhere. The law assigns applications of AI to three risk categories. First, applications and systems that create an unacceptable risk, such as government-run social scoring of the type used in China, are banned. Second, high-risk applications, such as a CV-scanning tool that ranks job applicants, are subject to specific legal requirements. Lastly, applications not explicitly banned or listed as high-risk are largely left unregulated.

  • Others (IAIS, BCBS 239): International bodies (IAIS, Basel Committee on data aggregation) similarly emphasize model/data integrity. For example, BCBS 239 on risk data calls for reliable model outputs, complementing E‑23’s data and explainability rules. The common theme worldwide is clear: institutions must treat AI models as key risk drivers and govern them rigorously.

  • China has been rapidly developing and implementing regulations related to AI, covering technology, cybersecurity, privacy, and intellectual property. One significant regulation is the Provisional Administrative Measures of Generative Artificial Intelligence Services (Generative AI Measures), which were published by the Cyberspace Administration of China (CAC) and took effect Aug. 15, 2023. These measures apply to the use of generative AI technology to provide services for generating text, pictures, sounds, videos, and other content within the territory of China. They impose various obligations on generative AI service providers, including the prohibition of generating illegal content, taking measures to prevent the generation of discriminatory content, and not infringing on others’ rights, including privacy rights and personal information rights.

    China has enacted several comprehensive laws aimed at protecting personal information, like the Personal Information Protection Law (PIPL) and the Internet Information Service Algorithmic Recommendation Management Provisions. These regulations mandate data minimization, user consent, and transparency in algorithm decision-making.

    Practical Implementation Considerations

  • Resourcing: Achieving E‑23 compliance will likely require hiring or training model-risk staff (quantitative validators, data scientists, AI ethics experts). Many firms are already augmenting teams for new analytics. Smaller institutions may need to outsource certain functions (e.g. complex model validation) to keep up.

  • Tooling: Maintaining an up-to-date model inventory at scale will require tools (e.g. MRM software). Some firms are implementing workflow and inventory platforms to track models across lines of business. Automated monitoring dashboards and data lineage tools can ease ongoing governance.

  • Policies and Governance: Firms must update MRM policies to map roles explicitly to E‑23 roles (owner/reviewer/approver) and ensure independence where needed. Separate committees may be needed for high‑risk AI models. Policies should codify what constitutes a model change and how to document it, per E‑23.

  • Validation Teams: Given the volume of models (including spreadsheets and apps that now count as models), dedicated validation resources will be strained. Institutions should prioritize the highest‑risk models and consider sampling or statistical validation methods for lower‑risk ones. Scenario analysis and stress testing (like those under Basel) may complement validation for certain model classes.

  • Training: Business users and IT staff must be trained on new model governance processes. Model developers should document everything clearly. Executives and board members may need briefing on AI/model risk so they can oversee reporting.

Requirement (OSFI E-23)

Recommended Institutional Actions Governance & Accountability: Board oversight; clear roles; risk-aligned policies. Establish a Model Risk Committee or integrate into existing risk committees; define owner/reviewer roles; update governance charter; schedule regular reports to senior management/board.

Model Inventory: Enterprise inventory of all non-negligible models. Create or enhance a centralized inventory system; catalog every model (ID, use, owner, risk rating); update inventory processes (use automated discovery tools if needed).

Risk Classification: Risk-rating framework (quantitative & qualitative). Develop and document risk-rating criteria (e.g. impact, complexity, data issues); apply ratings to all models; review ratings periodically and on model changes.

Documentation: Full documentation for model design, data, assumptions. Standardize development documentation templates; include model purpose, logic, data sources, assumptions, limitations; ensure documentation is stored securely and linked in inventory.

Independent Validation: Risk-based validation of models. Build or strengthen a validation unit independent from development; schedule validations based on model risk (e.g. quarterly for high-risk models); maintain review checklists.

Model Deployment: Change control and pre-launch testing. Define deployment procedures (test plans, sign-off); implement IT change-control and versioning; require deployment approvals; test models end-to-end in production environment.

Monitoring (Drift): Ongoing performance checks & drift thresholds. Set up automated monitoring (performance metrics vs. benchmarks); define alert thresholds; create workflow to investigate anomalies; plan triggers for retraining or remediation.

Data Governance: Fit-for-purpose data, lineage, quality checks. Enforce data quality controls (profiling, cleansing); map data lineage for key models; integrate model data governance with enterprise data strategy; document synthetic data use.

Explainability & Fairness: Ensure outputs are understandable; test for bias. For each model, define required transparency (e.g. feature importance output, decision rules); perform bias/fairness testing (e.g. subgroup analysis); document rationale of outputs.

Third-Party Models: Due diligence and oversight per B-10. Vendor contracts: require model info access; include audit/validation clauses; apply model review even if vendor-supplied; track third-party models in inventory; align with B-10 processes.

Reporting & Escalation: Incident logs; board reporting. Establish incident management for model failures; ensure risk reports on models to risk committees; maintain audit logs of all governance actions; escalate major events to executives.

Retention & Auditability: Archival of models/docs; audit trails. Retain historical model versions, data snapshots, documentation per records policy; perform periodic audits of the MRM process; prepare for regulator reviews.

Each action item should be incorporated into the institution’s model risk policies and procedures in advance of May 2027. OSFI expects clear evidence (e.g. right down to meeting minutes, reports) that these practices are in place.

ISO/IEC 42001 is the first international standard specifically for managing artificial intelligence systems.

It was published in December 2023 by International Organization for Standardization.

ISO 42001 establishes a management system for Artificial Intelligence (AI) similar to how:

  • ISO 27001 governs information security

  • ISO 9001 governs quality management

The official name:

ISO 42001: Artificial Intelligence Management System (AIMS)

The standard helps organizations:

  • govern AI responsibly

  • manage risks from AI systems

  • ensure transparency and accountability

  • implement ethical AI practices

It applies to organizations that:

  • develop AI

  • deploy AI

  • use AI in operations

Organizations must create:

  • AI oversight policies

  • accountability structures

  • internal controls

Requires identifying risks such as:

  • algorithmic bias

  • data privacy violations

  • unsafe automation

  • unintended outcomes

Organizations must document:

  • how AI systems work

  • data sources

  • decision logic

  • limitations

AI systems must allow:

  • human review

  • intervention

  • auditability

Closing comments;

AI is an international technology, from its sources of data and talent to its implementation and use. It is therefore not unreasonable to expect some movement toward international regulation or at least agreed best practices. Organizations that adopt ISO 42001 demonstrate responsible AI Governance.

This topic can seem overwhelming for certain, but it is important we “get it right”. A key guardrail for the entire exercise is “humans in the loop”. The current state of play is fraught with risk. Generative AI (GAI) LLMs (large language models) currently often present erroneous or false context in response to prompts.

“Confabulation” refers to a phenomenon in which GAI systems generate and confidently present erroneous or false content in response to prompts. These phenomena are colloquially also referred to as “hallucinations” or “fabrications.” MIT researchers have mathematically proven that ChatGPT’s built-in sycophancy creates a phenomenon called “delusional spiraling” (you ask it something it agrees. You ask again and it more vehemently agrees until one believes things that are false).

Note. This is study preparation for my April 2026 examination for the Risk and AI (RAI) certification. Wish me luck! GARP's RAI cert

This AI craze has certainly made it interesting to choose a career path for younger generations! It would appear that “number” people are on the front line, right behind software engineers that now spend their time scrubbing data and prompting chatbots to ring further efficiency (model training), to the point where adding more data provides a less optimal result (over-fitting). The “word” people seem safe for a bit, given the slop many models currently produce.

JCG

No posts

Read the original on calebgibbons.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.