Every day, a new medical AI tool shows up.
It promises to predict readmissions, flag silent hypoxia, summarize patient notes, or help clinicians triage faster. But how do you know whether the tool was trained fairly? That it performs equitably across racial groups or age brackets? That it won’t quietly drift into irrelevance after a few model updates?
The truth is: most users don’t get a clear answer. Even the vendors building these models often don’t have the internal capacity to apply full ethics frameworks. Worse, the few that do exist often require specialized staff, dedicated committees, or regulatory teams most startups simply don’t have.
That’s the gap the SAFE-AI Framework was built to close. We at MTN are very excited to have played a part.
SAFE-AI stands for Scalable Agile Framework for Execution in AI. It was designed through a collaborative effort between MTN, the Data Science Alliance, the University of Utah, and Nemsee LLC. The goal is highly practical: embed lightweight but testable ethical oversight into the way small vendors and product teams already work.
Built to fit real Agile/Scrum workflows (no ethics bureaucracy required)
Designed for resource-limited teams who require hospital-grade in transparency and fairness
Aligns with emerging federal rules (HHS §1557, FDA SaMD)
Makes it easy for hospitals to ask vendors the right questions—and get clear, auditable answers
A preprint of the paper is available through this link: http://arxiv.org/abs/2507.01304
Hospitals are already integrating AI-powered tools into core operations. But most of these tools are:
Developed by small-to-mid-sized teams
Trained on siloed or non-representative datasets
Tuned and re-tuned rapidly in live environments
These facts create risk. Not just regulatory or reputational risk, but real clinical and operational risks: alarm fatigue, performance drift, and unequal outcomes for vulnerable patients.
SAFE-AI offers a way to make sure every model change, deployment, and update includes an ethics checkpoint—without slowing product delivery.
SAFE-AI is not a checklist, it’s a repeating lifecycle that maps directly to Agile build cycles.
Figure 1. SAFE-AI Summary Workflow - highlights each core phase and feedback loop, from prioritization to deployment and monitoring.
Select and prioritize among potential projects and their degree of alignment with organizational interests. Identify compliance and regulatory issues, affected stakeholders, and match the level of ethical scrutiny to the real-world impact.
Define Acceptance Criteria (e.g. sensitivity, latency), Fairness Metrics (subgroup performance), and Transparency Metrics using the SPAMM approach (more on that below).
Build, tune, test tolerance metrics. Most importantly, embed those metrics directly into product backlogs. Everyone from data scientists to QA owns ethics.
Set re-entry rules: every model retrain, data update, or environment change triggers another lightweight ethics cycle. Think of it as “post-market surveillance” for your AI signals.
→ For tech teams: Implementation is modular and audit-friendly.
Figure 2. Detailed Ethical Evaluation Process SAFE-AI - conducted continuously throughout product development. The process begins with a discussion of priorities and alignment and proceeds through model building and inference implementation, emphasizing appropriate tolerance levels at each iteration.
SAFE-AI helps vendors (and hospitals) move beyond vague principles by using three metric categories:
🟢 Acceptance Criteria: Stakeholder-defined performance thresholds—like “AUC ≥ 0.85 on under-40 patients.”
🟡 Fairness Metrics: Gap analyses across subgroups, like “False-negative rate for Black patients should not exceed overall FNR by >3%.”
🔵 Transparency Metrics: Narrative scenario testing using SPAMM (Scenario-Based Probability Analogy Mapping). Instead of abstract confidence scores, you get human-readable summaries such as:
“In 70% of patients with these vital trends, respiratory failure occurred within 6 hours. In the other 30%, the model missed early signs, mostly in those with comorbidities.”
One of the paper’s key innovations is the SPAMM technique, which makes it easier for clinicians, executives, and even patients to understand what an AI model is doing, and where it might fail.
This matters because tools like SHAP or LIME (popular explainability techniques) don’t always translate meaningfully to frontline decision-makers. SAFE-AI emphasizes narrative transparency, telling the story of how the model behaves across contexts, including error modes.
Hospitals can demand this kind of scenario framing as part of every AI procurement or evaluation process.
Many ethics approaches assume a waterfall-style dev cycle and dedicated review boards. That’s fine for pharma, but it doesn’t match the pace or constraints of AI product teams.
SAFE-AI maps ethics tasks onto existing tools like Jira, GitHub Projects, and sprint planning boards. That means:
No separate documentation silos
No ethics teams working in isolation
No extra quarterly committee meetings
Instead, you get “responsibility metrics” that your vendors can track just like latency or uptime—so you can track them too.
Every time a model is retrained—even if the inputs/outputs don’t change—the SAFE-AI framework requires a re-check of fairness and transparency metrics. That ensures drift doesn’t creep in silently.
This also future-proofs your health system. As AI regulations evolve (including under HHS, ONC, and the AI Bill of Rights), you’ll already have an audit trail showing due diligence and governance maturity.
If you operate in a regulated environment and are considering AI-powered tools or already deploying them at the edge, in the EHR, or in predictive analytics stacks, SAFE-AI provides a simple question:
“Do you have an Agile-compatible ethics process, with defined fairness and transparency metrics, and do you revisit them after every model update?”
If the answer is “no,” SAFE-AI can make it “yes.”
We’re currently piloting SAFE-AI with several systems and partners. More on that to come!
Preprint is available at: http://arxiv.org/abs/2507.01304
This post is public so feel free to share it.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.