I debated two key opinion leaders last week. I will tell you topic later; it doesn’t matter for what I want to argue now.
In both debates, my academic colleagues did not focus on the evidence underpinning the procedure. They mentioned evidence, but in only a guideline-sort of way. As if to say, it’s in the guidelines, it’s settled.
In one of the debates, my opponent cited the vast number of procedures that have been done since approval. The message: Mandrola, the majority of your colleagues love this procedure. You are an outlier? Are you smarter than the masses?
In the other debate, my opponent said he just knew the device worked. The studies I cited were old, and my analysis was technical. “John, most doctors can’t see the details in studies like you can.”
I’ve been stewing about these attitudes for a few days. I’ve heard them before, especially the one about doctors not being able to appraise studies.
This is embarrassing for our field. Of course doctors should be able to assess evidence. We are in the evidence-assessment business. When a patient seeks our help, we use our evidence-assessment skills to sort out the problem and the fix.
We should expect nothing less when it comes to assessing the evidence underpinning our therapies.
But that’s not what I see out there—in the hospital, clinic and speaking circuit.
What I see is that doctors simply accept what guideline writers put into those colored boxes. Which would be fine if guidelines were written by neutral Martians free of dualities of interests. But that’s not the case.
Let me show you some examples of dubious evidence, which a practicing doctor ought to know about and be able to use in clinic.
Example 1: Left atrial appendage occlusion proponents should be able to explain their enthusiasm for a procedure that had two failed regulatory trials against warfarin. PROTECT AF failed to pass FDA muster, and PREVAIL failed to meet noninferiority in its co-primary endpoint for stroke, embolism and CV death. Slide below.
I’ve been told that this is too nuanced for doctors. I disagree.
What this shows is that the upper bound (worst case) is that the Watchman device could be 2x worse than warfarin. And the NI margin was a whopping 1.75—meaning the trial had such a low bar that Watchman could be 74% worse and still meet noninferiority. But it didn’t. It was even worse.
Example 2: In heart failure, sacubitril/valsartan (Entresto) is preferred over basic generic ACE-I. I contend that every cardiologist should be able to explain their love of this drug given two facts: one is that its regulatory trial (PARADIGM HF) used maximum dose valsartan against medium dose enalapril, and we know that dosage of this drug class matters. The second is that no other trial of sacubitril/valsartan has delivered significant results. Why not?
Example 2a: Also in heart failure, the new mineralocorticoid receptor antagonist (MRA) finerenone is being promoted as the next new thing in treating patients with heart failure and a preserved ejection fraction (HFpEF). The regulatory trial called FINEARTS HF was positive, but this trial compared the drug to placebo. Yet I think every cardiologist who uses this drug should explain why they think it is better than the generic MRA drug spironolactone. That a new drug for HFpEF gets on the market without proving superiority to a generic is not too technical for a cardiologist.
Example 3: In some centers, cath lab staff are not sent home until the last stress test is read as negative. That’s because nearly all positive stress tests get sent to the cath lab for angiography. The sky is blue and you get an angiogram for a positive stress test. Well, I don’t think it too nuanced for doctors to explain why this is the norm when our government spent $100 million on the ISCHEMIA trial, where 5000 patients with positive stress tests were randomized to immediate angiography or conservative medical management and there was no difference in hard outcomes.
Example 4: Primary care clinicians don’t get off the hook. I would expect them to be able to explain the NORDICC Trial of screening colonoscopy. In this NEJM-published study, more than 84,000 people were randomized to be invited to screening or not invited. In the end, the risk of dying of colon cancer was 0.28% in the invited group vs 0.31% in the non-invited group and risk of dying was 11.0% in both arms. I do not think it is too technical for a primary care clinician ordering a colonoscopy to understand that the difference in colon cancer death in this most modern trial is 0.03%. And there was no survival advantage. (The response should not be that those who actually had the screening did better; if this is your answer, you fail critical appraisal 101. Because people who get the screen are different from those who did not.)
I will stop now. But my point is that doctors (and advanced practice professionals) are about to face a crisis. That is, when a patient comes to you with an AI-driven thread of questions regarding the evidentiary support of a proposed treatment, the answer better not be…it’s in the guidelines, or, I know it works, or, 500,000 procedures have already been done.
What will doctors be for, if we cannot be better at evidence appraisal than a large language model?
I propose that we reinvigorate the skill of critical appraisal.
If we can learn the biochemistry and organic chemistry, we can surely understand that putting efficacy and safety endpoints into a non-inferiority primary endpoint is an easy way to bias a trial.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.