RSSAmplifier

Blog

Statistical Thinking

fharrell.comRSS feed ↗20 posts

Latest posts

Modernizing Clinical Trial Design and Analysis to Improve Efficiency & Flexibility

UCLA Cardiology Grand Rounds 2020-10-23 | Video (better video below) Vanderbilt University Department of Biostatistics 2020-11-18 Vanderbilt Translational Research Forum 2021-11-04 | Video Consilium Scientific 2024-03-14 | Video and here Seventh Annual Janice Pogue Lectureship in Biostatistics , Population Health Research Institute, Hamilton, Ontario, Canada 2024-12-05 National Heart and Lung…

Parametric Models vs. Empirical Distributions and Ordinal Regression

Background In modeling data we frequently need to estimate more than a mean. We may want to estimate quantiles, dispersion, or tail probabilities. These require model assumptions to be correct when the model is a smooth parametric one. Even when estimating a mean, for inference to be accurate we need distributional assumptions we make to be satisfied. The central limit theorem offers no protection…

Statistical Models Answer the Fundamental Clinical Question and Provide Clinical Trial Estimands

Goals and estimands should flow from the study design while respecting the lack of exchangeability of the participants and leading to estimators with data-verifiable assumptions. Statistical models are approximations of reality. It is not rational to fear transparent assumptions required by them. Avoiding models will almost surely lead to worse results. It is impossible to model an individual…

The Unifying Capabilities of Cumulative Probability Semiparametric Models

Disclaimer: No data were harmed during the making of this article. Any use of binning was strictly prohibited. Background Consider the oldest statistical regression model, the linear model, named from the fact that it is linear in the parameters . Suppose that the outcome variable is continuous, and conditional on a vector of baseline descriptors has a normal distribution. Let denote a vector of…

Implications of the Draft FDA Bayesian Guidance

Events AstraZenica Global Statistical Forum and Washington Statistical Society 2026-05-21 Slides Video

Thoughts About the Roles of AI for Statistics

Events Regression Modeling Strategies Course 2026 2026-05-19 Slides

Questions We Forget To Ask When Designing an RCT

Events JHU Trial Innovation Center DIDACT Symposium 2026-04-16 Slides

Causal by Design

Background Two research methods are being used with increasing frequency: causal inference and target trial emulation . There are many complex situations where a formal causal calculus such as that developed by Judea Pearl is needed to allow one to infer than an effect is caused by a specific variable such as the use of a treatment or exposure to a specific agent. Practitioners of causal calculus…

Goal-Driven Flexible Bayesian Design

Events Vanderbilt Department of Biostatistics Seminar, Nashville TN USA 2025-09-17 ACTStats 2025 Annual Meeting Keynote Talk , Nashville TN USA Video ; see also this Slides Details

Measures of Central Tendency for an Asymmetric Distribution, and Confidence Intervals

Measures of Central Tendency For symmetric normal-like distributions there is a clear winner for measuring central tendency: the sample mean. The mean has the highest precision/efficiency and is also representative of a typical observation from the population distribution. The mean is not robust, e.g., is too affected by extreme values, when the distribution is heavy-tailed or asymmetric. For…

Bootstrap Confidence Limits for Bootstrap Overfitting-Corrected Model Performance

Background The goal here is strong internal validation after fitting a pre-specified regression model or one that was derived using backwards step-down variable selection such that the same variable selection procedure can be repeated afresh for each bootstrap repetition. So strong internal validation means estimating a variety of model performance measures in a way that does not reward them for…

Minimal-Assumption Estimation of Survival Probability vs. a Continuous Variable

Background This article considers the following setting. Suppose we have one continuous predictor and an outcome variable and we wish to estimate a smooth, usually nonlinear, relationship between and some property of such as the mean or the probability that exceeds some specified value. When there is no censoring on , one can estimate such a smooth relationship nonparametrically using a standard…

Bayesian Thinking

Janice Pogue Lecture in Biostatistics , Department of Health Research Methods, Evidence, and Impact, McMaster University, Hamilton, Ontario, Canada 2024-12-06 Center for Biostatistics, Dept. of Population Health Science and Policy, Icahn School of Medicine at Mount Sinai, New York, 2025-03-18. Department of Biostatistics, Vanderbilt University School of Medicine, 2025-04-23 Slides

Statistical Computing Approaches to Maximum Likelihood Estimation

Overview Maximum likelihood estimation (MLE) is a gold standard estimation procedure in non-Bayesian statistics, and the likelihood function is central to Bayesian statistics (even though it is not maximized in the Bayesian paradigm). MLE may be unpenalized (the standard approach) or various penalty functions such as L1 ( lasso , absolute value penalty), and L2 ( ridge regression; quadratic)…

Ordinal State Transition Models as a Unifying Risk Prediction Framework

Event: International Chinese Statistical Association Applied Statistics Symposium , Nashville, Tennessee USA 2024-06-17 CANSSI Ontario STatistics Seminars (CAST) , Virtual, 2024-11-18 Slides

Adjudication and Statistical Efficiency

Background In clinical and epidemiologic studies one is frequently tasked with maximizing accuracy when assessing the presence of clinical conditions (symptoms, diagnoses, syndromes, etc.) or verifying outcome events such as stroke, myocardial infarction, or death from a specific cause. Prospective studies have the advantage of standardizing definitions of clinical conditions, minimizing bias, and…

The Burden of Demonstrating Statistical Validity of Clusters

Background Clustering of patients to find new “phenotypes” is now a fad. For example, repeating the false assertion that diabetes was ever a binary diagnosis , Ahlqvist et al claimed to have found 5 diabetes subtypes using a purely statistical analysis not driven by clinical knowledge. What they found is likely just inefficient prognostic stratification that could be improved upon by directly…

Hosting Web Content

One of my best decisions was to build my own web sites hbiostat.org and fharrell.com so that I have total control of content and formatting and can easily and quickly post content updates. I want to share a few things I’ve learned. While your organization’s web pages are great for static content, my public-facing content evolves rapidly with constant improvements made to course web pages,…

Tips for Biostatisticians Collaborating with Non-Biostatistician Medical Researchers

Slides

Rare Degenerative Diseases & Statistics:Methods for Analyzing Composite Patient Outcomes

Event: Consilium Scientific Slides Video