RSSAmplifier

Blog

Lucy D'Agostino McGowan

Recent content on Lucy D'Agostino McGowan

/RSS feed ↗211 posts

Latest posts

Symmetric vaccine efficacy

Traditional measures of vaccine efficacy (VE) are inherently asymmetric, constrained above by 1 but unbounded below. As a result, VE estimates and corresponding confidence intervals (CIs) can extend far below zero, making interpretation difficult and potentially obscuring whether the apparent effect reflects true harm or simply statistical uncertainty. The proposed symmetric vaccine efficacy (SVE)…

Symmetric Vaccine Efficacy: Interpretable Estimation and Inference for Vaccine Trials

Using large language models to enhance clinically-driven missing data recovery algorithms in electronic health records

Objectives Electronic health record (EHR) data are prone to missingness and errors. Previously, we devised an enriched chart review protocol where a “roadmap” of auxiliary diagnoses was used to recover missing values. Still, chart reviews are expensive and time-intensive, limiting the number of patients whose data can be reviewed. Now, we investigate the accuracy and scalability of a…

Mind the Gap: Causal Inference is Not Just a Statistics Problem

Causal Inference in R

Causal Inference in R

Miss(ing) Congeniality: The Cost of Incompatible Imputation Models

The Role of Congeniality in Multiple Imputation for Doubly Robust Causal Estimation

Causal Inference in R

Causal Inference in R

The Role of Congeniality in Multiple Imputation for Doubly Robust Causal Estimation

The Art of Data Refinement: Severance Analyses

Exploring the Potential of Large Language Models in Generating Saturated DAGs for Causal Inference

Causal Inference Is Not Just A Statistics Problem

Causal Inference in R

Causal Inference in R

Immaculate vibes, questionable code

Building Strong ASA Student Chapters: Strategies for Growth, Inclusivity, and Maximizing Value

Understanding Statistics in Medical Literature

Understanding Statistics in Medical Literature

Partitioned Local Depth (PaLD) Community Analyses in R

Partitioned Local Depth (PaLD) is a framework for holistic consideration of community structure for distance-based data. This paper describes an R package, pald, for calculating Partitioned Local Depth (PaLD) probabilities, implementing community analyses, determining community clusters, and creating data visualizations to display community structure. We present essentials of the PaLD approach,…

Quantifying the Alignment of a Data Analysis between Analyst and Audience

Using Mathlink Cubes to Introduce Data Wrangling with Examples in R

This article explores an innovative approach to teaching data wrangling skills to students through hands-on activities before transitioning to coding. Data wrangling, a critical aspect of data analysis, involves cleaning, transforming, and restructuring data. We introduce the use of a physical tool, mathlink cubes, to facilitate a tangible understanding of datasets. This approach helps students…

One Simple Way to Get Better at Reading Data

The Case for Deterministic Imputation in Predictive Modeling

Combining Straight-Line and Map-Based Distances to Investigate the Connection Between Proximity to Healthy Foods and Disease

Healthy foods are essential for a healthy life, but accessing healthy food can be more challenging for some people than others. This disparity in food access may lead to disparities in well-being, potentially with disproportionate rates of diseases in communities that face more challenges in accessing healthy food (i.e., low-access communities). Identifying low-access, high-risk communities for…

The Why Behind Including Y in your Imputation Model

Untangling Causal Effects: Understanding the Limits of Statistics

Data Jamboree: A Party of Open-Source Software Solving Real-World Data Science Problems

The evolving focus in statistics and data science education highlights the growing importance of computing. This paper presents the Data Jamboree, a live event that combines computational methods with traditional statistical techniques to address real-world data science problems. Participants, ranging from novices to experienced users, followed workshop leaders in using open-source tools like…

It's ME hi, I'm the collider it's ME

Including the outcome in your imputation model -- why isn't this 'double dipping'?

Power and sample size calculations for testing the ratio of reproductive values in phylogenetic samples

The Art of the Invite: Crafting Successful Invited Session Proposals

Evaluating the Alignment of a Data Analysis between Analyst and Audience

Why You Must Include the Outcome in Your Imputation Model (and Why It's Not Double Dipping)

Causal Inference in R

Causal Inference in R

Statistical Rigor in Academic Medicine: Practical Strategies for Improvement

A Framework for Developing AI-powered Negotiation Agents to Examine Producer-Consumer Dynamics in Data Science

Partnering with Authors to Enhance Reproducibility at JASA

When to Include the Outcome in Your Imputation Model: A Mathematical Demonstration and Practical Advice

Bridging the Gap Between Theory and Practice: When to Include the Outcome in Your Imputation Model

The 'Why' behind including 'Y' in your imputation model

Missing data is a common challenge when analyzing epidemiological data, and imputation is often used to address this issue. Here, we investigate the scenario where a covariate used in an analysis has missingness and will be imputed. There are recommendations to include the outcome from the analysis model in the imputation model for missing covariates, but it is not necessarily clear if this…

Power and sample size calculations for testing the ratio of reproductive values in phylogenetic samples

The quality of the inferences we make from pathogen sequence data is determined by the number and composition of pathogen sequences that make up the sample used to drive that inference. However, there remains limited guidance on how to best structure and power studies when the end goal is phylogenetic inference. One question that we can attempt to answer with molecular data is whether some people…

Bridging the Gap Between Imputation Theory and Practice

Causal Inference is Not Just a Statistics Problem

Integrating Design Thinking in the Data Analytic Process

The Study of the Epidemiology of Pediatric Hypertension Registry (SUPERHERO): Rationale and Methods

Analytic Design Theory: Framework for Alignment Between Analyst and Audience

Causal Inference is not just a statistics problem

This paper introduces a collection of four data sets, similar to Anscombe’s Quartet, that aim to highlight the challenges involved when estimating causal effects. Each of the four data sets is generated based on a distinct causal mechanism: the first involves a collider, the second involves a confounder, the third involves a mediator, and the fourth involves the induction of M-Bias by an…