The above is the title of a recent a newspaper op-ed (pdf here) by columnist David French. He’s asking a good question! My quick take is that there are four things going on: 1. People want improvement. If you’re at … Continue reading →
With some sustained effort, I cleaned out my entire inbox! Now I’m ready to do research again. Actually, cleaning out the inbox entailed doing a lot of research and writing, lots of interesting things I’d put off for a while. … Continue reading →
In a distressing update to an already distressing story, James Smoliga writes: Q-collar is ramping up its fall sports campaign. Lots of buzz on Facebook, with misleading comments from the company, new ways of spinning the FDA disclaimers, and plenty … Continue reading →
This came in the mail: My name is ** and I am a junior at ** High School. I am writing to you as a part of a summer assignment for my upcoming research class where I had to choose … Continue reading →
Mitzi Morris sends along the paper trail (I love archaic idioms) for her data spelunking in advance of her all-day introductory Stan tutorial next week at StanCon (we’ll see you in Uppsala). She was trying to reproduce the results from … Continue reading →
The Q collar, I hope you don’t remember, is a neck accessory that is advertised as “proven” to protect athletes’ brains, but the studies purporting to offer this proof are suspect. This is an interesting example of science-adjacent incompetence or … Continue reading →
OK, this one’s time constrained so I’ll post it right away, not on the usual lag. Christian Robert points to this conference announcement: We invite proposals for the BayesComp 2027 mirror event in Aussois, France, that will take place on … Continue reading →
I was in the Playroom the other day and noticed this book on the shelf, The Makers of Public Policy, by R. J. Monsen and M. W. Cannon, copyright 1965. Here’s the table of contents: It’s interesting to see here … Continue reading →
Last week Andrew commented that we need a more transparent workflow for survey statistics. So I looked in the new Bayesian Workflow book: Chapter 19 “Building up to a hierarchical model: Coronavirus testing” is a case study about a 2020 … Continue reading →
Regular readers of this blog will trace my short career as a Freud expert to a post from 2012, Economics now = Freudian psychology in the 1950s: More on the incoherence of “economics exceptionalism”. But recently I thought we need … Continue reading →
This post offers two recommendations and a story. First, the recommendations. • Here’s the book referred to in the title of the post. I highly recommend it! My only regret is that we didn’t call it Applied Regression and Causal … Continue reading →
After quoting Donald Trump at a political rally with Robert Kennedy Jr., That is why today I am repeating my pledge to establish a panel of top experts, working with Bobby, to investigate what is causing the decades-long increase in … Continue reading →
Picking up a thread from awhile ago, another item from my review of that Daniel Davies book: Davies writes, “Denial, when you are not part of it, is actually a terrifying thing. One watches one’s fellow humans doing things that … Continue reading →
Parthsarthi Joshi writes: There is a stark problem that I have observed with how data science is taught in most books and tutorials – the concepts are taught on an individual level but fail to give a complete idea. The … Continue reading →
Jonathan “No Trump” Falk writes: Your post this week on Integration and Differentiation got me thinking, never a good sign. Adding to this, I only really learned in January what Attention means, the foundation of LLM. And what it means … Continue reading →
Maciej Cegłowski writes: Unlike the Moon, which hangs in the sky like a lonely grandparent waiting for someone to visit, Mars leads a rich orbital life of its own and is not always around to entertain the itinerant astronaut. There … Continue reading →
Last week, Raphael K shared a concern: adjusting for lots of variables can lead to very large weights. So today let’s dive into Si et al. 2020, who saw this in constructing survey weights for the NYC Longitudinal Study of … Continue reading →
We are happy to announce the official release of Walnutpie version 0.0.1. Walnutpie is an MCMC sampler for continuously differentiable densities coded in Python, accepting models coded in Stan, PyMC, NumPyro, JAX, and plain old Python. Walnutpie is not an … Continue reading →
Paul Alper writes: Regarding your blog of today, believe it or not, I knew Marshall Brickman pretty well before he became famous. He was an undergrad at the University of Wisconsin in Madison when I was a grad student there. … Continue reading →
In our bathroom we have this book, “And here’s the kicker: Conversations with 21 top humor writers on their craft,” edited by Mike Sacks. It’s well suited for the throne, as you can dip in and read as little or … Continue reading →
Miha Gazvoda shares this post with the above title and the subtitle, “Using Bayesian multilevel models to correct bias and calibrate uncertainty.” He’s using the chickens model from our Slamming the Sham paper in the more general setting of placebo … Continue reading →
I was thinking about the above saying in the context of bad regression discontinuity analyses. Statistical methods can be characterized in terms of how they can go wrong. Some common modes of “how can things go wrong” include: – biased … Continue reading →
In reaction to my article with Andy King proposing post-publication review, Dan “Fast and Frugal” Goldstein writes: Your process limits information search, computation, and time so it seems fast and frugal to me. Happy you still associate me with that … Continue reading →
Gaurav Sood writes: I was reading ‘The Bestseller Code.’ The book reports results from some regressions of the form: bestseller or not ~ features of content This got me thinking about what you can recover from such an exercise. Say … Continue reading →
When I first started working on posterior predictive checking back in 1988, it was as a device for determining equivalent degrees of freedom for a chi-squared test for a model with constrained parameters–in that case, positivity restrictions in an image … Continue reading →
Alex Malinowski has a question about evaluating the calibration of interval forecasts: We publish interval forecasts under a pre-registration scheme: each forecast is serialised, hashed and timestamped into a Bitcoin block before publication, so the stated interval cannot be adjusted … Continue reading →
Following up on Richard Yates (see last year’s post, Double Feature: Revolutionary Road and That Darned Chatbot), I came across his collection of short stories from 1962, Eleven Kinds of Loneliness. These stories are wonderful and deserve all the praise … Continue reading →
Last month we saw that the Times/Siena Poll is now using energy balancing weights (Huling & Mak, 2024). In a toy example, we saw under which outcome models these weighting methods might do well. I was inspired by Little 2004, who … Continue reading →
You know how we talk about the two modes of microeconomic reasoning? For example, here: The logic of social science can work in two directions: generative modeling predicts behavior given assumed preferences, and inferential reasoning deduces preferences given observed behavior. … Continue reading →
Scott Cunningham writes: You’ll appreciate this I think. I ran Claude code on 96 specs for a popular difference-in-difference estimator with the identical specifications, ranging covariates only, for three languages (R, Python and Stata) and 2 packages for each. Different … Continue reading →
At the end of the second world war, Jews were a bit over 3% of the U.S. population, voted at a high rate, and were concentrated in the swing state of New York. Jews had two big issues–Israel and political … Continue reading →
Someone pointed to this post from last year, “If only Arxiv required researchers to sign at the top rather than the bottom of the page, none of this would’ve happened,” and asked about this statement of mine: “Seriously, though, setting … Continue reading →
Recently in the sister blog: Generic language, that is, language that refers to a category as an abstract whole (e.g., ‘Girls like pink’) rather than specific individuals (e.g., ‘This girl likes pink’), is a common means by which children learn … Continue reading →
In reaction to my recent post, If Cuomo had been able to run against Mamdani head-to-head, would he have won?, sociologist Kieran Healy posted a pair of maps showing precinct-level results from the recent New York mayoral election. One of … Continue reading →
OK, this is absolutely horrifying: The ill effects of ingested lead and other heavy metals had been known since the 1920s, when employees at TEL [tetraethyl lead]-refining plants began hallucinating butterflies and going into convulsions of violent insanity (at least … Continue reading →
Poststratification uses population data on X to estimate E(Y) via E(E(Y | X, R = 1)), where R = 1 are survey respondents who provide Y and X. When the inner expectation “E” is estimated via Multilevel Regression, this is … Continue reading →
Joshua Brooks writes: I know you’ve posted on the topic more generally but don’t recall if you’ve discussed this in particular. Given the timing in relation to cuts in food assistance, It seems a particularly egregious example of the politicization … Continue reading →
In a post entitled, “You’re Writing a Book. So Stop Writing a Movie,” Rebecca Makkai writes: You want to set your movie in a futuristic New York where every building has a flying car port on top and there are … Continue reading →
This is quite possibly the stupidest thing in the Epstein files. Lord knows there’s lots of competition from the likes of Soon-Yi “Woody” Allen, Larry “Lawrence” Summers, Nathan “Clippy” Myhrvold, and Columbia’s own Richard “Axel” Foley, but I think this … Continue reading →
Tom Ferguson came across this news article, Who Really Has the 2026 Midterms Cash Edge?, and was disappointed to see this completely wrong graph: The problem here is not the inclusion of no-longer-candidate Platner, as that’s noted in a footnote. … Continue reading →
I came across this webpage by Sheeva Azma entitled, “Here’s every scientist I have found in the Epstein Files so far.” She’s missing a few big fish: Dan Ariely (professor at MIT and Duke, Ted talk star, and teller of … Continue reading →
I recently learned from a blog comment that Herman Chernoff passed away last week at the age of 103. He was born the same year as my dad. I first met Chernoff–it’s not like he was a particularly formal guy, … Continue reading →
Roughly speaking, Bayesian Workflow is to Bayesian Data Analysis in 2026 what Bayesian Data Analysis was to earlier Bayesian books in 1995: it builds upon everything that came before. With Bayesian Data Analysis, the big steps forward were: Going beyond … Continue reading →
Evan Rosenman writes: The implosion of Graham Platner’s Senate campaign in Maine has upended a marquee Senate race, leaving the state Democratic party just a few weeks to choose a substitute nominee. A planned nominating convention on July 25th has … Continue reading →
We’ve talked about uncertainty in polls (see Margin of Error, Total Margin of Error, Total Margin of Error II) and we’ve talked about ranked data (see exploded logit !). A new paper, Rosenman & Liang 2026, looks at uncertainty in … Continue reading →
I took a look at the above-titled book by economists Duncan Foley and Ellis Scharfenaker. It’s an interesting read, in many ways a throwback to the 1950s when a group of mathematicians brewed a heady mix of operations research, game … Continue reading →
The following came in the email the other day: I’m reaching out to introduce the Voter Impact Index, a new data tool from PowerMoves that assigns every U.S. zip code a voter impact score based on the recent competitiveness of … Continue reading →
Apropos of our recent discussion on the estimation of historical population sizes, Sean Manning writes: Some archaeologists have measured house sizes for Gini-coefficient-style studies aside from studying human remains to measure nutrition and rates of illness. I think that was … Continue reading →
In an abstract entitled, “Statistical dust and sweeping claims about maternal warmth,” John Richters and Everett Waters write: Alley and colleagues draw on mediation analyses of longitudinal data from Millennium Cohort Study to argue that their findings “highlight the critically … Continue reading →