One morning session on optimal transport after first-hand witnessing the impressive ballet of orderly lines entering the subway at Nagoya Station (and a very early run along the river and a high humidity rate, hence the picture of empty street at 5am). With Hugo Lavenant exhibiting optimal rates for Bayesian nonparametrics using Wasserstein distances, Pierre Jacob coupling MCMC chains, and Anya Katsevich investigating non-Gaussian asymptotic distributions in high dimensions (beyond Bernstein-von Mises). Then I tried to attend the session on Bayesian Uncertainty Quantification and Posterior Sampling for Large-Scale Generative Models, but it proved too popular for the number of seats, and I ended up discussing with others. After a nap related to my jetlag induced, early, rise I went back to chair Sid Chib’s Foundation Lecture, where he discussed the use of an orbit of models in Bayesian model choice, rather than the (MAP) most likely one. Based on their 2018 JASA paper which I already discussed in Paris with Anna Simoni presenting. Hence reminding me of points I presumably already made, from the issue of having too many models to realistically explore to constructing coherent priors across them, with the fractional, empirical, proxy of using a (same) fraction of the sample as a learning sample, to more philosophical issues like missing a utility function about having to chose a model, especially with all models being wrong, missing an uncertainty quantification on the evidence itself, rather than using most likely models (MAP!), called an orbit by Sid (which requires some calibration). Since the uncertainty represented by the sample induces an uncertainty in the ranking of models. And a most appropriate, almost local, occurrence of the Rashomon principle!!! And I finished the day mixing with many friends in the poster session¹, where Darren Wraith presented our ongoing work on novel, adaptive, importance, sampling.
Archive for evidence
ISBA 2026²
Posted in Books, pictures, Running, Statistics, Travel, University life with tags Bayesian model averaging, Bayesian model choice, Bernstein-von Mises theorem, Chib's approximation, evidence, ISBA 2026, ISBA World Meeting, Japan, map, model uncertainty, Nagoya, Rashomon, Sid Chib, Wasserstein distance on July 1, 2026 by xi'aninverse probability weighting
Posted in Books, pictures, Running, Statistics, Travel, University life with tags Boston, Charles river, Charles Stein, Chris Sims, Comptes Rendus de l'Académie des Sciences, evidence, Harvard University, Horvitz-Thompson estimator, inverse probability, James-Stein estimator, Jaroslav Hájek, Larry Wasserman, New England, New England Journal of Statistics in Data Science, noise-contrastive estimation, Riemann sums, Robins-Wasserman paradox, self-normalised importance sampling, sunrise, survey sampling on May 4, 2026 by xi'an
Quite recently, Jyotishka Datta and Nick Polson published a fairly interesting [imho] paper in The New England Journal of Statistics in Data Science, entitled Inverse Probability Weighting: From Survey Sampling to Evidence Estimation that (obviously) caters to my own interests! They bring three threads together. First, they recall the long debate between using [normalized] Horvitz–Thompson and [self-normalized] Hájek estimators in survey sampling, pointing out that the latter is “usually the better estimator, despite estimation of an a priori known quantity” (citing from Särndal & al., 2003). Which is also my experience with importance sampling, as in this 1995 Note aux Comptes Rendus with George. The mathematical paradox of “estimating” a constant is central to other advances in the area, like noise-contrastive estimation à la Gutmann & Hyvärinen (2005) or the measure estimation of Kong & al. (2003). Even more interestingly, Datta & Polson consider there is a link with the inconsistent Bayesian (counter)example of Larry Wasserman and Jamie Robins, where the censoring probability increases with the value of the parameter of interest, paradox in which Chris Sims also got involved. (I was unaware that he had passed away last month.) And, lo and behold!, with the Stein “paradox” of my PhD years (and beyond).
Their central argument stands with the missing data link between Horvitz-Thompson survey sampling and Monte Carlo integration. (With a reference to our Riemann sum papers with Anne Philippe!, making me realise the authors had recently published an extension on that idea.) This reminds me very much of the missing measure approach of Kong et al. (2003). (Actually the reference appears in the final discussion.) The authors go over several paradoxes like Basu’s circus estimate (1988), Larry’s inconsistent Bayes estimate (2004), where Horvitz–Thompson performs nicely under compactness assumptions, the Bayesian answers (which include nested sampling even though I do not see the connection). Especially Li’s (2010) solution. The attached numerical experiment displays a consistent underperformance of the Horvitz-Thompson estimator, in contrast with the theory…
In conclusion, while enjoying very much revisiting so many examples and papers I came across in the past decades, I remain somewhat puzzled by the lack of overall message.


