I ve switched the host of my blog to Ghost, so my blog is now located at https://bounded-regret.ghost.io/. This is different from the previous re-hosting on my personal website, which is no longer maintained. Feedly users can subscribe via RSS here.Or click the bottom-right subscribe button on this page to get e-mail notifications of new posts.
[Update: This post is out of date. My blog has now moved again, to https://bounded-regret.ghost.io/.] I ve decided to move this blog to my personal website. It s now located here: https://jsteinhardt.stat.berkeley.edu/blog/, including all the old posts and comments, plus some new and upcoming posts :).
Suppose that we want to construct subsets with the following properties: for all for all The goal is to construct as large a family of such subsets as possible (i.e., to make as large as possible). If , then up to constants it is not hard to show that the optimal number of sets is [ ]
I ve spent much of the last few days reading various ICML papers and I find there s a few pieces of feedback that I give consistently across several papers. I ve collated some of these below. As a general note, many of these are about local style rather than global structure; I think that good local style [ ]
In my previous post, “Latent Variables and Model Mis-specification”, I argued that while machine learning is good at optimizing accuracy on observed signals, it has less to say about correctly inferring the values for unobserved variables in a model. In this post I’d like to focus in on a specific context for this: inverse reinforcement [ ]
Here is interesting linear algebra fact: let be an matrix and be a vector such that . Then for any matrix , . The proof is just basic algebra: . Why care about this? Let s imagine that is a (not necessarily symmetric) stochastic matrix, so . Let be a low-rank approximation to (so consists of [ ]
Consider the following statements: The shape with the largest volume enclosed by a given surface area is the -dimensional sphere. A marginal or sum of log-concave distributions is log-concave. Any Lipschitz function of a standard -dimensional Gaussian distribution concentrates around its mean. What do these all have in common? Despite being fairly non-trivial and deep [ ]
Machine learning is very good at optimizing predictions to match an observed signal for instance, given a dataset of input images and labels of the images (e.g. dog, cat, etc.), machine learning is very good at correctly predicting the label of a new image. However, performance can quickly break down as soon as we [ ]
In my post on where I plan to donate in 2016, I said that I would set aside $2000 for funding promising projects that I come across in the next year: The idea behind the project fund is [to] give in a low-friction way on scales that are too small for organizations like Open [ ]
The following explains where I plan to donate in 2016, with some of my thinking behind it. This year, I had $10,000 to allocate (the sum of my giving from 2015 and 2016, which I lumped together for tax reasons; although I think this was a mistake in retrospect, both due to discount rates and [ ]