In a previous blog, we experimented with reinforcement learning (RL) for solving a car control challenge. We found RL methods had difficulties converging to optimal values, despite being successful in a toy cart simulation. Learning and exploring simultaneously with complex real-world noise is hard.
Laplace wrote this great essay on probability1. This was back in 18th c. when mathematicians prose was the language of math instead of formulas, and there was no differentiation between philosophy and the practical application of math from its theory. https://en.wikipedia.org/wiki/A_Philosophical_Essay_on_Probabilities
A cool property of ML is that we can interpret the logits, outputs of the model, as anything we want. This leads to all sorts of rich interpretations . Well, almost anything. The interpretations from logits don’t actually come from thin air, but derived from a few simple assumptions about the prior distribution and linear relationship. So from here our interpretation of logits, and our subsequent…