Preamble The simplest explanation is always the best among the explanations that are representative. It is called the principle of parsimony, Occam's razor . It is the bedrock of scientific enlightenment. There is a recent development that people start to do a category error to discard this principle based on the following setting that lead to misunderstanding: if we say that the simplest model…
Figure: Dual Tomographic Compression Performance, Süzen, 2025. Preamble Randomness is elusive and its probably one of the outstanding concepts for human scientific endeavour, along with gravity. Kolmogorov complexity, appears to be so novel in trying to answering "what is randomness?". The idea that the length of the smallest model that can generate the random sequence determines its complexity…
Preamble Bekenstein's Information (Wikipedia) Black holes are not too esoteric anymore after LIGO's success and successful imaging efforts. Their entropy behaves much different than the entropy of an ordinary matter. This leads to incredible discovery of so called holographic principle . The principle stating that we live on a projection of higher-dimensional manifestation of universe. On the…
Preamble Gaussian Curvature (Wikipedia) One of the core concepts in Physics is so called metric tensor . This object encodes any kind of geometry. Combined genius of Gauß, Riemann and their contemporaries lead to such a great idea, probably one of the achievements of human quantitative enlightenment. However, due to notational aspects and lack of obvious pedagogical introduction, making object…
Preamble David Hume (Wikipedia) Experimental design is not a new concept and randomised control trials (RCTs ) are our solid gold standard of doing quantitative research, when no apparent physical laws are available to validate observations. However, it is very expensive to design RCTs, not ethical or either not possible due to logistical reasons in some cases. Then we fall into Causal Inference's…
Preamble Occam (Wikipedia) A misconception that overfitted model can be identified with the amount of generalisation gap between model's training and test sets over its learning curves is still out there. Even in some prominent online lectures and blog posts, this misconception is now repeated without critical look. In general, this practice unfortunately diffuse into some academic papers and…
Boltzmann (Wikipedia) Post covers the papers: H-theorem do-conjecture, M. Süzen, arxiv:2310.01458 (2023) Preamble Probably the most important achievement for humans is the ability to produce scientific discoveries, that helps us objectively understand how nature works and build artificial tools where no other species can. Entropy is an elusive concept and one of the crown achievements of human…
Jacob Bekenstein (Wikipedia) Preamble Thermodynamics of black holes has appeared as one of the most interesting areas of research in theoretical physics [ Wald1994 ], specially after LIGO's massive success. The striking results of Jacob Bekenstein [ Bekenstein1973 ] in proposing a formulation of entropy for a black hole was on of the most striking turning point in building explanations for the…
Preamble Babylonian Tablet for square root of 2. (Wikipedia) Prediction implies a mechanics, as in knowing a form of a trajectory over time. Strictly speaking a predictive system implies knowing a solution to the path, set of variable depending on time, time evolution of the system under consideration. Here, we define semi-informally how a prediction system is defined mathematically and show how…
Preamble Walt Disney Hall, Los Angeles (Wikipedia) Unfortunately, it is still thought in machine learning classes that overfitting can be detected by comparing training and test learning curves on the single model's performance. The origins of this misconception i s unknown. Looks like an urban legend has been diffused into main practice and even in academic works the misconception taken granted.…
Preamble The Tilled Field , Joan Miró (Wikipedia) One of the core concepts in data sciences is conditional probabilities, $p(x|y)$ appear as logical description of many of the tasks, such as formulating regression or as a core concept in Bayesian Inference . However, there is operationally no special meaning of a conditional or joint probabilities as their arguments are no more than a…
Preamble Sample space is the primary concept introduced in any probability and statistics books and in papers. However, there needs to be more clarity about what constitutes a sample space in general: there is no explicit distinction between the unique event set and the replica sets. The resolution of this ambiguity lies in the concept of an ensemble. The concept is first introduced by American…
Preamble Figure: Moon patterns human brain invents. (Wikipedia) Detecting overfitting is inherently a comparison problem of the complexity of multiple objects, i.e., models or an algorithm capable of making predictions. A model is overfitted ( underfitted ) if we only compare it to another model. Model selection involves comparing multiple models with different complexities. The summary of this…
Solar Eclipse of 1919 (wikipedia) Preamble Cool ideas in theoretical physics are ofter opaque for general reader whether if they are backed up with any experimental evidence in the real world. The success of LIGO (Laser Interferometer Gravitational-wave Observatory) definitely proven the value of interferometry for advancement of cool ideas of theoretical physics supported by real world measurable…
Preamble George Boole (Wikipedia) Agent, AI agent or an intelligent agent is used often to describe algorithms or AI systems that are released by research teams recently. However, the definition of an intelligent agent (IA) is a bit opaque. Naïvely thinking, it is nothing more than a decision maker that shows some intelligent behaviour. However, making a decision intelligently is hard to quantify…
Preamble The White Rabbit (Wikipedia) A novice analyst or even experienced (data) scientist would have thought that the bar notation $|$ in representing conditional probability carries some different operational mathematics. Primarily when written in explicit distribution functions $p(x|y)$. Similar approach applies to joint probabilities such as $p(x, y)$ too. One could see a mixture of these,…
Simionescu Function (Wikipedia) Preamble The holy grail of machine learning appears to be the empirical risk minimisation . However, on the contrary to general dogma, the primary objective of machine learning is not risk minimisation per se but mimicking human or animal learning . Empirical risk minimisation is just a snap-shot in this direction and is part of a learning measure, not the primary…
Preamble Figure 1: Two observable's approach to ergodicity for Bernoulli Trials. Ergodicity appears in many fields, in physics, chemistry and natural sciences but in economics to machine learning as well. Recall that, ergodicity in physics and mathematical definition diverges significantly due to Birkhoff's statistical definition against Boltzmann's physical approach . Here we will follow…
Figure: Maxwell's handwritings, state diagram (Wikipedia) Preamble The modern statistics now move into an emerging field called data science that amalgamate many different fields from high performance computing to control engineering . However, the emergent behaviour from researchers in machine learning and statistics that, sometimes they omit naïvely and probably unknowingly the fact that some of…
Preamble Dali (1931), The Persistence of Memory (Wikipedia) One of the new mathematical concepts arise due to understanding of deep learning is called periodic spectral ergodicity (PSE). The cascading PSE (cPSE) propagates over deep learning layers which can also be used as a complexity measure. cPSE actually can also predict the generalisation ability. In this post, we review this interesting…
Preamble Figure: Monalisa on Eigenvector grids (Wikipedia) In the post, A New Matrix Mathematics for Deep Learning : Random Matrix Theory of Deep Learning , we have outlined a new mathematical concepts that are aimed at deep learning but in general belonging to applied mathematics. Here, we dive into one of the concepts, spectral ergodicity. We aimed at conveying what does it mean and how to…
Preamble Figure: Definition of Randomness (Compagner 1991, Delft University) Development of deep learning systems (DLs) increased our hopes to develop more autonomous systems. Based on the hierarchal learning of representations , deep learning defies the basic learning theory that beg the question of still rethinking generalisation . Even though DLs lacks severely the ability to reason without…
Preamble Progress in machine learning, specifically so-called deep learning , last decade was astonishingly successful in many areas from computer vision to natural language translation reaching automation close to human-level performance in narrow areas, so-called narrow artificial intelligence. At the same time, the scientific and academic communities also joined in applying deep learning in…
Kindly reposted to KDnuggets by Gregory Piatetsky-Shapiro with the title Data science is not about data -applying Dijkstra principle to data science and enhancements. Prelude Dijkstra in Zurich, 1984 (Wikipedia) Edsger Dijkstra was a Dutch theoretical physicist turned computer scientist, and probably one of the most influential earlier pioneers in the field. He had deep insight in what is computer…
Optimal learning : Meta-optimization Many papers directly equate “machine” learning problem, algorithmic learning oppose to human or animal learning, with optimisation problem. Unfortunately, contrary to common belief machine learning is not an optimisation problem. For example, take optimal learning strategy, a replace learning with optimisation and we end up having and absurd terms of optimal…