The central limit theorem can be informally summarized in few words: The sum of x 1 , x 2 , ... x n samples from the same distribution is normally distributed, provided that n is big enough and that the distribution has a finite variance. to show this in an experimental way, let's define a function that sums n samples from the same distrubution for 100000 times: import numpy as np import…
This year I decided to join the March Machine Learning Mania 2021 - NCAAW challenge on Kaggle. It proposes to predict the outcome of each game into the basketball NCAAW tournament, which is a tournament for women at college level. Participants can assign a probability to each outcome and they're ranked on the leaderboard according to the accuracy of their prediction. One of the most attractive…
I recently published on a wrapper around The Dictionary of Obscure Words (originally from this website http://phrontistery.info ) for Python and in this post we'll see how to create a visualization to highlight few entries from the dictionary using the dimensionality reduction technique called T-SNE. The dictionary is available on github at this address https://github.com/JustGlowing/obscure_words…
Have you ever heard of the Travelling Salesman Problem? I'm pretty sure you do, but let's refresh our mind looking at its formulation: "Given a list of points and the distances between each pair of points, what is the shortest possible path that visits each point and returns to the starting point?". What makes this problem so famous and so studied is the fact that it has no "quick" solution as the…
Using regularization has many benefits, the most common are reduction of overfitting and solving multicollinearity issues. All of this is covered very well in literature, especially in (Hastie et all) . Howerver, wihout touching too many details we can have a very straigthforward interpretation of regularization. Regularization is a way to constrain a model in order to learn less from the data. In…
Lately there's a bit of attention about charts where the values of a time series are plotted against the change point by point. This thanks to this rather colorful and cluttered Tornado plot . In this post we will see how to make one of those charts with our favorite plotting library, matplotlib, and we'll also try to understand how to read them. Let's start loading the records of the…
Not too long ago I've been gifted a Raspberry Pi camera, after taking some pictures I realized that it produced very weird colors and I discovered that it was a NoIR camera! It means that it has no infrared filter and that it can take pictures in the darkness using an infrared LED. Since I never found an application that required taking pictures without proper lighting I started wondering if I…
What makes a word beautiful? Answering this question is not easy because of the inherent complexity and ambiguity in defining what it means to be beautiful. Let's tackle the question with a quantitative approach introducing the Aesthetic Potential , a metric that aims to quantify the beaty of a word w as follows: where w + is a word labelled as beautifu, w - as ugly and the function s is a…
A Ridgeline plot (also called Joyplot) allows us to compare several statistical distributions. In this plot each distribution is shown with a density plot, and all the distributions are aligned to the same horizontal axis and, sometimes, presented with a slight overlap. There are many options to make a Ridgeline plot in Python ( joypy being one of them) but I decided to make my own function using…
In this post we will see how to organize a set of movie covers by similarity on a 2D grid using a particular type of Neural Network called Self Organizing Map (SOM) . First, let's load the movie covers of the top 100 movies according to IMDB (the files can be downloaded here ) and convert the images in samples that we can use to feed the Neural Network: import numpy as np import imageio from glob…
Let's say that we want to study the time between the end of a marked point and next serve in a tennis game. After gathering our data, the first thing that we can do is to draw a histogram of the variable that we are interested in: import pandas as pd import matplotlib.pyplot as plt url = 'https://raw.githubusercontent.com/fivethirtyeight' url += '/data/master/tennis-time/serve_times.csv' event =…
In the past we have covered Decision Trees showing how interpretable these models can be (see the tutorials here ). In the previous tutorials we have exported the rules of the models using the function export_graphviz from sklearn and visualized the output of this function in a graphical way with an external tool which is not easy to install in some cases. Luckily, since version 0.21.2,…
In this post we will see a snippet about how to plot a part of the results of the eurobarometer survey released last March. In particular, we will focus on the responses to the following question: Please tell me whether the following statement evokes a positive or negative feeling for you: Immigration of people from other EU Member States. The data from the main spreadsheet reporting the results…
Let's have a look at how to create a visualization that shows how CO2 concentrations evolved in the atmosphere. First, we fetched from the Earth System Research Laboratory website like follows: import pandas as pd data_url = 'ftp://aftp.cmdl.noaa.gov/products/trends/co2/co2_weekly_mlo.txt' co2_data = pd.read_csv(data_url, sep='\s+', comment='#', na_values=-999.99, names=['year', 'month', 'day',…
Lately, on invitation of my right honourable friend Michal , I've been trying to solve some problems from the Euler project and felt the need to have a good way to find prime numbers. So implemented the the Sieve of Eratosthenes . The algorithm is simple and efficient. It creates a list of all integers below a number n then filters out the multiples of all primes less than or equal to the square…
The trend of time series is the general direction in which the values change. In this post we will focus on how to use rolling windows to isolate it. Let's download from Google Trends the interest of the search term Pancakes and see what we can do with it: import pandas as pd import matplotlib.pyplot as plt url = './data/pancakes.csv' # downloaded from https://trends.google.com data =…
Raveling and unraveling are common operations when working with matricies. With a ravel operation we go from matrix coordinate to index coordinates, while with an unravel operation we go the opposite way. In this post we will through an example how they can be done with numpy in a very easy way. Let's assume that we have a matrix of dimensions 4-by-4, and that we want to index of the element (1,…
We have previously seen how to implement KMeans . However, the results of this algorithm strongly rely on the choice of the parameter K. According to statistical folklore the best K is located at the 'elbow' of the clusters inertia while K increases. This heuristic has been translated into a more formalized procedure by the Gap Statistics and in this post we'll see how to pick K in an optimal way…
And here's a function to plot a compact calendar with matplotlib: import calendar import numpy as np from matplotlib.patches import Rectangle import matplotlib.pyplot as plt def plot_calendar(days, months): plt.figure(figsize=(9, 3)) # non days are grayed ax = plt.gca().axes ax.add_patch(Rectangle((29, 2), width=.8, height=.8, color='gray', alpha=.3)) ax.add_patch(Rectangle((30, 2), width=.8,…
Have you ever wanted to check carbon emissions in the UK and never had an easy way to do it? Now you can use the Official Carbon Intensity API developed by the National Grid. Let's see an example of how to use the API to summarize the emissions in the month of May. First, we download the data with a request to the API: import urllib.request import json import pandas as pd import numpy as np import…
Isolation Forest is an algorithm to detect outliers. It partitions the data using a set of trees and provides an anomaly scores looking at how isolated is the point in the structure found, the anomaly score is then used to tell apart outliers from normal observations. In this post we will see an example of how IsolationForest behaves in simple case. First, we will generate 1-dimensional data from…
Lately I've been working a lot with dates in Pandas so I decided to make this little cheatsheet with the commands I use the most. Importing a csv using a custom function to parse dates import pandas as pd def parse_month(month): """ Converts a string from the format M in datetime format. Example: parse_month("2007M02") returns datetime(2007, 2, 1) """ return pd.datetime(int(month[:4]),…
In this post we will see how to create a heatmap with seaborn. We'll use a dataset from the Wittgenstein Centre Data Explorer . The data extracted is also reported here in csv format. It contains the ratio of males to females in the population by age for 1970 to 2015 (data reported after this period is projected). First, we import the data using Pandas: import pandas as pd import numpy as np…
In this post we will see how to create a Multi Layer Perceptron (MLP), one of the most common Neural Network architectures, with Keras. Then, we'll train the MLP to tell apart points from two different spirals in the same space. To have a sense of the problem, let's first generate the data to train the network: import numpy as np import matplotlib.pyplot as plt def twospirals(n_points, noise=.5):…
It's a while that there are no posts on this blog, but the Glowing Python is still active and strong! I just decided to publish some of my post on the Cambridge Coding Academy blog . Here are the links to a series of two posts about Regression Analysis with Decision Trees: 1. Getting started with Regression Analysis and Decision Trees 2. From simple Regression to Multiple Regression with Decision…