RSSAmplifier

Blog

Reasonable Deviations

reasonabledeviations.comRSS feed ↗44 posts

Latest posts

Man and machine: GPT for second brains

In this post, I discuss how I used GPT embeddings to build a smart search tool for my second brain note-taking system. Try this if the video isn't working for you Introduction I spend a lot of time thinking about how to improve my processes for learning and knowledge management; I’ve written extensively about my second-brain note-taking system, Molecular Notes. I use Molecular Notes on a daily…

Molecular Notes: Practice

In this Part 2, I discuss the practical implementation of Molecular Notes in Obsidian. I explain how I organise my Second Brain (Tags, Folders, Topics) and present detailed workflows for ingesting different types of content. I also explain how one can extend Obsidian, giving the example of Polymer – a spaced repetition application that I built on top of my Second Brain. After reading this post, if…

Molecular Notes: Principles

In this post I present Molecular Notes, a note-taking system I created to help me learn from diverse sources (books, textbooks, articles, courses), distil insights, and synthesise new ideas. Molecular Notes is how I approach my Second Brain – the body of concepts and ideas that are relevant to my understanding of the world, both personally and professionally. In general, I try to seek out mental…

Convexity in DCFs

In this post, we revisit classical DCFs through the lens of convexity. This leads to the counterintuitive finding that increased uncertainty about an asset’s fundamentals can sometimes be a good thing! Overview of DCFs DCF analysis is used to value cash flow generating assets. The key idea is that we prefer cash today to cash a year from now, such that £100 next year is only worth $£100 \nu$ to us…

How I Read Books

I recently launched a website to open-source some of my book reviews. To accompany this, I’d like to share some thoughts on my philosophy of reading books, and how my current workflow reflects this philosophy. This post primarily focuses on non-fiction because it is an area in which I think I have “edge” – I am confident in my ability to extract value from non-fiction books. I do enjoy fiction…

Probability matching and Kelly betting

In this post, we discuss a cognitive bias called probability matching, explaining how it is rational from a population perspective. We then make an analogy to the Kelly criterion, a betting strategy that finds widespread use in both gambling and finance. Probability matching I probably don’t need to tell you that humans aren’t perfectly rational utility-maximisers – thanks to the work of…

How I use Notion

This post is a fairly comprehensive discussion of how I use Notion (a free personal knowledge management app) to organise various aspects of my life: project management, reading, academics, plans/goals, investing, and more. The post is not designed to be read linearly – pick and choose the bits that are relevant to you. Inevitably, this is going to sound like one massive ad for Notion. I have no…

Hypothesis testing in quant finance

At its core, science is about making falsifiable hypotheses about the world (Popper), testing them experimentally, then using the experiment outcomes to refute or refine the hypotheses. The scientific method is an integral part of quantitative finance; it provides a framework we can use to identify and analyse trading signals or anomalies. In this short post, we discuss a general method for…

COVID-19 Beta

In this short post, we compute and visualise “COVID-19 betas” for stocks in the S&P500 index, to quantitatively and visually understand which companies were most affected (positively and negatively) by COVID-19. For those of you who just want to see the (interactive!) result, here it is. Click on any sector to zoom in on its constituents: If you would like to generate this plot for yourself,…

Option-implied probability distributions, part 2

In Part 1 of this series, we demonstrated that the prices of option butterfly spreads imply a probability distribution of prices for the underlying asset. In this post, we will first examine the limiting case of butterfly spreads. Then, we will tackle the industry-standard approach for constructing PDFs from option prices: interpolating in volatility space to generate a volatility surface,…

Option-implied probability distributions, part 1

The key idea in Expectations Investing , a well-known book by Alfred Rappaport and Michael Mauboussin, is that to profitably invest in stocks one needs to find situations in which your view is variant to consensus expectations – a difficult task when you don’t know what consensus expectations are. There are several ways to figure out what the Street is thinking, the most common of which is to read…

Equity Investing Checklist

This page represents a work-in-progress checklist for equity investing. It is continuously being refined based on my reading, as well as my actual experience in researching and investing. The hope is that this provides transparency and accountability (for myself), to better understand my cognitive biases and weak spots. An important idea underlying this checklist is that time is finite, and there…

A critical look at Greenblatt's Magic Formula

As the saying goes, when something sounds too good to be true, it probably is – all the more so when it comes to investing. In this short post, we look at the Magic Formula of Joel Greenblatt, as described in The Little Book That Still Beats the Market , critically examining the strategy and attempting to quantify its alpha. What is unusual about this Magic Formula is that the person behind it is…

The Big Shorts: what is the smart money betting against?

Hedge funds in the UK are legally obligated to disclose to the Financial Conduct Authority (FCA) whenever their net short position in a particular listed company reaches 0.5% of the issued equity capital. In this post, I investigate a publicly available dataset containing information about these large institutional short positions in UK equities and attempt to understand the value that this data…

Statistical arbitrage in closed-end funds

Sometimes, it is cheaper to buy a basket of assets than it is to buy the assets in the basket. In this post, we discuss closed-end funds and why they often trade at a discount to their net asset value. Furthermore, we explore whether this could be the basis for an algorithmic trading strategy. Closed-End Funds vs ETFs Most of us are probably familiar with exchange-traded funds (ETFs) – baskets of…

A Tanker Trade

April 2020 has been a volatile month for oil. Last week, the May WTI contract traded at a low of minus \$40 a barrel. In a desperate search for storage space, people have been chartering oil tankers to use as floating storage units, leading to a price surge in shares of tanker companies like Nordic American Tanker (46%), Teekay (30%), and Scorpio Tankers (59%). In this post, we aim to build a…

Understanding the market's expectations of COVID-19

One of the reasons why I find markets fascinating is that they are capable of integrating huge amounts of information, misinformation, hope, fear and uncertainty into a single number – the price of an asset. In this post, we work backwards, quantitatively examining what the current price of an asset can tell us about its future prospects. The importance of expectations For an investment to be…

Rebuilding PyPortfolioOpt: an open source adventure

A few weeks ago, a user raised an issue on the GitHub repository for PyPortfolioOpt , my open-source portfolio optimisation software library. In this nontechnical post, I discuss why a seemingly innocuous error resulted in a ground-up rebuild of a large chunk of PyPortfolioOpt, and share some reflections on open-source in general. The actual bug that was reported is not particularly important, but…

Black-Litterman allocation in algorithmic trading

In December 2019, I released a major update to PyPortfolioOpt , my python portfolio optimisation package. The most significant addition was an implementation of the Black-Litterman (BL) method. Although BL optimisation is commonly used as part of a pipeline to optimise a multiasset/equity portfolio, in this post I argue that BL is particularly well suited to the problem of optimally weighting…

An asymmetric bet on interest rates

In a classic scene of No Country For Old Men , Javier Bardem’s character ominously asks a shopkeeper: “what’s the most you ever lost on a coin toss?”. The shopkeeper says that he doesn’t know – this is probably quite a reasonable response given that for a fair coin, one has little reason to make a bet since your expected value (EV) is zero. Yet retail investors seem to make coin-toss bets all the…

How predictive is the historical volatility?

One of the things that makes markets exciting (or frightening) is that prices move around a lot. It is important to be able to describe and predict the range of possible price movements over a given time horizon since some investors might desire assets whose prices don’t move up and down too much. We can quantify this by computing the volatility , which is commonly defined to be the standard…

Implementing k-means clustering from scratch in C++

I have a somewhat complicated history when it comes to C++. When I was 15 and teaching myself to code, I couldn’t decide between python and C++ and as a result tried to learn both at the same time. One of my first non-trivial projects was a C++ program to compute orbits – looking back on it now, I can see that what I was actually doing was a (horrifically inefficient) implementation of Euler’s…

What we learnt building an enterprise-blockchain startup

It has been almost a year since the idea of HyperVault was first conceived. In that time, we built HyperVault up from a single sentence, gained and lost team members along the way, developed a functional proof-of-concept over the short winter holidays, crashed out of a few competitions (also won a couple of prizes), and finally decided to open source. This post aims to be an honest reflection on…

Graph algorithms and currency arbitrage, part 2

In the previous post (which should definitely be read first!) we explored how graphs can be used to represent a currency market, and how we might use shortest-path algorithms to discover arbitrage opportunities. Today, we will apply this to real-world data. It should be noted that we are not attempting to build a functional arbitrage bot, but rather to explore how graphs could potentially be used…

Graph algorithms and currency arbitrage, part 1

Arbitrage is the holy grail for traders and the bedrock of financial academia. Let’s say you are in an open marketplace with Alice selling oranges for \$1 each and Bob buying them for \$2. As a cunning trader, you realise you can buy an orange from Alice and immediately sell it to Bob for \$1 of “risk-free” profit. However, as you keep reaping this \$1 profit by buying up Alice’s oranges, she…

Portfolio optimisation: lessons learnt

Over the past few months I have been busy doing a mixture of blockchain consulting and quantitative finance research. In particular, I have had the opportunity to investigate the interesting problem of portfolio management for cryptoassets – it was not my first experience with portfolio optimisation, having implemented efficient frontier portfolios at a roboadvisor startup, but this time I took…

Exponential Covariance

For the past few months, I have been doing a lot of research into portfolio optimisation, whose main task can be summarised as follows: is there a way of combining a set of risky assets to produce superior risk-adjusted returns compared to a market-cap weighted benchmark? The answer of Markowitz (1952) is in the affirmative, with some major caveats. Given the expected returns and the covariance…

Stormy Seas for Proof of Work

In this post we will be examining one of the main problems with Proof of Work (PoW) – not the energy inefficiency (as it is debatable how much of a problem this really is), but something more fundamental with the consensus process. In the past couple of months we have seen a number of cryptocurrencies fall victim to 51% attacks. Verge, Bitcoin Gold, ZenCash, and Electroneum are just a few coins…

Evolving cellular automata to solve problems, part 2

We will be picking up where the previous post left off. As a brief summary, we are attempting to replicate the results of Evolving Cellular Automata with Genetic Algorithms (Mitchell, Crutchfield and Das 1996), dealing with the density classification task for 1D binary cellular automata (CAs). To put it simply, we are trying to design a ruleset such that the final configuration of a cellular…

Evolving cellular automata to solve problems, part 1

Recently I finished reading Complexity: A Guided Tour , by Melanie Mitchell, which reminded me a lot of Gödel, Escher Bach (indeed, the book is dedicated to Douglas Hofstadter). It has reminded me that emergence is an incredibly fascinating concept – simple individual units somehow coming together to result in complex behaviour that cannot really be explained in terms of the components. In this…

Classifying financial time series using Discrete Fourier Transforms

A financial time series represents the collective decisions of many individual traders; it seems reasonable to me that the nature of these decisions may differ based on the underlying asset. For example, a company with a higher market cap may be more liquid, and subject to larger individual buy/sell orders including institutional investment. Thus, there is a case to be made that information such…

DIY MachineLearningStocks

I recently released a machine learning stock prediction project on GitHub , unimaginatively named MachineLearningStocks . It is a project that I’ve put quite a lot of time into, and is in fact a simplified version of a system that I’ve been using to live trade. This post doesn’t really offer anything on top of the existing readme, but I figured it would be good to have a copy (with some minor…

Creating a stock price database with MariaDB and python

One of my interests is exploring the applications of machine learning to financial markets. As part of this hobby, I’ve spent many more hours parsing and processing data than I have actually applying machine learning. I’ve worked broadly with two datasets in particular: historical financial statistics (e.g. P/E ratio, price/book) make up the features that my algorithms learn from, but the actual…

Learning Machine Learning

Two years ago I was an absolute novice at machine learning: I had read around the subject a little bit, could probably rattle off a few of the buzzwords, and had some appreciation of the general idea, but there was no way I could have developed a predictive model beyond linear regression. I was somewhere near the peak of Mount Stupid (from a great chart by SMBC ): Fast forward to the present day –…

Gradient tree boosting and XGBoost

Decision trees make for pretty vanilla classifiers: they do an unspectacular job with most machine learning tasks, and you’d be forgiven for overlooking them when deciding on a classification algorithm. But decision trees happen to be the cornerstone of a powerful class of learning algorithms: gradient tree boosting methods. I will try to elucidate the (short) history of gradient tree boosting,…

8-bit Julia set art in python

You may have heard a mathematician or physicist (or more likely your maths teacher) describe mathematics as beautiful . What could they mean by this? There is just something mysteriously attractive about the purity, complexity, interconnectedness, and underlying truth of it all (“Beauty is truth, truth beauty” - Keats). I can’t really say more than that, so I will leave you with a quote from the…

Retrieving historical stock prices from Yahoo Finance with no API

Yahoo Finance has long been an excellent free financial resource with a wealth of data and a convenient API, allowing open source programming libraries to access stock data. But not any more. As of May 2017, they have discontinued their API , probably as a result of Yahoo’s pending acquisition by Verizon. This means that excellent tools like pandas-datareader are now broken, much to the dismay of…

The Leibniz integral rule in electrodynamics

I have been slowly working through David Griffith’s widely-used textbook, Introduction to Electrodynamics . It is known for having a large number of reasonably difficult exercises, which are instrumental in conveying some concepts not directly addressed in the main text. I found an elegant shortcut to one of the questions, which was not noted in the solution manual. The shortcut involves (ab)using…

Intuiting the gamma function, part 3

We ended Part 2 with the stunning result that \[x! = \int_0^\infty t^x e^{-t}dt\] Of course, this is not exactly the same as the gamma function, because the gamma function is defined as: \[\Gamma (x) = \int_0^\infty t^{x-1}e^{-t} dt\] A direct observation of the above leads to the conclusion that: \[\Gamma (x) = (x-1)! \qquad \text{or alternatively} \qquad x! = \Gamma (x+1)\] This ‘shift’ is a…

Intuiting the gamma function, part 2

In Part 1 , we showed that repeated differentiation gives rise to a factorial. In this second post of Intuiting the gamma function , we are going to show that integration by parts can also produce a factorial – an instrumental step in generalising the idea of a factorial to non-integers. Just a couple of minor points regarding the presentation of this post. I will abbreviate ‘integration by parts’…

Intuiting the gamma function, part 1

What is the factorial of a half? This series of posts builds up from middle-school algebra to the enigmatic gamma function in an attempt to answer this simple question. When high-school students study the Binomial Theorem, a classic problem is the expansion of a binomial with a fractional power, such as $(1+x)^{1/2}$. Speaking from experience, a student’s first attempt is often to naïvely…

Combinatorial optimisation with a pseudo-genetic algorithm

a python approach to XKCD’s Social Seating problem Social groups are remarkably complex affairs, and on many occasions this complexity can lead to awkward situations. Randall Munroe humorously portrays one such example in XKCD #173. This comic naturally begs the question – how do we find the optimal linear seating arrangement for a given social group? If you’re not a fan of people trying to…

Conway's Game of Life in python

In this short post, I explain how to implement Conway’s Game of Life in python, using numpy arrays and matplotlib animations. I emphasise intuitive code than performance, so it could be a useful starting point for somebody to understand the logic before implementing a more efficient version. Cellular automata A cellular automaton (pl. automata) consists of a grid of cells ; each cell has a state —…

The Cambridge Natural Sciences Interview

In December 2015, just after finishing my IB exams, I went to Cambridge for my (Physical) Natural Sciences Interview. The interview has a reputation for being incredibly demanding and intense, and you’ve probably seen on the internet some of the ‘crazy’ questions that people get asked. Although it is definitely stressful, the Cambridge interview is almost purely academic (and quite reasonable). So…