letterpaths: or how LLMs can be good even when they're bad
Using visualisations and disposable GUIs to give faster feedback to LLMs
Probabilistic record linkage, Data Deduplication, Data Science, Engineering and the Environment
Using visualisations and disposable GUIs to give faster feedback to LLMs
Practical notes on getting DuckDB to run efficiently on large EC2 instances
Letterpaths is a free and open source library for powering educational apps
A behind the scenes look at winning the Civil Service innovation award
How to use LLMs in team settings without damaging team health
Some thoughts on how to measure the accuracy
An interactive explanation of how a fault tolerant trie can be used for address matching
A bag of tricks to improve the accuracy of geocoding
How to work past the limits of LLMs to build more complex apps
Why DuckDB has become my go-to tool for data processing, offering simplicity, speed, and powerful features.
A second equivalent mental model to help think about how we arrive at predicted probabilities in the Fellegi Sunter model
A demo of a live Splink model in the browser
Graph editor for illustrating clustering concepts
My mental model of LLMs, their strengths and shortcomings
The emerging impact of LLMs on productivity
How a small team of MoJ analysts built data linking software that’s used by governments across the world
A visualisation of how the connected components algorithm works
A calculator for converting between match weights, probabilities, and Bayes factors
Evaluating 1 billion record comparisons to deduplicate 7 million records in two minutes
How to ensure that all available information is used to make predictions
What will be the impact of LLMs on knowledge workers
Using treemaps to visualise updating the prior with information about a scenario in the Fellegi Sunter model
A set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article shows how to compute the model from an algorithmic perspective
Deep dive into the role and interpretation of m and u probabilities in the Fellegi-Sunter model for probabilistic linkage. Learn how these probabilities impact match weights and how to quantify the strength of evidence in favor or against a record match.
Partial match weights in the Fellegi-Sunter model. Part of an explorable, interactive introduction to probabilistic record linkage (data deduplication) theory
Visualising the correspondence between match weights, probabilities, Bayes factors and their intuitive explanations
Splink and the open source dividend
SQL should be the first option considered for new data engineering work. It’s robust, fast, future-proof and testable. With a bit of care, it’s clear and readable.
Open data should be served as CORS-enabled parquet files rather than using a custom API
The phrase 'why don't you just' is problematic
An intuitive explanation for how the Expectation Maximisation algorithm is able to produce unsupervised estimates of Splink model parameters
Splink 3 now offers support for Python and AWS Athena backends, in addition to Spark. It's now easier to use, faster and more flexible, and can be used for close to real time linkage.
Generate m and u probabilities to input into Splink. Part of the introduction to Fellegi Sunter series.
How good is Splink: Are more complex probabilistic linkage models more accurate?
How good is Splink: Are more complex probabilistic linkage models more accurate?
The Thorniest Problem of Building an Analytical Platform: Enabling collaborative development of the platform itself without losing control of complexity.
What is the comparative carbon footprint of electric cars? As an existing petrol ICE car owner, should you switch to an electric car
Generate m and u probabilities to input into Splink. Part of the introduction to Fellegi Sunter series.
An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. The dependencies between match weights.
An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article discusses match weights.
An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article presents a way of visualising how the model works.
An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article shows how to compute the model
A set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article shows the derivation of the mathematical formulation of the model
The first in a series of interactive, explorable explanations of the Fellegi-Sunter model, providing an introduction to probabilistic record linkage (data deduplication).
The Downfall of Command and Control Data Leadership - why new big bang data platforms fail
Demystifying Apache Arrow - some observations from a data scientist. Learning more about a tool that can filter and aggregate two billion rows on a laptop in two seconds
Test how good you are at identifying UK birdsong recordings
Listen to UK birdsong using the xeno-canto API