RSSAmplifier

Blog

Robin Linacre's homepage: Probabilistic record linkage, Data Science and Data Engineering

Probabilistic record linkage, Data Deduplication, Data Science, Engineering and the Environment

robinlinacre.comRSS feed ↗66 posts

Latest posts

letterpaths: or how LLMs can be good even when they're bad

Using visualisations and disposable GUIs to give faster feedback to LLMs

Optimising DuckDB performance on large EC2 instances

Practical notes on getting DuckDB to run efficiently on large EC2 instances

Letterpaths - a free library for teaching cursive writing

Letterpaths is a free and open source library for powering educational apps

Splink or Swim: The Benefits of Being Small Enough To Fail

A behind the scenes look at winning the Civil Service innovation award

Respectful use of AI in software development teams

How to use LLMs in team settings without damaging team health

Measuring the accuracy of record linkage

Some thoughts on how to measure the accuracy

Using a fault tolerant trie for address matching

An interactive explanation of how a fault tolerant trie can be used for address matching

Building Accurate Address Matching Systems

A bag of tricks to improve the accuracy of geocoding

Putting Scaffolding Around Vibe Coding to Build More Complex Apps

How to work past the limits of LLMs to build more complex apps

Why DuckDB is my first choice for data processing

Why DuckDB has become my go-to tool for data processing, offering simplicity, speed, and powerful features.

An alternative way to think about predicted probabilities in the Fellegi Sunter model

A second equivalent mental model to help think about how we arrive at predicted probabilities in the Fellegi Sunter model

Live DuckDB WASM Splink model

A demo of a live Splink model in the browser

Graph editor for illustrating clustering concepts (graph playground)

Graph editor for illustrating clustering concepts

AI probably won't replace me in 2025

My mental model of LLMs, their strengths and shortcomings

The emerging impact of LLMs on my productivity

The emerging impact of LLMs on productivity

Splink: Transforming data linking through open source collaboration

How a small team of MoJ analysts built data linking software that’s used by governments across the world

Connected components visualisation

A visualisation of how the connected components algorithm works

Match weight calculator

A calculator for converting between match weights, probabilities, and Bayes factors

Super-fast deduplication of large datasets using Splink and DuckDB

Evaluating 1 billion record comparisons to deduplicate 7 million records in two minutes

Why Probabilistic Linkage is More Accurate than Fuzzy Matching For Data Deduplication

How to ensure that all available information is used to make predictions

Thoughts and questions about the short term impact of LLMs on knowledge workers

What will be the impact of LLMs on knowledge workers

Visualising updating a prior

Using treemaps to visualise updating the prior with information about a scenario in the Fellegi Sunter model

Computing the Fellegi Sunter model

A set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article shows how to compute the model from an algorithmic perspective

m and u values in the Fellegi-Sunter model

Deep dive into the role and interpretation of m and u probabilities in the Fellegi-Sunter model for probabilistic linkage. Learn how these probabilities impact match weights and how to quantify the strength of evidence in favor or against a record match.

Partial match weights

Partial match weights in the Fellegi-Sunter model. Part of an explorable, interactive introduction to probabilistic record linkage (data deduplication) theory

The relationship between probabilities, match weights and Bayes factors

Visualising the correspondence between match weights, probabilities, Bayes factors and their intuitive explanations

Splink and the Open Source Dividend

Splink and the open source dividend

SQL should be the default choice for data transformation logic

SQL should be the first option considered for new data engineering work. It’s robust, fast, future-proof and testable. With a bit of care, it’s clear and readable.

Why parquet files are my preferred API for bulk open data

Open data should be served as CORS-enabled parquet files rather than using a custom API

Why don't you just

The phrase 'why don't you just' is problematic

The Intuition Behind the Use of Expectation Maximisation to Train Record Linkage Models

An intuitive explanation for how the Expectation Maximisation algorithm is able to produce unsupervised estimates of Splink model parameters

Splink 3: Fast, accurate and scalable linkage in Python

Splink 3 now offers support for Python and AWS Athena backends, in addition to Spark. It's now easier to use, faster and more flexible, and can be used for close to real time linkage.

m and u probability generator with starting values

Generate m and u probabilities to input into Splink. Part of the introduction to Fellegi Sunter series.

Are more complex probabilistic linkage models more accurate? Part 2, unsupervised learning

How good is Splink: Are more complex probabilistic linkage models more accurate?

Are more complex probabilistic linkage models more accurate? Part 1, supervised learning

How good is Splink: Are more complex probabilistic linkage models more accurate?

The Thorniest Problem of Building an Analytical Platform

The Thorniest Problem of Building an Analytical Platform: Enabling collaborative development of the platform itself without losing control of complexity.

The carbon impact of switiching to an electric car

What is the comparative carbon footprint of electric cars? As an existing petrol ICE car owner, should you switch to an electric car

m and u probability generator

Generate m and u probabilities to input into Splink. Part of the introduction to Fellegi Sunter series.

Dependencies between match weights

An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. The dependencies between match weights.

Understanding match weights in the Fellegi Sunter model

An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article discusses match weights.

Visualising the Fellegi Sunter model

An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article presents a way of visualising how the model works.

Maths of Fellegi Sunter (old version)

An set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article shows how to compute the model

The mathematics of the Fellegi Sunter model

A set of interactive, explorable explanations of the Fellegi Sunter model of probabilistic record linkage. This article shows the derivation of the mathematical formulation of the model

An Interactive Introduction to Record Linkage (Data Deduplication) in the Fellegi-Sunter framework

The first in a series of interactive, explorable explanations of the Fellegi-Sunter model, providing an introduction to probabilistic record linkage (data deduplication).

The Downfall of Command and Control Data Leadership

The Downfall of Command and Control Data Leadership - why new big bang data platforms fail

Demystifying Apache Arrow

Demystifying Apache Arrow - some observations from a data scientist. Learning more about a tool that can filter and aggregate two billion rows on a laptop in two seconds

Birdsong quiz

Test how good you are at identifying UK birdsong recordings

Birdsong recording finder

Listen to UK birdsong using the xeno-canto API

Comparing energy usage across countries

Filling the country with solar panels