Over the past two years I’ve quietly rebuilt major parts of the OldNYC photo viewer . The result: 10,000 additional historic photos on the map , more accurate locations, and a site that’s cheaper and easier to run—thanks to modern AI tools and the OpenStreetMap ecosystem. OldNYC had about 39,000 photos in 2016 . Today it has 49,000 . Most of these changes happened in 2024, but I’m only writing…
For me, 2025 has been the Year of Boggle. It started off as a revival of a 15 year-old project in an attempt to learn how to combine Python and C++. It turned into a months-long, all-consuming affair that eventually landed me on the front page of the Financial Times . It’s always hard to call a project “done,” but my interest in Boggle has trailed off since this summer and I don’t foresee…
The original motivation to revisit my Boggle project back in January was that I wanted to experiment with mixing Python and C++. This had the potential to be a best-of-both-worlds combination: the developer productivity and ergonomics of Python and the performance of C++ for the parts where it really matters. (For an explainer on the Boggle project, check out my announcement post or Ollie Roeder’s…
It’s been four months since I announced my big Boggle result. This post will recap what’s happened since then. What’s a Boffin? My announcement post got some play on Hacker News (it made it to the top five) but, if I’m honest, I was a bit disappointed that it didn’t blow up more. In a bit of a hail mary, I cold emailed Ollie Roeder , who’d written a book ( Seven Games ) about how computers have…
Exciting news! This is the best possible Boggle board: Boggle is a word search game. You form words by connecting adjacent letters, including along diagonals. Longer words score more points. Good words on this board include STRANGERS and PLASTERING. After you spend three minutes trying to find as many words as you can, you’ll be struck by just how good computers are at this. Using the ENABLE2K…
My last boggle post presented an exciting insight that yielded a 30x speedup. This meant two things: I could find the globally-optimal 3x4 Boggle board in ~19 hours using three cores on my laptop, rather than 8-9 hours on a 192-core cloud instance. My cost estimate for finding the globally-optimal 4x4 Boggle board dropped from $500,000→$15,000. That was still more than I was willing to pay, but it…
At the end of my last post on Boggle, I’d achieved perhaps a 10x speedup over my 2009 approach and run my code for 8-9 hours on a 192-core machine to definitively prove that, with 1,651 points , this is the highest-scoring 3x4 Boggle board for the ENABLE2K word list: S L P I A E N T R D E S I was happy with the work. All that was left was to write one final blog post reflecting back on the…
Over the past few weeks I’ve revisited a 15-year old project of mine: trying to find the globally optimal Boggle board. In the last post , I recapped the work I did in 2009 to find the globally-optimal 3x3 Boggle board. In this post, I’ll present a few optimizations I found in 2025 that add up to something like a 10x speed boost over the 2009 approach. Between better algorithms, faster CPUs, and…
Over 15 years ago (!) I wrote a series of blog posts about the board game Boggle. Boggle is a word search game created by Hasbro. The goal is to find as many words as possible on a 4x4 grid. You may connect letters in any direction (including diagonals) but you may not use the same letter twice in a word (unless that letter appears twice on the board). You score points for each word you find, with…
Over the past few weeks, I found the time to work on webdiff , an open source project of mine that I built over a decade ago . I still use it all the time, but hadn’t actively worked on it since 2015. Revisiting an old project is always an interesting experience, and this post presents my reflections on it. First off, what is webdiff? It’s a diff tool. Rather than running git diff , you run git…
I was dimly aware of the ongoing competition between AlphaGo and Lee Sedol , but I hadn’t paid much attention until I saw this chart on reddit : It’s hard to read “Lee Sedol’s brilliant attack (78th)” and not get curious! This led me into a deep dive on the competition. You can read more about the move or watch a 15 minute summary of the match. The full 6 hour match , including a press conference…
I recently added around 1,000 new photos to the map on OldNYC. Read on to find out how! At its core, OldNYC is based on geocoding : the process of going from textual addresses like “9th Street and Avenue A” to numeric latitudes and longitudes. There’s a bit of a mismatch here. The NYPL photos have 1930s addresses and cross-streets, but geocoders are built to work with contemporary addresses.…
I’ve just wrapped up my trip to NIPS 2015 in Montreal and thought I’d jot down a few things that struck me this year: Saddle Points vs Local Minima I heard this point repeated in a talk almost every day. In low-dimensional spaces (i.e. the ones we can visualize) local minima are the major impediment to optimizers reaching the global minimum. But this doesn’t generalize. In high-dimensional spaces,…
I haven’t written a substantial blog post on danvk.org since January . Instead, I’ve been writing over on my groups’s blog at hammerlab.org . Here are the posts I’ve written or edited: SVG→Canvas, the pileup.js Journey (13 Oct 2015) In which I explain why we changed from using SVG to using canvas for our genome browser (spoiler: it’s performance). This also introduces data-canvas , which…
Two weeks ago I launched my latest side project, OldNYC . It’s a collaboration with NYPL Labs which places around 40,000 historical photos of New York City on a map. Avid readers of this blog know that I’ve been working on this for years . The response to OldNYC has been completely overwhelming. Hundreds of thousands of people have used the site. Millions of images have been viewed. Users have…
Here’s the video of my talk from PyCon 2015, Make web development awesome with visual diffing tools : Here are the slides for the talk: The two tools referenced are: dpxdt for generating screenshots webdiff for viewing image diffs I used comparea and dygraphs as sample apps for the demos. If you enjoyed my talk, you might also enjoy Brett’s talk from a few years ago, The Secret of Safe Continuous…
In the last post , we walked through the steps in the Ocropus OCR pipeline. We extracted text from images like this: The results using the default model were passable but not great: O1inton Street, aouth from LIYingston Street. Auguat S, 1934. P. L. Sperr. NO REPODUCTIONS. Over the larger corpus of images, the error rate was around 10%. The default model has never seen typewriter fonts, nor has it…
In the last post , I described a way to crop an image down to just the part containing text. The end product was something like this: In this post, I’ll explain how to extract text from images like these using the Ocropus OCR library. Plain text has a number of advantages over images of text: you can search it, it can be stored more compactly and it can be reformatted to fit seamlessly into web…
As part of an ongoing project with the New York Public Library, I’ve been attempting to OCR the text on the back of the Milstein Collection images. Here’s what they look like: A few things to note: There’s a black border around the whole image, gray backing paper and then white paper with text on it. Only a small portion of the image contains text. The text is written with a tyepwriter, so it’s…
I recently switched back to iOS after a few years using Android. One aspect of Android that always bothered me was that I had trouble finding a great Podcasting app. I’d happily used Instacast on iOS, but I couldn’t find anything quite like it for Android. I eventually settled on Podcast Addict . It epitomizes a stereotype of Android apps: tons of features, gajillions of options and a UI that was…
I recently read Douglas Crockford’s JavaScript: The Good Parts . It’s a classic (published in 2008) which is credited with reviving respect for JavaScript as a programming language. Given its title, it’s also famously short. One very specific thing it cleared up for me was what to do with all of JavaScript’s various substring methods: String.prototype.substr (start[, length])…
I’ve released webdiff 0.8.0 , which you can install via: pip install --upgrade webdiff The most interesting new features are GitHub pull request integration and expanded image diffing modes. You can view a GitHub Pull Request in webdiff by running something like: webdiff https://github.com/hammerlab/cycledash/pull/175 Any github Pull Request URL will do. This will pull down the files from GitHub…
It’s been almost exactly six months since I ended an eight year run at Google. One of the biggest reasons to do this was to come back up to speed with the open source ecosystem and to experience a different working environment (sample size 1→2!). When I joined Google, it seemed like our tech stack was years ahead of anything else. This included tools like BUILD files for managing dependencies and…
My danvk.org site is now fully hosted on GitHub pages. I changed the DNS entry last night. My hope was to do this without breaking anything. That didn’t prove to be possible, but I came close. And overall, the process wasn’t too bad! It was helpful to make a census of material that was on my old site using access logs. This turned up a few redirects I wouldn’t have thought of, and also reminded me…
The data store for Comparea is a giant 23MB GeoJSON file. Most of the space in that file is taken up by the giant lists of coordinates which define the boundaries of each shape. But there’s also some interesting metadata hidden amongst all those latitudes and longitudes: { "features" : [ { "geometry" : { "type" : "Polygon" , "coordinates" : [ [ [ -69.89912109375001 , 12.452001953124963 ], […
There was a spike of traffic to Comparea over the weekend: Awesome! All of it came from Facebook and went to a comparison of France vs Australia : I can easily get insight into who tweeted it . But Facebook is a big black box. In theory, I can use Facebook Insights to track this. It claims that 50 actions have led to ~2400 visits to my site, but declines to say anything more: I understand that…
I’m going to try hosting my site and blog on GitHub pages. My hope is that blogging using GitHub and Markdown will lower the barrier to writing, and that GitHub pages will eliminate any worries about performance and security while hosting my own site. This is all very much a work in progress, so feedback is welcome!