Building Bibliographic Superwork Clusters for Discovery with Local LLMs Judging work relationships in 88K clusters Aug 13 2026 Intro Links and Clusters Local LLM + Data Relationship Vocabulary Quality Evaluation OPAC Browse Interface Summary Intro In bibliographic description you are able to build relationships to external authorities and records in various ways. In MARC you can use controlled…
Jeopardy! Video Games of the 80s and 90s Analysis of 24K questions extracted Aug 3 2026 I have been looking at trivia game stuff recently and also been thinking about ROM emulation files, for old console gaming systems like Nintendo, as a data corpus. And it made me remember seeing these Jeopardy games around growing up and thinking how awful they must be. But now that I’m old I kind of…
Email Visualization Visualizing 14,000 Released Epstein Emails Nov 24 2025 Browsing the 20,000 Epstein documents released by the House Oversight Committee using tools people built such as this searchable interface or even network graph is interesting but still kind of opaque. I was thinking about what a browseable interface for thousands of emails would look like. If you have time as the x-axis…
LCNAF & Trie Storing +11M unique LCNAF names in 50MB Trie data structure Nov 12 2025 Intro Trie Data Structure & LCNAF Application - Client Lookup and MARC Reconciliation Application - Interface for Humans Application - OpenRefine, Simple API and Command Line Limitations Takeaways Intro The core of this post is looking at a pretty niche problem but it spans across a number of interconnected…
Giallo Using a vision language model to analyze Italian Giallo films Oct 31 2025 Intro Italian Giallo Films seem to be having a little resurgence recently. I first was exposed to them when the Criterion Channel app had a collection of them for Halloween a few years ago and now watch a few of them around this time of year mixed in with other spooky season movies. They are terrible movies, but are…
Book Bans 2025 Analysis of PEN America Book Ban Data Oct 17 2025 Intro PEN America released their Book Bans 2025 list which aggregates books that have been banned in US school districts for 2024-2025. They also kindly released the spreadsheet of data which lets me take the titles and author list and do some enrichment and analysis. I did something similar a couple years ago that can be found here…
Library of Congress & Flickr Commons Analysis of user interactions on 40,000 images Oct 6 2025 Intro Flickr Commons is a program to bring the visual collections of cultural heritage organizations to new audiences. Getting these resources in front of people where they are online as opposed to being siloed in their own website or not online at all. It was a pretty ground breaking project, the…
Building datasets from video collections using local & cloud LLMs Using Qwen2.5-VL, Gemini 2.5 and Whisper to build a Siskel and Ebert dataset Jul 30 2025 Intro I continue to experiment with LLM models finding very narrow tasks in my domain of cultural heritage where they could be useful and appropriate. I’m not going to go into it in this post but I find most pro-“AI” claims, discourse, and use…
Glitch Migration and Maintenance Jul 23 2025 Intro This month Glitch.com shut down, ending hosting for the projects there. They gave folks about a month to migrate their projects off the site. This also included paying customers, like me, who had the pro plan which cost $99 a year. A month is not a very long time at all, especially for paying customers, but that’s just like my opinion. The bottom…
HathiTrust 1929 Pubic Domain Data and tools to explore 50,000 1929/2025 public domain titles in HathiTrust Feb 5 2025 Intro Towards the end of the year HathiTrust quietly posts a collection of their resources that are set to enter the public domain in the new year. This is such a great resource and for the past few years I’ve done some visualistions or lists for 1925 , 1926 , 1927 and 1928 . This…
WoodBlockShop Using Segment Anything, LLaVA and other methods on a 14K image corpus Dec 16 2024 I came across people sharing the Plantin-Moretus Museum collection site and I really liked it. The 14,000 woodcuts are obviously all very different but the monochromatic cohesion of the collection ties it together. I thought it would be a great collection to try some of the various image based machine…
Banned Metadata Analysis of PEN America 2022-2023 Banned and Challenged Book Metadata Sep 27 2024 Intro I'm writing this at the end of Banned Book week 2024 , which in recent years has become much more poignant as there are literally books being banned in the US. PEN America released info that for 2023-2024 there is over 10,000 instances of books being challenged or banned . They haven’t released…
Lomax & Whisper.cpp Automatic transcription of Alan Lomax’s 1938 Midwest Folk Song Collection using Whisper.cpp Sep 12 2024 Sound to Data Whisper(.cpp) Alan Lomax Collection of Michigan and Wisconsin Recordings The Shape of Sound Many Models - Least Bad? Web Component Player LLM Enrichment Search Interface Mutable Metadata Now I'm going to sing a song like you're dying Code Links Sound to Data…
LCNAF Anagrams Apr 1 2024 LCNAF (Library of Congress Authority File) has over eleven million names of things ranging from people to geographic locations. I look at this data all the time, and it must have rotted my brain because I started thinking how there were probably a lot of names that if rearranged would spell other names in the file. Turns out there are—in just the personal and corporate…
Migrating Your Docker Wikibase Mar 20 2024 At Semantic Lab @ Pratt we were early-ish adopters of the Wikibase Docker distrubtuion (now deprecated and archived, replaced with the wikibase-release-pipeline ). It was nice to have a system that was ready to deploy with minimal configuration, we used version 1.31 for many years and it worked fairly well. Like all system maintenance eventually you need…
Using GPT on Library Collections Use cases for applying GPT3/3.5/4 on a full text collection March 30 2023 --> Intro Ground Rules Domain - Susan B Anthony Daybooks Correspondences Writings Comparing GPT3 / 3.5 / 4 Provenance Don’t call it Boutique Scale vs Voltage Should you use GPT? Conclusion Intro Here in early 2023 the breathless hype surrounding ChatGPT/GPT4 is honestly reaching annoying…
Jan 6th Overflow 2022 Watching the live stream of Trump's speech on Jan 6th 2021 on YouTube I was struck by the comment section on the video. Rendered pretty much useless by the sheer volume of comments flying up the window. It consisted of mostly Trump supporters shouting into the internet. Watching it trying to gauge some kind of zeitgeist of what was about to happen. At the time I lived only a…
Ley Lines An ongoing feed of drawing lines between things LinkNYC LinkNYC are kiosks throughout the street corners of New York City, they have big screens, wifi and do other computer stuff, like collect information on you, etc. I was thinking, what if you drew a line between all of them in a specific borough. You would end up with a big LinkNYC ley line. And of course if you found the center of…
Animated Gifs in US Elections Mining the Library of Congress web archives for political Gifs Dec 1 2021 --> Intro Problem Space Official Gifs User Content Gifs Getting Weird With User Content Gifs Technical - How to get started working with this dataset Intro A new dataset was announced by the prolific data releasing web archiving team and LC Labs at the Library of Congress. I really enjoy working…
Leeks: ParkMobile What's in a (car) name? June 9 2021 About the Leeks Series: This blog series uses publicly found dataset leaks, breaches and scrapes to do something interesting or fun. Personal identifying information will never be posted. Nor will any data be redistributed and you will never find any information on how to obtain the dataset. I will however use the data in aggregate for…
Maya Lin's Eclipsed Time Time telling in a pandemic Apr 12 2020 --> Repeated Hunt 2020, 7 minutes These layered 18th century hand colored engravings come from the collection of Caroline King Duer who gave them to the Cooper Hewitt. Duer was an author as well as an editor in the 1930s at Vogue, heading up the house furnishing section of the magazine. Hunting scenes are a common theme in 18th…
Maya Lin's Eclipsed Time Time telling in a pandemic Apr 12 2020 --> Example Tweet Example Tweet ISBN Uh-Oh Dec 2020, Twitter Bot ISBNs are numbers given books by publishers to help uniquely identify them. Every time a new book comes out or a new edition of a book (translation, 2nd edition, paperback, etc) it gets a new ISBN. These numbers are very useful for libraries, book sellers, book readers…
Maya Lin's Eclipsed Time Time telling in a pandemic Apr 12 2020 --> Ostraca 2020, 2 hours 34 minutes In the weeks after the 2020 presidential election you could stream live video of the Philadelphia vote counting process. You could watch as dozens of workers claid in safety vests sorted and counted ballots. This low paid or volunteer work is the invisible labor that makes the machinery of US…
Non-Renewed Copyright Analysis Digging into NYPL's digitized CCE Corpus Jul 22 2020 Intro Due to the bizarre nature of US copyright many works published in the mid-20th century are actually in the public domain. If a resource was published in a specific range of years and did not have its copyright renewed that material is likely open. The problem is determining this status because the records of…
Maya Lin's Eclipsed Time Time telling in a pandemic Apr 12 2020 “Time is broken” seems to be the thing everyone can agree with lately. The last few weeks folks I’ve talked to always seem to casually mention how they don’t know what day of the week it is. Co-workers, family, friends, students, it’s different for everyone, days feel too long, or seem too short, or a mixture of both. As we enter the…
Library of Congress Web Archives Big Picture Apr 23 2020 It was great to see the web archiving team at the Library of Congress recently celebrate twenty years of archiving. And they had a great article in the New York Times last month . I really like the LC web archive for a number of reasons. One is that it is heavily curated, meaning a lot of metadata is created for each site that is archived.…
Smithsonian Open Access Data Release A First Look Feb 28 2020 The Smithsonian Open Access release has put a trove of cultural heritage digital assets and data in the public commons. I wanted to show how you could start exploring and using the data. If you are comfortable running some Python scripts this post will quickly give you a head start. If not you might still be interested in following…
Repotting Old Digital Humanities Projects Two Test Cases Jan 31 2020 Knowing how to stop doing things is an important skill to learn. Kind of like learning to say no to things when you don’t have the capacity to do them. I've seen requirements around sustainability and sunsetting, that is how to wind down or retire a project become more prevalent on grant applications and funding processes. I've…
Hi there. I’m a librarian, researcher and developer. I work with cultural heritage organizations to build and implement technology solutions with an emphasis on metadata. My research focuses on linked open data, knowledge organization and data visualization. Quick Bio: I work for the Library of Congress helping with id.loc.gov and the Bibframe initiative. I previously worked for NYPL Labs and was…
Hi there. I’m a librarian, researcher and developer. I work with cultural heritage organizations to build and implement technology solutions with an emphasis on metadata. My research focuses on linked open data, knowledge organization and data visualization. Quick Bio: I work for the Library of Congress helping with id.loc.gov and the Bibframe initiative. I previously worked for NYPL Labs and was…
Hi there. I’m a librarian, researcher and developer. I work with cultural heritage organizations to build and implement technology solutions with an emphasis on metadata. My research focuses on linked open data, knowledge organization and data visualization. Quick Bio: I work for the Library of Congress helping with id.loc.gov and the Bibframe initiative. I previously worked for NYPL Labs and was…