(This is a post I've had unpublished since writing it in 2016. Just hitting publish without reviewing right now because it's something I find myself periodically looking at the charts for). As we get ready to launch the full Hathi Trust+Bookworm to allow tracking words across 13 million books, I've been working on fixing up the metadata from the original MARC records. This is useful information to…
I periodically write about Google Books here, so I thought I'd point out something that I've noticed recently that should be concerning to anyone accustomed to treating it as the largest collection of books: it appears that when you use a year constraint on book search, the search index has dramatically constricted to the point of being, essentially, broken. Here's an example. While writing…
I did a slightly deeper dive into data about the salaries by college majors while working on my new Atlantic article on the humanities crisis . As I say there, the quality of data about salaries by college major has improved dramatically in the last 8 years. I linked to others' analysis of the ACS data rather than run my own, but I did some preliminary exploration of salary stuff that may be…
NOTE 8/23: I've written a more thoughtful version of this argument for the Atlantic. They're not the same, but if you only read one piece, you should read that one. Back in 2013, I wrote a few blog post arguing that the media was hyperventilating a bout a "crisis" in the humanities, when, in fact, the long term trends were not especially alarming. I made two claims them: 1. The biggest drop in…
Historians generally acknowledge that both undergraduate and graduate methods training need to teach students how to navigate and understand online searches. See, for example, this recent article in Perspectives . Google Books is the most important online resource for full-text search; we should have some idea what's in it. A few years ago, I felt I had some general sense of what was in the Books…
Matthew Lincoln recently put up a Twitter bot that walks through chains of historical artwork by vector space similarity. https://twitter.com/matthewdlincoln/status/1003690836150792192. The idea comes from a Google project looking at paths that traverse similar paintings. This reminds that I'd meaning for a while to do something similar with words in an embedding space. Word embeddings and image…
This is a blog post I've had sitting around in some form for a few years; I wanted to post it today because: 1) It's about peer review, and it's peer review week ! I just read this nice piece by Ken Wissoker in its defense. 2) There's a conference on argumentation in Digital History this weekend at George Mason which I couldn't attend for family reasons but wanted to resonate with at a distance.…
Digging through old census data, I realized that Wikipedia has some really amazing town-level historical population data, particularly for the Northeast, thanks to one editor in particular typing up old census reports by hand. (And also for French communes, but that's neither here nor there.) I'm working on pulling it into shape for the whole country, but this is the most interesting part.…
I've been doing a lot of reading about population density cartography recently. With election-map cartography remaining a major issue, there's been lots of discussion of them: and the " Joy Plot " is currently getting lots of attention. So I thought I'd finally post some musings I wrote up last month about population density, the built environment, and this plot I made of New York City building…
Robert Leonard has an op-ed in the Times today that includes the following anecdote: Out here some conservatives aren’t even calling them “public” schools anymore. They call them “government schools,” as in, “We don’t want to pay for your damn ‘government schools.’ ” They’re afraid to send their kids to them. I'm pretty interested in the process of objects shifting from belonging to the "public"…
The Library of Congress has released MARC records that I'll be doing more with over the next several months to understand the books and their classifications. As a first stab, though, I wanted to simply look at the history of how the Library digitized card catalogs to begin with. A couple notes for the technically inclined: 1. the years are pulled from field 260c (or if that doesn't exist or is…
One of the interesting things about contemporary data visualization is that the field has a deep sense of its own history, but that "professional" historians haven't paid a great deal of attention to it yet. That's changing. I attended a conference at Columbia last weekend about the history of data visualization and data visualization as history. One of the most important strands that emerged was…
I want to post a quick methodological note on diachronic (and other forms of comparative) word2vec models. This is a really interesting field right now. Hamilton et al have a nice paper that shows how to track changes using procrustean transformations: as the grad students in my DH class will tell you with some dismay, the web site is all humanists really need to get the gist. Semantic shifts from…
This is a quick digital-humanities public service post with a few sketchy questions about OCR as performed by Google. When I started working intentionally with computational texts in 2010 or so, I spent a while worrying about the various ways that OCR--optical character recognition--could fail. But a lot of that knowledge seems to have become out of date with the switch to whatever post-ABBY,…
Like everyone else, I've been churning over the election results all month. Setting aside the important stuff, understanding election results temporally presents an interesting challenge for visualization. Geographical realignments are common in American history, but they're difficult to get an aggregate handle on. You can animate a map, but that makes comparison through time difficult. ( One with…
I'm pulling this discussion out of the comments thread on Scott Enderle's blog , because it's fun. This is the formal statement of what will forever be known as the efficient plot hypothesis for plot arceology . Noble prize in culturomics, here I come. Brief background: Enderle shows pretty persuasively that all the fundamental plot arcs described in a paper by a math-based computational story lab…
Word embedding models are kicking up some interesting debates at the confluence of ethics, semantics, computer science, and structuralism. Here I want to lay out some of the elements in one recent place that debate has been taking place inside computer science. I've been chewing on this paper out of Princeton and Bath on bias and word embedding algorithms. (Link is to a blog post description that…
Debates in the Digital Humanities 2016 is now online, and includes my contribution, "Do Digital Humanists Need to Understand Algorithms?" (As well as a pretty snazzy cover image …) In it I lay out distinction between transformations, which are about states of texts, and algorithms, which are about processes. Put briefly: Put simply: digital humanists do not need to understand algorithms at all.…
Some scientists came up with a list of the 6 core story types . On the surface, this is extremely similar to Matt Jockers's work from last year . Like Jockers, they use a method for disentangling plots that is based on sentiment analysis, justify it mostly with reference to Kurt Vonnegut, and choose a method for extracting ur-shapes that naturally but opaquely produces harmonic-shaped curves.…
I usually keep my mouth shut in the face of the many hilarious errors that crop up in the burgeoning world of datasets for cultural analytics, but this one is too good to pass up. Nature has just published a dataset description paper that appears to devote several paragraphs to describing "center of population" calculations made on the basis of a flat earth. "Spatializing 6,000 years of global…
I started this post with a few digital-humanities posturing paragraphs: if you want to read them, you'll encounter them eventually. But instead let me just get the point: here's a trite new category of analysis that wouldn't be possible without distant reading techniques that produces sometimes charmingly serendipitous results. I'll call it dopplegänger books. A dopplegänger is, for any…
A heads-up for those with this blog on their RSS feeds: I've just posted a couple things of potential interest on one of the two other blogs (errm) I'm running on my own site. One, " Vector Space Models for the digital humanities ," describes how a newly improved class of algorithms known as word embedding models work and showcases some of their potential applications for digital humanities…
Mitch Fraas and I have put together a two-part interactive for the Atlantic using Bookworm as a backend to look at the changing language in the State of Union. Yoni Appelbaum, who just took over this week, spearheaded a great team over there including Chris Barna, Libby Bawcombe, Noah Gordon, Betsy Ebersole, and Jennie Rothenberg Gritz who took some of the Bookworm prototypes and built them into a…
Far and away the most interesting idea of the new government college ratings emerges toward the end of the report. It doesn't quite square the circle of competing constituencies for the rankings I worries about in my last post, but it gets close. Lots of weight is placed on a single magic model that will predict outcomes regardless of all the confounding factors they raise (differing pay by…
Before the holiday, the Department of Education circulated a draft prospectus of the new college rankings they hope to release next year. That afternoon, I wrote a somewhat dyspeptic post on the way that these rankings, like all rankings, will inevitably be gamed. But it's probably better to bury that off and instead point out a couple looming problems with the system we may be working under soon.…