I was talking on Bluesky[^1] about why I dislike the widespread use of alphabetical ordering for states on the y-axis of charts. There are better ways! My favorite is detailed in [this notebook](https://observablehq.com/@bmschmidt/useful-linear-orders-for-countries-and-states#linear_us_state_order), where I talk through some methods for treating paths. I have an interactive tool for building out…
Although I've given up on historically professing myself, I still have a number of automated scripts for analyzing the state of the historical profession hanging around. Since a number of people have asked for updates, it seems worth doing. [As a reminder](https://benschmidt.org/post/2020-10-01-jobs-update/), I'm scraping H-Net for listings. When I've looked at job ads from the American Historical…
Last week we released a big data visualiation in collaboration with the [Berens Lab](https://www.eye-tuebingen.de/berenslab/) at the University of Tübingen. It presents a rich, new interface for exploring an extremely large textual collection. Because I can I'll simply embed it below--but you'll have a better experience reading it [at the original site](https://static.nomic.ai/pubmed.html).…
Yesterday was a big day for the Web: [Chrome just shipped WebGPU without flags in the Beta for Version 113.](https://developer.chrome.com/blog/webgpu-release/) Someone on Nomic's GPT4All discord asked me to ELI5 what this means, so I'm going to cross-post it here---it's more important than you'd think for both visualization and ML people. (thread) So: GPUs are processors on basically every…
_This is a [Twitter thread from March 14](https://twitter.com/benmschmidt/status/1635692487258800128) that I'm cross-posting here. Nothing massively original below. It went viral because I was one of the first to extract the ridiculous paragraph below from on the release of GPT-4, and because it expresses some widely shared concerns. _ {.tweet} ::: I think we can call it shut on 'Open' AI: the…
Recently, Marymount--a small Catholic university in Arlington, Virginia--has been in the news for a draconian plan to eliminate a number of majors, ostensibly to better meet student demand. I recently learned the university leadership has been circulating one of my charts to justify the decision, so I thought I'd chime in on the context a bit. My understanding of the situation, primarily informed…
I sure don't fully understand how large language models work, but in that I'm not alone. But in the discourse over the last week over the Bing/Sydney chatbot there's one pretty basic category error I've noticed a lot of people making. It's thinking that there's some _entity_ that you're talking to when you chat with a chatbot. [Blake…
I attended the American Historical Association's conference last week, possibly for the last time since I've given up history professorin. Since then, the collapse of the hiring prospects in history has been on my mind more. See [Erin Bartram](https://contingentmagazine.org/2023/01/07/a-profession-if-you-can-keep-it/), [Kathryn…
The collapse of Twitter under Elon Musk over the last few months feels, in my corner of the universe, like something potentially a little more germinal; unlike in the various Facebook exoduses of the 2010s, I see people grasping towards different models of the architecture of the Web. Mastodon itself (I've ended up at [@benmschmidt@vis.social](https://vis.social/@benmschmidt) for the time being)…
I'm excited to finally share some news: I've resigned my position on the NYU faculty and started working full time as Vice President of Information Design at [Nomic](https://nomic.ai), a startup helping people explore, visualize, and interact with massive vector datasets in their browser. This will be a big shift. I've spent my whole career up to this point in academic institutions; but right now,…
When you teach programming skills to people with the goal that they'll be able to _use_ them, the most important obligation is not to waste their time or make things seem more complicated than they are. This should be obvious. But when I'm helping humanists decide what workshops to take, reviewing introductory materials for classes, or browsing tutorials to adapt for teaching, I see the same…
It's not very hard to get individual texts in digital form. But working with grad students in the humanities looking for large sets of texts to do analysis across, I find that larger corpora are so hodgepodge as to be almost completely unusable. For humanists and ordinary people to work with large textual collections, they need to be distributed in ways that are actually *accessible, not just open…
{.thread} ::: {.tweet} ::: I've never done the "Day of DH" tradition where people explain what, exactly, it means to have a job in digital humanities. But today looks to be a pretty DH-full day, so I think, in these last days of Twitter, I'll give it a shot. (thread) ::: {.tweet} ::: We'll start it at the beginning--1:30 or so AM, finally sent out an e-mail I'd been procrastinating on to the…
 There are programming languages that people use for money, and programming languages people use for love. There are Weekend at Bernie's/Jeremy Bentham corpses that you prop up for the cash, and there are "Rose for Emily" corpses you sleep with every night for decades because it's too painful to admit that the best version of your…
I've been spending more time in the last year exploring modern web stacks, and have started evangelizing for [SvelteKit](https://kit.svelte.dev/), which is a new-ish entry into the often-mystifying world of web frameworks. As of today, I've migrated this, personal web site from Hugo, which I've been using the last couple years, to sveltekit. Let me know if you encounter any broken links,…
Scott Enderle is one of the rare people whose [Twitter pages](https://twitter.com/scottenderle) I frequently visit, apropos of nothing, just to read in reverse. A few months ago, I realized he had at some point changed his profile to include the two words "increasingly stealthy." He had told me he had cancer months earlier, warning that he might occasionally drop out of communication on [a…
This [article in the New Yorker](https://www.newyorker.com/magazine/2021/03/15/genre-is-disappearing-what-comes-next) about the end of genre prompts me to share a theory I've had for a year or so that models at Spotify, Netflix, etc, are most likely not just removing artificial silos that old media companies imposed on us, but actively destroying genre without much pushback. I'm curious what you…
I've been yammering online about the distinctions between different entities in the landscape of digital publishing and access, especially for digital scholarship on text. So I've collected everything I've learned over the last 10 years into one, handy-to-use, chart on a 10-year-old meme. The big points here are: 1. HathiTrust and JSTOR are not for-profit cartels, and I can't count the number of…
I mentioned [earlier](http://benschmidt.org/post/2021-03-07-bookworm-caching/bookworm-caching/) that I've been doing some work on the old Bookworm project as I see that there's nothing else that occupies quite the same spot in the world of public- facing, nonconsumptive text tools. That codebase is *old*--pieces of it [date back to this blog post from a decade…
 I've recently been getting pretty far into the weeds about what the future of data programming is going to look like. I use pandas and dplyr in python and R respectively. But I'm starting to see the shape of something that's interesting coming down the pike. I've been working on a project that involves scatterplot visualizations at a massive scale--up to 1 billion…
I used to blog _everything_ that I did about a project like [Bookworm](https://github.com/bookworm-project), but have got out of the habit. There are some useful changes coming through through the pipeline, so I thought I'd try to keep track of them, partly to update on some of the more widely used installations and partly The core work on Bookworm happened in 2011-2013 when I was at Harvard…
I last looked at the H-Net job numbers [about a month ago.](/post/2020-10-01-jobs-update/2020-10-01-jobs-update/) Since then, the news isn't exactly good, but it's also probably as good as anyone could expect. For most of September and October, history jobs were at about 25% of their average for the 2010s; this was slightly worse than we're seeing in the approximate numbers in--for…
Out of a train-wreck curiosity about what's been happening to the historical profession, I've been watching the numbers on tenure-track hiring as posted on H-Net, one of the major venues for listing history jobs. \[Update 10-2: switching to US and Canada only. An earlier version of this included other countries, even though I said it didn't.\] We're now into October. Usually--I know now--this is…
I've been doing a lot of my data exploration lately on Observable Notebooks, which is--sort of--a Javascript version of Jupyter notebooks that automatically runs all the code inline. Married with Vega-Lite or D3, it provides a way to make data exploration editable and shareable in a way that R and python data code simply can't be; and since it's all HTML, you can do more interesting things. Of…
Every year, I run the numbers to see how college degrees are changing. The Department of Education released this summer the figures for 2019; these and next year's are probably the least important that we'll ever see, since they capture the weird period as the 2008 recession's shakeout was wrapping up but before COVID-19 upended everything once again. But for completism, it's worth seeing how…
{#ranking-graduate-programs} # Ranking Graduate Programs While I was choosing graduate programs back in 2005, I decided to come up with my own ranking system. I had been reading about the Google PageRank algorithm, which essentially imagines the web as a bunch of random browsing sessions that rank pages based on the likelihood that you--after clicking around at random for a few years--will end up…
As I often do, I'm going to pull away from various forms of Internet reading/engagement through Lent. This year, this brings to mind one of my favorite stray observations about digital libraries that I've never posted anywhere. As part of the 2016 Republican Primary, Jeb! Bush released a website enabling exploration of e-mails related to his official accounts as governor of Florida in the early…
Since 2010, I've done most of my web hosting the way that the Internet was built to facilitate: from a computer under the desk in my office. This worked extremely well for me, and made it possible to rapidly prototype a lot of of websites serving large amounts of data which could then stay up indefinitely; I have a curmudgeonly resistance to cloud servers, although I have used them a bit in the…
Some news: in September, I'll be starting a new job as Director of Digital Humanities at NYU. There's a wide variety of exciting work going on across the Faculty of Arts and Sciences, which is where my work will be based; and the university as a whole has an amazing array of programs that might be called "Digital Humanities" at another university, as well as an exciting new center for Data…
Critical Inquiry has posted an article by Nan Da offering [a critique of some subset of digital humanities](https://www.journals.uchicago.edu/doi/abs/10.1086/702594) that she calls "Computational Literary Studies," or CLS. The premise of the article is to demonstrate the poverty of the field by showing that the new structure of CLS is easily dismantled by the master's own tools. It appears to have…
I wrote this year's report on history majors for the American Historical Association's magazine, _Perspectives on History_; it takes a medium term view of at the significant hit the history major has taken since the 2008 financial crisis. You can read it…
As part of the _Creating Data_ project, I've been doing a lot of work lately with interactive scatterplots. The most interesting of them is [this one about the full Hathi collection](https://t.co/erWeUkR9Fk). But I've posted a few more I want to link to from here: - [An exploration of co-occurring street names in the United States](http://creatingdata.us/etc/streets/) - A [description of the…
I have a new article on dimensionality reduction on massive digital libraries this month. Because it's a technique with applications beyond the specific tasks outlined there, I want to link to a few things here. - [The article](http://culturalanalytics.org/2018/09/stable-random-projection-lightweight-general-purpose-dimensionality-reduction-for-digitized-libraries/) in _Cultural Analytics_. - [A…
I'm switching this site over from Wordpress to Hugo, which makes it easier for me to maintain. It may also confuse the RSS feed a bit. This should be hopefully be a one-time occurrence.
I have a [new article in the Atlantic](https://www.theatlantic.com/ideas/archive/2018/08/the-humanities-face-a-crisisof-confidence/567565/) about declining numbers for humanities majors.
. In short, it's been bad enough to make me recant earlier statements of mine about the long-term health of the humanities discipline.
This is some real inside baseball; I think only two or three people will be interested in this post. But I'm hoping to get one of them to act out or criticize a quick idea. This started as a comment on Scott Enderle's blog, but then I realized that Andrew Goldstone doesn't have comments for the parts pertaining to him... Anyway. Basically I'm interested in feature reduction for token-based…
I've gotten a couple e-mails this week from people asking advice about what sort of computers they should buy for digital humanities research. That makes me think there aren't enough resources online for this, so I'm posting my general advice here. (For some solid other perspectives, see here). For keyword optimization I'm calling this post "digital humanities.” But, obviously, I really mean the…
Practically everyone in Digital Humanities has been posting increasingly epistemological reflections on [Matt Jockers' Syuzhet package](https://github.com/mjockers/syuzhet) since Annie Swafford posted a [set of critiques of its assumptions](https://annieswafford.wordpress.com/2015/03/02/syuzhet/). I've been drafting and redrafting one myself. One of the major reasons I haven't is that the…
Just some quick FAQs on my[ professor evaluations visualization](http://benschmidt.org/profGender): adding new ones to the front, so start with 1 if you want the important ones. -3 (addition): The largest and in many ways most interesting confound on this data is the gender of the _reviewer_. This is not available in the set, and there is strong reason to think that men tend to have more men in…
I promised Matt Jockers I'd put together a slightly longer explanation of the weird constraints I've imposed on myself for topic models in the Bookworm system, like t[hose I used to look at the breakdown of typical TV show episode structures.](http://sappingattention.blogspot.ca/2014/12/typical-tv-episodes-visualizing-topics.html) So here they are. The basic strategy of Bookworm at the moment is…
Just a quick follow-up to my [post from last month on using Markdown for writing lectures](http://benschmidt.org/2014/09/05/markdown-historical-writing-and-killer-apps/). The [github repository for implementing this strategy is now online](https://github.com/bmschmidt/MarkdownLectures). The goal there was to have one master file for each lecture in a course, and then to have scripts automatically…
I've been thinking a little more about how to work with the [topic modeling extension](https://github.com/bmschmidt/Bookworm-Mallet) I recently built for bookworm. (I'm curious if any of those running installations want to try it on their own corpus.) With the movie corpus, it is most interesting split across _genre;_ but there are definite temporal dimensions as well. [As I've said…
I've been seeing how deeply we could integrate topic models into the underlying Bookworm architecture a bit lately. My own chief interest in this, because [I tend to be a little wary of topic models in…
This is a post about several different things, but maybe it's got something for everyone. It starts with 1) some thoughts on why we want comparisons between seasons of the Simpsons, hits on 2) some previews of some yet-more-interesting Bookworm browsers out there, then 3) digs into some meaty comparisons about what changes about the Simpsons over time, before finally 4) talking about the internal…
Like many technically inclined historians (for instance, [Caleb McDaniel](http://wcm1.web.rice.edu/hacks.html), [Jason Heppler](http://jasonheppler.org/2012/11/20/using-markdown-like-an-academic/), and [Lincoln Mullen](http://chronicle.com/blogs/profhacker/markdown-the-syntax-you-probably-already-know/35295)) I find that I've increasingly been using the plain-text format Markdown for almost all of…
I thought it would be worth documenting the difficulty (or lack of) in building a Bookworm on a small corpus: I've been reading too much lately about the Simpsons thanks to the FX marathon, so figured I'd spend a couple hours making it possible to check for changing language in the longest running TV show of all time. For some thoughts on how to build a bookworm, read "prep”: otherwise, skip to…
Here's a very technical, but kind of fun, problem: what's the optimal order for a list of geographical elements, like the states of the USA? If you're just here from the future, and don't care about the details, here's my favorite answer right now: But why would you want an ordering at all? Here's an example. In the baby name bookworm, if you search for a name, you can see the interaction of…