A few years ago, I located a copy of John Guillory’s Cultural Capital in my local, used book store. Despite his former employ at New York University, I claim no bias here. I had heard of his book of field-specific criticism on canon formation, but that I found myself reading it was random chance. I also cannot claim that I have read it in its entirety, but for a few weeks afterward I found myself reading whole sections which seemed to have familiar overlaps with a problem I was confronting in my dissertation. There are digital, cultural landscapes that have been defined in the last thirty or so years that have natural relationships with the ones Guillory describes and critiques. When I wrote about the siting of our house/dataset within a literary landscape in my last post, I was not just referring to the physical extents of an individual work or collection, but rather a broader field of ideas in which the digital edition you are considering exists.
When one studies literature I would like to think that all of us are confronted with a particular question that can undercut the feeling that you are objectively approaching a work, a set of works, a period, etc. And that is: “How is it that I have come to find myself reading this particular work, edition, set of works, period, etc.?” This cognizance is a fundamental characteristic of literary criticism. Just as with a physical edition of a work, practicality often weighs down our consideration of if we are reading the best edition or version of a work. E.g. What edition(s) am I able to find? Part of my aim is to qualify for digital humanists what it actually means when we say we are working with “the best” edition or version of a work, and there are stipulations in that declaration too.
The first is that “the best” is not always necessary. If we are incorporating a work into a language model, for example, and our aim is to utilize its central tendency-related statistics, it might be that a small percent difference between the works will not matter much.
The second stipulation is that working out what “the best” edition/version of a work is is not straightforward. One sense that you should take caution of is the “prestige” of an edition. Sometimes this implicit understanding can bias our choices in favor of well-funded, institutional editions. As I will go into further detail in a later post, there are other metrics by which to judge the quality of an edition, and that includes not just quantifiable ones but more qualitative ones like accessibility.
And the third stipulation is that there is a distinction between a physical version and a digital edition (beyond the obvious corporeal differences). There can be multiple digital editions of the same physical version of a work. This last point is important to understand because one might come across digital editions of the same physical version of a work where each digital edition has slight differences that are imperative to document and consider. Digital editions are not throwaway objects; they are the very materials (and the closest ones) from which we base our modeling, analysis, and interpretations.
These considerations and the kind of comparisons and measurements that follow them open up digital humanities research to a new sensibility for a digital landscape. They bring forth concrete understandings of provenance and the qualities of the actual objects we are studying: digital editions. If you are already engaged in a study or are considering one, I would give you this exercise. Look at the digital edition of whatever it is you are basing your study from and take a hard look at it. Explore and note its characteristics. What are its bounds? What are its faults? Where is it from? When was this digital object created? Who created it and why? Not all of this information is always available, but you might be surprised. Perhaps it is time to reconsider other digital editions of the same physical object.
Digital objects found at what we tend to think of as open access sites with little prestigious or verifiable provenance have more bibliographic information than is initially obvious. Digital editions found on the Internet Archive and Project Gutenberg, for instance, both hold metadata on their creation. Internet Archive lists downloadable metadata alongside each digital edition they host. Project Gutenberg, often considered a source of texts without provenance, not only keeps a file trail of its editions but has a documented process for text creation on its main site and at its other site Distributed Proofreaders – where volunteers labor collaboratively, discussing and noting the progress of the creation of digital texts.
I will be delving deeper into this topic of the broadened digital landscape in future posts. It’s a new-ish area of study for the digital humanities. Borrowing a term from information science, I refer to it as “data quality”. But as you will see, that term expands into the perspectival when it meets humanistic spaces.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.