full copy of my slides with speaker text as annotations
Good afternoon, everyone. When finalizing a date for this, I really did not think enough about how it’s only a week after the national elections and how the fraught leadup and outcomes could impact my writing this or your attending. So thank you for being here, or watching after. I hope that what I’ve brought will be useful to you in some ways, even though I know some of us will not be in the best mental place for it.
As Marianne said when introducing me, I work on discovery, the library catalog, and linked data projects at Penn State. My official title is “Cataloging Systems and Linked Data Strategist.” It sounds either exciting or faddish, but I appreciate being able to lean on the “strategist” part of my title. My job isn’t to “make linked data happen” on some efficient timeline. Instead, I evaluate and decide what would actually be meaningful for us to focus on and what isn’t, or isn’t currently, going to be useful to Penn State, given our local needs, goals, and capacity.
I’ve been in this role for seven years now, listening, thinking, writing, experimenting, and evaluating, and I have quite a few thoughts. I’m not going to pitch you the idea that linked data is just around the corner or that AI is coming and when that happens we won’t have to worry about cataloging any more. I do think that we’re coming closer to having cataloging tools which support graph-oriented and BIBFRAME-informed cataloging. That’s not the same as having a true linked data ecosystem, but I can see a real benefit in having new data creation tools which prioritize the use of controlled values and terms and reduce some kinds of data entry errors.
But while I’m not exactly bullish on a brilliant linked data future being just around the corner, I love the rare glimpses of what happens when we have collaboratively-built, interlinked nodes, although they’re still rarely functional outside single systems.
When it comes to linked data, I think we’re in what I’d describe as a period of long liminality. The word “liminal” comes from the Latin word for threshold. It’s been used to describe various kinds of transitions – the point at which something becomes perceptible or the period around a rite of passage from one status to another, such as childhood to adulthood.
Today I’m thinking of another variation of how it’s used, in what are sometimes called “liminal spaces” – thresholds where we’re existing between two different things. These are often depicted as something eerie, an empty street or a hotel hallway. They’re not our starting point or our destination, but somewhere in between or unsettling. And liminal spaces don’t have to be empty to leave us disoriented.
You know that point in waiting for a plane where you’ve been in the airport, in the gate for hours, your flight’s been pushed back several times? Things start to feel a little surreal, like nowhere exists except the airport and you’re never going to get out of it? Sometimes, it feels like that’s where we are when we talk about linked data. I still remember back in 2019 at an LD4 conference we got to talking about “Linked data fatigue,” how you keep hearing about it, hearing that it’s coming, but you’re not there. Someone shared that they’d heard about it in library school 10 years before and they felt like it was still far in the future.
Five years later we’ve seen a lot more progress toward things like BIBFRAME actually be usable in daily cataloging but, at best, we’re still in a “between” space where the average cataloger at the average library is only going to encounter it in workshops and theoretical spaces. And if you’ve been hearing about it for 15 years now, the fatigue is real and it’s reasonable. I know people who have taken the position that they’ll get back into it when the message is, “we’re starting linked data cataloging tomorrow.” They have no interest in being dragged back into that airport, full of uncertainty and repeated reschedulings. That makes sense to me.
So when Marianne asked me about speaking, I clarified that the group wasn’t interested in my coming in to hype things up. But I do have thoughts about what we do in this liminal space and how we assess whether we should actually start gathering our bags and getting ready to board the metaphorical plane. I shared some of those thoughts in 2022, when I spoke at the Potomac Technical Processing Librarians annual meeting about BIBFRAME implementation. Specifically, I outlined the key things we would need to have in place for the profession to move from our current practice to a BIBFRAME environment. I identified which ones we had or partly had and which we were still lacking. I’ll share a link at the end. So, good news I guess, I’m not going to talk much about BIBFRAME specifically here because I feel like I’ve already said what I have to say on that.
And since I gave that talk, we’ve seen the eruption of generative AI, or what’s marketed as AI, along with a whole set of new promises about what they’ll mean for libraries and for library data. Most of those tools are right now are pattern recognition and replication tools. They’re certainly extraordinary. They can be quite useful.
Leaving aside other very important critiques, I think they’re actually in a similar place right now to linked data. We hear pitches for how transformative they could be to our workflow or to discovery. We see interesting prototypes. But for day-to-day work they’re still lacking enough, say, in generating a catalog record, that a trained cataloger would have to review every single field (including and especially the fixed fields) to make an LLM-generated MARC record ready for prime time. If copy already exists, you’re going to spend a lot less time and effort (and cause less environmental stress) to pull in a record from Z39.50 or the WorldCat database that doesn’t have a dozen weird pattern replication errors in it.
Pattern recognition and replication is not the same as skill or overall knowledge. And it’s certainly not the same as judgment, insight, or discernment. These are qualities that we still need in libraries – maybe even more than we did before our spaces were spinning up with predictive-text realities.
But even as LLMs and more useful forms of machine-learning added a new facet to our vision of the future, I think that’s what it really is, a facet. Because when we come back to actual problems and questions we’re trying to solve in library descriptive work, a lot of them don’t actually need big, power-hungry systems. Much of the time, they don’t even need linked data.
In fact, for all its flaws, MARC and the LMS, or even the traditional ILS, can handle a lot of the basics:
I then want a smooth path to getting my hands on it or getting it onto my device.1
Even some of the things I’ve heard thrown around for linked data are actually things that you can do with MARC, or could do if you had better MARC (which is something I’m going to come back to), but we don’t do because of our systems, capacity, and time. If the developers and I were given time for it, we could build a catalog index that not only let you see, say, every film that Vincent Price had been in but gave you a list of his most frequent costars (at least from the materials we have), neatly linked to searches for their work, or even a social graph showing which of them had appeared regularly together.

We just don’t because that’s not the kinds of questions we’re most commonly asked and “realize all Ruth’s fun ideas” is not a strategic goal of my institution.
I’m not saying this because I think MARC is the future of library bibliographic description, but because I think before I say anything else, it’s important to keep in mind that we2 build systems to do things and we don’t build them to do other things, even if we could have chosen to do so. Our current systems are not a representation of what’s fully possible in the present moment. And when we build new systems, making sure they do everything isn’t necessarily a good idea. For example, I mentioned that people want to find books by titles. Think of the frustration caused by book reviews in our discovery systems. That’s not an inherent problem of an everything-index, but the vendors running them didn’t put in the care to skew the index toward ranking books above book reviews when the two kinds of records match on title keywords.
Where this gets back to the data is that when we want to build systems to do things or answer questions, it’s informed by data. It’s informed by the data’s:
And then our outcome is informed by whether the system we use is the right one for the purpose and what we program those systems to do with the data.
To touch back again on LLMs, Emily Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell3 coined the term Stochastic Parrot in 2021 as a way of describing the ways that LLMs repeat back the most probable or common sequence of words.
As a simple example, in a recent church committee meeting, someone asked GPT-4 where the 2025 meeting of Mennonite Church USA was going to be held. It said Kansas City. Several folks nodded. The assertion made sense to everyone. The biennial meeting had been held there in 2019 and 2023, with a 2021 being a hybrid event with the in-person in Cincinnati. Given the history of recent repeats, it would be entirely reasonable for it to be held there in 2025 as well.
The statistical prediction made sense. The LLM would have ingested conference websites, blog posts, articles, retrospectives, and the like, maybe only a few dozen actual objects, the most recent of which would have said – Mennonite Church USA meeting? Associate that with Kansas City! But annual meetings are shaped by more than statistical probability.
Suspicious of this answer, I used DuckDuckGo, which sent me to the actual conference website, which said nope, next time it’s being held in Greensboro, North Carolina.
A few weeks ago, some friends and I were discussing outputs of Alma’s AI assistant – I’m not at an Alma library, but someone who is was graciously sharing her experimentation – and we noticed something similar. While the tool helpfully transcribed some of the actual information from photos she entered, it also generated some statistically probable information for fields like the 300:
xviii, 195 pages : ǂb illustrations ; ǂc 24 cm.
Whereas the actual data was:
ix, 189 pages : ǂb maps ; ǂc 25 cm
Experience tells us that illustrations are more common than maps. For the other two, I queried my own catalog and found that a dimension of 24 cm was more than twice as common as 25 cm. 195 pages was slightly more common. Like most LLM tools and Ex Libris products, it’s a bit of a black box, so we can only know so much about how it’s trained to choose between common sizes, pages, etc., but perhaps that’s a little insight.
For purpose, AI-type tools are pretty good at getting a bunch of transcription done. Additional pattern-recognition training can make them good at selecting title, subtitle, and author info from an image because publishers place those in similar locations on materials and emphasize them using font sizes or letter positions. But we’re also describing actual real world objects, not their most statistically probable forms, so such a system can only assist to a certain point.
I’ve been impressed by work a friend has done on her own tool to generate stub records from scans of an uncataloged batch of low-priority materials. It’s a different way of tackling something we did a decade ago with student workers typing info about some uncataloged microfilm into a spreadsheet and generating MARC. But again, it goes as far as it goes and can’t go farther because it can’t examine the material.
Or LLMs are known to struggle with what’s sometimes called the problem of Tom Cruise’s mother. They can tell you the name of Tom Cruise’s mother but not of Mary Lee Pfeiffer’s son. That’s because, in the texts they’ve processed, people describe her as his mother but don’t describe him as her son.4 Unlike the previous case where the system was returning the most common kind of assertion, in this case, the data doesn’t exist. Or rather, it doesn’t exist formatted in a way that the system has been designed to process and replicate. Even though the same LLM could probably tell you that mother and son are a reciprocal relationship and that, as a man, Tom Cruise would be described as a son vs. as a daughter, they are not currently set up to do an additional level of inferencing.
On the other hand, this is a case where a well-designed and well-implemented linked data platform and data would shine. Not because a linked data platform “knows” what a mother or son is any more than an LLM does, but because instead of deriving the most statistically-likely scenarios, it’s programmed for these kinds of relationships – reciprocity, subclass, and inference.
A really great example of transitive relationships would be a Wikidata query that I was able to help a friend set up for her repository. She’d gone through and applied the “archives at” property to Wikidata records for the people whose papers were held where she works. When she was looking for ways she could use Wikidata to understand her holdings better, I set up a query so that she could identify all the people whose archives were at her institution and who had attended HBCUs, just as an example of one way she could explore it.
Now, their records don’t say they went to HBCUs, they say that this person went to Howard University and that person went to North Carolina Central University. And then the records for those institutions say that they’re instance of HBCUs. Each of these statements were probably assigned by different people – one person adding the info that Alice Walker was educated at Spelman and another person adding the info that Spelman is an HBCU. When such data exists and when it’s accurate, there’s an extraordinary power in chaining associations, better than the LLMs do it and with more kinds of info than we can have in our MARC.
As we think about the descriptive work we do, what we can do now, and what our future might hold, I think the first thing for us to keep in mind is that:
our data always has been and always will be multi-faceted.
Sometimes our systems only need minimal data – like a title to match search strings, a creator to display so the patron can confirm it’s the object they’re looking for, and then system-managed data to help them obtain it. Sometimes we benefit from data pattern analysis to solve bigger problems. And sometimes what we need to answer questions is really thorough, structured, and well-defined data.
So wherever we’re going, we need good data.
As Marianne mentioned in my introduction, my current research is actually focusing on library systems and the people who use and maintain them. More concretely, I’m currently wrapping up a sabbatical in which I’ve been studying the impacts of ILS/LMS migration on library workers. You may have seen or even participated in a survey I sent out last spring or been one of my interviewees. I think it’s an understudied area that actually impacts workers today.
One of the recurring themes in both survey responses and interview conversations was the complicated nature of legacy data. Excepting perhaps the Innovative/III sequence of INNOPAC, Millennium, and Sierra, no two ILS/LMSes structure data precisely same way. Nor do the variety of front-end systems present it the same way. It’s all well and good and quite useful to add an 856 to your printed government records until it suddenly isn’t because somewhere in the migration workflow every record with an 856 second indicator 0 got transformed into Alma’s e-resource specific format.
And so two things I’ve heard over and over again from survey respondents and interviewees are:
While an ILS migration is a very different kind of transition, I think the idea of transitioning to linked data structures is related and thinking about the former can help orient us toward what is useful to do be doing during our long liminality.
Based on these two ILS migration data regrets, I’m going to spend the rest of my talk with some thoughts about:
So first, I want to talk about approaches, or perhaps “mindsets,” though that’s not word I love. When we undergo an entire system shift, it often requires us to reorient parts of our work. In ILSes, this might be where we get our records, like Alma’s community zone, or how they’re structured into bibs, holdings, and items, when previously we’d only done holdings for serials. When it comes to a shift to linked data description, I think there are three things, approaches that, to some degree, we already take, that are going to inform our work even more than they do now. These are:
The appeal of linked data as an underlying technology for description, as opposed to something as straightforward as MARC, is its focus on relationships. Not all of bibliographic description is relational, of course. For example, we have a page count or other kind of length. That’s a datapoint, but not a relationship. Abstracts are abstracts, helpful for keywords and for evaluating the work once you’ve found it. But a book has a publisher relationship as well as author and possibly illustrator and editor relationships. A film has relationships to actors, directors, producers, etc.
While our MARC systems are designed to search for the simplest of relationships, all films with a particular actor or all books by a particular author, the nature of a linked data system should allow one to traverse more relationships. As I said to the Potomac Technical Processing Librarians, this is the kind of make or break of a BIBFRAME system. We change how we encode our data, but we will also need systems which make use of it. Let me search all the books by a person who’s won a Hugo award, for example.
Thinking relationally doesn’t mean stuffing every relationship we can think of into a particular record. I have encountered relational thinking gone wrong – at a previous job, I was working on a collection of digitized architectural lantern slides and records created for them by a contracted visual resources cataloger. These slides were from the end of the 19th century, yet I found the keyword “Stalin” on several images of Russian churches. You see, the person doing this job had attempted to be exceptionally thorough. Sure, Stalin was a child at the time the image was taken, but some thirty or forty years later, he’d destroyed the church depicted. Well, I doubt he had personally, or even necessarily ordered it, but he’d set the process in motion (I avoided going on a deep dive to find out if there was any specific speech or legislation behind the destruction of churches in this era). I ended up doing massive cleanup on the keywords, paring them down to descriptions of the images themselves. And it’s true, our system had no way of letting you find images of churches destroyed during Stalin’s leadership of the USSR.
If we were recording as linked data statements, and please note this is just a mockup vs. a recommendation:
And, again, the system would have to let us search for a list of churches destroyed during Stalin’s period. That might be done by dates and locations or by some kind of event statement regarding the implementation of his orders to destroy Orthodox churches. We can see flashes of how this might be accomplished in the wikimedia system, both Wikidata and the ways it’s integrated into things like Wikipedia or Wikimedia Commons.
Thinking relationally isn’t some giant pivot from what we’re already doing. More, it’s an awareness that what we’re recording aren’t a series of standalone descriptions but something which we could transform into a traversable network. It’s a reason to strive for completeness, we can only traverse relationships which are already recorded, but it also should prompt us to ask whether this record is the right place for this relationship to be recorded.
Next, we have working collaboratively. Since the earliest days of library automation and computerized description,catalogers have worked collaboratively, perhaps moreso than any other kind of library worker. It makes sense – if we all have a copy of Tony Horowitz’s Midnight Rising, do we all need to retype the same record?
Early on, we loaded MARC tapes from the Library of Congress or used pre-internet network technologies to pull in records from RLIN and OCLC. Many places still use OCLC but may also pull in and augment records using Z39.50 or use Community Zone records. Some of us load large batches of files for electronic resources into the catalog while others activate collections in their ERM systems.
MARC was designed to produce fairly small files, especially with the practices we were using around its start. This was critical when systems might have a few megabytes of storage space – which cost them hundreds of thousands of dollars. Modern computers have allowed us to think comparatively expansively. We’ve used this to extend what’s in our records, more than 3 subjects, long text blurbs, etc. When BIBFRAME was designed, it didn’t really consider resource constraints. While the 9 million records in my catalog can be exported into a file that’s about 12GB, the same records transformed into BIBFRAME come to over a hundred GB. Now, there’s a difference between a full export and how the data’s actually stored, but it’s certainly going to expand when we move out of the hyper-efficient MARC format.
I expect a move to more centralized description to be driven as much by this desire for storage efficiency as by the desire to reduce duplicated work. But we’ve also seen downsides, whether it’s something wild in an OCLC record or a case I know of where person who kept editing Community Zone records to re-add a deprecated term, causing immense frustration to catalogers elsewhere.
People have experimented with ways of mitigating these downsides. In traditional MARC, we have protected fields we don’t import or export – nobody else needs Penn State’s binding notes. There’s data models such as the MARC format for holdings. The SHARE-VDE project has tried assigning provenance to different statements, which might be a way forward for linked data. In this model, you could decide that you’re fine with things added by Penn State catalogers, but reject everything from University of Pennsylvania because your cataloging departments have different philosophies. But it might also get messy and overengineered. What if the Penn State catalogers did great subject analysis but the Penn catalogers added the page length, publication info, etc.
I hope that, ultimately, working in shared systems can prove to be more rewarding than it is frustrating. From my own work in Wikidata and from talking to people who contribute to or upgrade records in NACO or OCLC, I know it can be rewarding to create something which others later expand or to flesh something out so that it’s high-quality and available for reuse. Just adding something like those “archives at” statements, for example, generates a whole new potential web for us or others to traverse.
And on that note, just as not everything has to be in the bib record and records don’t have to be duplicated with minor variations across all our systems, if we have actual linked data systems we may not be doing all of our work within our local catalogs or in library spaces at all. Some of us already do multi-locational description through things like NACO records. When it comes to systems, one of my greater disappointments is that in the post-RDA world, where more data is being added to these records, our systems haven’t already incorporated things like author characteristics for search.
There have also been projects that take us outside the library descriptive world, like the PCC Wikidata Pilot project or its older ISNI project. Our libraries don’t own or need to describe everything, but we can work in tandem with other systems. Penn’s Deep Backfile project, led by John Mark Ockerbloom, focuses on identifying and sharing public domain serials content. They’ve used Wikidata as one of their project tools, adding ISSNs to existing records when they’re not present, improving completeness of these records, or creating new records when one doesn’t exist. Of course, it’s an enormous task and in 2019 they invited others to help out as well. Project participants put data from the project and other library records into Wikidata. They’ve then incorporated Wikidata identifiers into their own tool. There isn’t a snazzy discovery integration, but this makes it much simpler to do large queries of either their data or Wikidata and augment it with info from the other.
The potential in external systems and integrations is part of why I don’t think that greater centralization of description means there is going to be less work or should be fewer people doing it. Not only is there a lot of collective work to be done on creating and augmenting these records, on identifying that the record is the correct one for your institution, there is a lot of work we can do across systems to ensure a more complete representation of the works, of the people whose works we’re describing, of the subjects, etc. For example, if the data doesn’t exist somewhere that Silvia Moreno-Garcia has won a Locus Award, her works won’t show up if we did that hypothetical search for works by authors who’ve won a particular award. In fact, the info about her win for Mexican Gothic wasn’t on her Wikidata profile, so I paused while writing this to add it.
I’m aware that talking about our approaches to the work, about how elements of our current work would be carried on and enhanced if and when we move in to linked-data centric systems, doesn’t really give us concrete things to do now. And I did promise “making use” of long liminality. So I’m going to try to put on my tech services person hat and talk about a few things we can do now to our own data.
These aren’t going to be transformative because if there were a transformative proposal, someone else would’ve shared it by now. But I hope that they’ll be useful and that my outlining them here will help you in making a case for doing this work.
Because the data we have is the data we carry forward.
And I need to pause here for something that isn’t as specific to MARC fields and datapoints as the rest of what I’m going to say but is vital to any linked data future.
One of the longtime promises of linked data is that anyone can say anything about anything.
And the data we have is the data we carry forward; it informs the statements we’ll make.
This doesn’t just mean pre-AACR or minimal records. This includes the choice of not just words but worldviews in the subjects used, in abstracts or notes, the way books are classified, and other practices that librarians and, particularly, catalogers have been talking about for ages.
As with the types of projects I’ll propose next, this work isn’t just about linked data. It matters now, in the spaces where your community encounters your materials, whether that’s in the catalog or discovery system or in ways classification drives shelving. But its impact multiplies the more our data is shared inside and outside our community. If you’re publishing records to a shared database or onto a network of nodes, what assertions are made in your records and are they the ones you want to put out into the world?
This is an area in which I strive to be informed but am not an expert. My recommendation, if you’re not already following these conversations or if you’re uncertain whether you’re up-to-date, is that the Cataloging Lab has an excellent compilation of resources.
I’m now pivoting back to where I have more expertise, which is the use and reusability of MARC data and places where most catalogs can be improved. I’ll identify three types of cleanup projects and why each is important as we think about transitioning from one type of descriptive cataloging system to another. I’m going to give examples of each and I have a couple links at the end for resources with even more ideas. This talk will also be posted on my website, so don’t worry about missing one while I’m talking.
I’ve divided the types of data cleanup work into three different things we need from our data. What we need is data that is:
That’s not really a list, more of a messy Venn diagram, so you’ll see overlaps as I talk specifics.
Whether or not we get systems which make full use of linked data, one of the best things to come out of these tools is the way they incorporate controlled sources into the descriptive process. Even though we should be able to validate certain kinds of fields and see when we’ve transposed the “th” in “Smith,” that doesn’t mean we don’t have things that come in through batch or older data or just times when we didn’t notice that an authorized field didn’t validate. Depending on whether you do authority control and any existing workflows, you might also conduct a large-scale review of unauthorized headings in your system and identify those which seem to have typos in them.
Or, because again it’s such a typing and text-based format, one place I’ve encountered a variety of issues is in the 336, 337, and 338. You’d think that with templates, we’d have the right terms in the right fields, but we’ve identified quite a few cases where the wrong type of term is showing up in the record. Fortunately, searching the subfield 2 to find all the instance of rdacarrier in the 336, etc, let us identify those (or find out when they were correct but the subfield 2 was inaccurate).
You might also hunt for records with problems or inconsistencies in the leader and fixed fields. For example, do you have records which have unusual format types, such as an 007 with videorecording formats like “g” - Laserdisc? While you might have some Laserdiscs, they’re at least worth a review. In our case, a new catalog exposed about 2500 reused records where the cataloger, maybe ours, maybe someone contributing to WorldCat, forgot to change that part. Fortunately the rest of the record has useful fields like the 338, 340, 347, and/or 538, which helped us sort into “actually a VHS,” “actually a DVD,” etc. We could also check call numbers and item types for under-described records.
Language and place of publication codes in your 008 should come from controlled lists. I would suggest taking a tiered approach if you review these. The easiest project is reviewing cases where the codes don’t exist in the controlled lists or where they’re blank. Next, if you’re feeling ambitious, check out the less probable codes. But keep in mind, improbable doesn’t mean inaccurate. We have 8 records for materials which are at least partly in “Yupik languages.”
And then, only if you’ve noticed a problem in your catalog where, say, a default template was wrong for some years and hasn’t been fixed. With language, you could compare with 041 fields, when they exist, and 546 fields. For place of publication, you could do a report of 26xs. But since both of these are present in almost all our records, that’s a level of deep dive I’d only recommend if you knew the data was likely to be wrong and I’d recommend scoping any reports to a period when you knew inaccurate data was being created.
Or another known issue I tackled recently – did you change URLs to local resources you’ve described? We had redirects in place for ages, but relying on those was going to cause problems farther down the line, so I updated them to the appropriate current version.
If you’re feeling a little more ambitious, last year, a librarian and grad student from USC published in Code4Lib Journal about a massive serials holdings cleanup project they did. This is an overlap with the next topic, completeness, because sometimes incomplete records are fundamentally inaccurate.
Due to the scale and history of our catalogs, it’s not reasonable to assume all of our records will be complete – and even when we’re doing cleanup projects, I don’t think any of us have the time available to re-catalog all the older materials in the library.
These are the kind of things you can triage by how important the missing data is. For example you might also have records without an 008 at all. That one came up just last week in conversation with a systems person who’s reviewing post-migration data and I know we have a similar problem, missing 007s on records that should have them. To revisit the language codes I mentioned previously – not only might these be inaccurate, but you may find works with the language code “mul,” for multiple languages but which don’t have an 041 field specifying which languages are used. Those are going to be more important than nice-to-have things like Sound Characteristics (346) or a Table of Contents note that uses subfield t instead of subfield a and hyphens. For many of these, you may just be able to overlay with newer and better records.
There are also times when adding a recent MARC field can be very meaningful. I won’t go down a full ISSN rabbit hole, but the cluster ISSN is a comparatively new concept and a very new MARC field that finally lets us do something we’ve wanted for forever, which is link a bunch of associated serials. Until now, we’ve been able to get some possible matches out of the 022y, but that’s actually for “incorrect ISSN” and contains actual mismatches as well as ISSN for the same serial in a different format or iteration.
Accurate and complete data is not inherently machine-usable data. Most of the time, the more granular our data becomes, the easier it is for machines to parse and use. Even in our current environment, our systems make use of a lot of granular data.
Catalog records serve a variety of purposes. Parts are meant to be read by patrons, overlapping parts need to be indexed for discovery, and other parts are used for machine integrations. Book covers? Those are generated using some combination of OCLC number, ISBN or ISSN, and LCCN. Format icons? Those often come off some combination of leader and fixed fields with fallbacks at various points in the record. The language facet is going to make use of the 008[35-37] and the 041a.
Over time, both our MARC and descriptive standards have evolved and sometimes our older data isn’t quite as granular as it could be. An example that I recently worked on was our ISBNs. Adding a new integration with ILL exposed just how many contained a parenthetical in the 020a, meaning that the field not only contained an ISBN, they also contained qualifying information such as:
(hardcover)
(electronic bk.)
etc.
Now, the subfield q, for qualifying information, was only introduced in 2013. And I think we thought we’d gotten them moved over by one of our cataloging vendors, but after that new integration, which pulled straight from our 020a vs. our catalog or elsewhere, went live, our ILL department found themselves with a bunch of records that glitched out ILLiad’s automated processes. Now, we could’ve redone the integration to just strip out anything after the ISBN. But that would just be pushing the problem down the road. Our Analytics tools revealed that we had 600,000 MARC records with this problem. So instead, I spent about a month doing batch jobs in our Data Control tool and moved these into the appropriate subfield q.
This is just one example. As you work in your own catalog, you may notice inconsistencies in subfield use. You can always check the history of MARC subfields on the Library of Congress’s MARC pages. Many suggestions in previous areas are also granular. When that data can be identified with regular expressions or in other ways that make it easy to move into a newer and more granular format, it improves both the consistency of the data in your current catalog and that data’s future usability.
Whenever possible, this isn’t the kind of cleanup work to do record-by-record. Fortunately, between tools built into your system, like Alma’s normalization rules or Symphony’s Data Control, and MarcEdit, which you can use on data you’ve exported, you can work with things like regular expressions or at least move quickly between records. So if those tools aren’t your strong suit, spending time getting to know them, experimenting (safely), and learning how to use them is a good start to making use of this period of liminality. You may also need to team up with someone who’s got the skills to find the records so that you can fix them.
Besides what I’ve written here and will post on my blog, there are a few other resources I’ve saved over the years for possible data cleanup projects. The first, is Zemkat’s “Looking for Trouble” project site. It’s especially useful for Alma folks, since she’s shared related reports in a folder that can be used by others in the system. But even for the rest of us, it’s a good starting point. There’s also the DC Public Library’s 2020 Cataloging/Metadata Remote Work document, which I saved more in anticipation of an ILS migration data cleanup than pre-linked data, but has some other good suggestions. I also asked John Ockerbloom at Penn about the Deep Backfile and he says a lot of work’s been done, but it’s not yet complete. He encouraged anyone who’s interested in contributing to get in touch.
These are a few examples. When looking at your own records and prioritizing:
If you’ve got specialized knowledge of some area, consider using it to enhance things like Wikidata. I know this group has hosted events which can help you become familiar with the system.
To return to the metaphor of the airport gate, we have been stuck in this in-between place together for quite a while now. But we’re not quite as powerless as we are when we’re passengers. For one, we’re all in this together and can share resources and ideas, like the Cataloging Lab or the Looking for Trouble cleanup project site.
We can anticipate that our existing practices will evolve, perhaps in ways we don’t anticipate, but not beyond recognition. The approaches we need are extensions of ones we already have. And if we want to ensure our data becomes better in the future, there’s plenty of remediation work we can do now.
That’s one we haven’t solved well, but not something that any of our emerging technologies promise to help us do in a way that isn’t convoluted, invasive of privacy, or both. ↩︎
Well, unfortunately mostly vendors, which means that we do not have or take much agency. ↩︎
Writing as Shmargaret Shmitchell. Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell, “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜”, FAccT ‘21: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610-623. https://doi.org/10.1145/3442188.344592. ↩︎
Although if enough people write about it, they’re going to catch up through sheer ingestion and predictive patterning, but only for this one instance. ↩︎