Unicode in DOIs
The previous post looked at how long DOIs are. One of the questions was: 
 
 Do UTF-8 encodings from Unicode characters make any difference on the statistics around DOI length?
Recent content in Home on Pardalotus
The previous post looked at how long DOIs are. One of the questions was: 
 
 Do UTF-8 encodings from Unicode characters make any difference on the statistics around DOI length?
In 2024 DataCite released their first public data file. It’s easy to get a copy . Crossref have made a data dump available for the past few years. 
 Having both files available opens up some interesting possibilities in comparing and combining the two data sources. 
 The most obvious place to look is the DOIs themselves… 
 … and the simplest question you could ask…
Are you interested in the technical side of scholarly publishing technology and
infrastructure? If so, you’re welcome to join the SPPIES Slack group. 
 SPPIES is short for Scholarly Publishing Programmers and Infrastructure
Enthusiasts, in homage to programmers’ love of acronyms. One of the ‘P’s is silent.
DOI s, or Digital Object Identifiers, are everywhere, for a
given value of ’everywhere’. They are the identifiers used to identify and link
research outputs, and a lot more besides. 
 Humans are good at spotting patterns, and with something as ubiquitous as DOIs,
there are plenty of patterns to spot. However, with hundreds of millions of DOIs
and decades of history,…
The word ’technical’ hasn’t always had a great reputation. Think of phrases like ’technically correct’, ’technical details’, or ‘acquitted on a technicality’. It’s easier to think of more negative uses than positive ones. I think there are a couple of reasons for that.
I presented these five principles at the altmetrics18 workshop. You can read the paper submitted to the workshop here . This post is a few years old, but all the ideas still stand up. At the time I was building Crossref Event Data, and discussing what it would take to build an data model that would support community-generated bibliometrics. A lot has changed since, but I think the principles are…
This is a command-line tool for working with DataCite and Crossref data dumps. It can convert between formats, combine snapshots and produce statistics.
A Rust library for parsing, validating and
normalising a selection of scholarly identifiers (DOI, ORCID, ROR, ORCID, ISBN),
etc. It’s work in progress, with more identifiers being added.
Tool to keep a local SQLite database up to date with changes from Crossref.
Hello! I’m Joe Wass. I’m an experienced technical leader and software developer.
I have over a decade’s experience in the open scholarly infrastructure space,
and over 15 years in various types of software development. 
 Find me on 
 
 LinkedIn: https://www.linkedin.com/in/joewass/ 
 Mastodon: https://fosstodon.org/@joewass 
 Email joe@afandian.com…
I can help with: 
 
 planning and refining technical roadmaps 
 building, extending and maintaining software systems, especially concerned with scholarly metadata 
 prototyping new systems 
 
 Get in touch! Email joe@afandian.com