Jonathan Reeve: Computational Literary Analysis · Dec 5, 2017
Computationally Identifying Similar Books in Project Gutenberg
0Sign in to vote or save
This page cannot be shown here. You can still read it on the original site — the toolbar below keeps your place in the directory.
As one of the first digital libraries, Project Gutenberg has lived through a few generations of computers, digitization techniques, and textual infrastructures. It’s not surprising, then, that the corpus is fairly messy. Early transcriptions of some electronic texts, hand-keyed using only uppercase letters, were succeeded by better transcriptions, but without replacing the early versions. As such,…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.