bitsgalore.org · Jul 10, 2018
Crawling offline web content: the NL-menu case
0Sign in to vote or save
This site took too long to answer. You can still read it on the original site — the toolbar below keeps your place in the directory.
In a previous blog post I showed how we resurrected NL-menu , the first Dutch web index. It explains how we recovered the site’s data from an old CD-ROM, and how we subsequently created a local copy of the site by serving the CD-ROM’s contents on the Apache web server . This follow-up post covers the final step: crawling the resurrected site to a WARC file that can be ingested into our web…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.