The third issue of IRRJ contains excellent research papers: CRAWLDoc, a system for contextual ranking and bibliographic metadata extraction from web resources by Fabian Karl and Ansgar Scherp; Simple techniques for efficient top-k batch query processing by Zhixuan Li and Joel Mackenzie; An exploration of online content accessibility from an autism-informed perspective by Hrishita Chakrabarti and…
Unlocking the Value of Data with Vertical Federated Learning by Afsana Khan Organisations generate and store increasing amounts of data through their daily activities. Customer interactions, financial records, medical information, and digital services all produce data that can help organisations understand patterns, improve decisions and support their operations. Machine learning has become an…
by Benard Wanjiru, Patrick van Bommel, Djoerd Hiemstra Automated grading tools can save time and effort while ensuring consistency in evaluating assignments. At our university, we have implemented SOCOLES for evaluating students’ SQL answers using parse tree comparisons between student statements and one or more correct statements. It can grade Data Query Language (SELECT), Data Continue reading…
Web search engines are essential for navigating the web. Suppose we look at the web as a service that is provided by public utility companies, a service similar to electricity, water or telephone. To make sure that everyone has access to the web, public utility companies have to be subject to public control and regulation. Continue reading "Towards a shared infrastructure for assembling web search…
by Gijs Hendriksen, Djoerd Hiemstra, and Arjen P. de Vries We propose to redesign the access to Web-scale indexes. Instead of using custom search engine software and hiding access behind an API or a user interface, we store the inverted file in a standard, open source file format (Parquet) on publicly accessible (and cheap) object Continue reading "Open Web Indexes for Remote Querying"
We organize the third International Workshop on Open Web Search (#WOWS) at ECIR 2026 with two calls for contributions: The first call targets scientific contributions on collaborative search engine building, including crawling, deployment, evaluation, and use of the web as a resource by researchers and innovators. The second call is for the WOWS-Eval shared task, Continue reading "The 3rd…
In the second issue of IRRJ, Paul Kantor, writes an editorial arguing for a more critical adoption of generative AI in information retrieval (IR). He puts his concerns under three distinct headings: consistency, confidence, and completeness. Kantor is the founder of the predecessor Information Retrieval Journal, and together with co-founder Stephen Robertson, his advise helped Continue reading…
by Chris Kamphuis Finding relevant information in a large collection of documents can be challenging, especially when only text is considered when determining relevancy. This research leverages graph data to express information needs that consider more information than just text data. In some cases, instead of using inverted indexes for the data representation in our Continue reading "Chris…
by Benard Wanjiru, Patrick van Bommel and Djoerd Hiemstra Automated grading systems for SQL courses can significantly reduce instructor workload while ensuring consistency and objectivity in assessment. At our university, an automated SQL grading tool has become essential for evaluating assignments. Initially, we focused on grading Data Query Language (SELECT) statements, which constitute the core…
Welcome to Part B, Databases! We will resume Tuesday 4 November with a lecture at 15:30h. in HG00.304 The Databases part contains mandatory, individual quizzes, for which the following honour code applies: You do not share the solutions; The solutions to the quizzes should be your own work; You do not post the quizzes, nor the Continue reading "Welcome to Databases 2025!"