RSSAmplifier

Blog

Friedrich Lindenberg

pudo.orgRSS feed ↗68 posts

Latest posts

A little tour of aleph, a data search tool for reporters

In a short story by Jorge Luis Borges, the Aleph is a point in space that contains all others. To those who see it, it presents the entire universe at once - an investigative reporter’s dream. Over the past six months, I’ve been working for OCCRP to productise a tool named after this mythical object . It’s based on a prototype I hacked up as part of my 2014 Knight Fellowship, and it has now grown…

A Poor Journalists's Text Mining Toolkit

How can you search and analyze collections of documents on your own computers with simple tools? At last weekend’s DataHarvest, Robert and I ran a workshop to answer that question. As people were seemed interested, here’s a write-up of the two key tools we worked with: Apache Tika for content extraction and regular expressions in Sublime Text as an advanced search tool. So, here’s the scenario: a…

Against Decentralization

In the free software/open web community, the notion that the web should be decentralized is more than a shared ideal, it is a piece of dogma. It is key to all the characteristics of the web that we are most proud of: diversity, innovation, the competition of ideas rather than bank accounts. There are plenty of great reasons to strive for decentralization. Zittrain’s The Future of the Internet and…

Keeping stock: investigative data warehouses

In this post I want to describe a design for investigative databases. Unlike the tooling that I’ve been working on for influence mapping projects, this approach is intended to be simple, reliable and extensible. The basic idea is to make sure that all data sources are loaded as tables in a shared, relational database. This includes both large data sources (e.g. company registries) and small…

SpenDB, a data analysis tool for government finance, looking for testers!

Earlier this month, I wrote about the prototype of SpenDB , a light-weight analytical tool for government financial data. After nearly a month of work, the tool has now matured into a beta release, ready for thorough testing with your data! Checking out the largest suppliers of World Bank projects in South Africa. The two most obvious changes are the re-designed user interface with a much simpler…

On Hacks/Hackers, Google and community building

A few weeks ago, the US team of Hacks/Hackers announced their plans to turn the network of journalism innovators into a collaboration with Google News Labs, starting with an event in Berlin. I tweeted about this , and today, Phillip Smith wrote a thoughtful reaction titled What is Hacks/Hackers? Given this invitation to debate, I wanted to outline my criticism in more detail, too. I’ve always…

SpenDB, a light-weight tool for government financial data

One day I will find the right words, and they will be simple. - Jack Kerouac SpenDB is a prototype-stage, light-weight data loading tool and analytical API for government financial data. Over the past few months, I have spent my weekends simplifying and modernizing the OpenSpending codebase to create this tool. So far, I’ve been able to make it easier for non-technical users to submit and model…

Who’s got dirt? - What if robots could do cross-border investigations?

This is a cross-post from IJNet . Cross-border investigative journalism has been much discussed in recent years. Working with the African Network of Centers for Investigative Reporting (ANCIR), I got a chance to observe that process when two reporters from Italy’s investigative reporting project (IRPI) came to visit South Africa to trace the business interests of the Italian Mafia in the country.…

8 things you probably believe about your data standard

Developing open data standards is all the rage. IATI, EITI, OCDS, GTFS, XBRL, SDMX, BDP, HDX - if your sector doesn’t have a cryptic-sounding data initiative yet, it probably will soon. In fact, chances are that you’re drawing one up right now (I am). In that case, here’s a list of things you may believe about your data standard. They are probably not true: Policy and tech people on your team mean…

A Tale of Two Networks

As part of my fellowship, I’ve had the chance to contribute to two influence mapping projects in South Africa and Mozambique. While both projects focus on finding possible conflicts of interest within a small group of politically exposed persons (PEPs), their approach has been very different. Modelling a country In South Africa, the Siyazana project had the goal of collecting and integrating data…

Data doesn't grow in tables: harvesting journalistic insight from documents

When we discuss data journalism, we often tend to think of nicely formatted spreadsheets full of financial data or crime stats. Yet most journalistic source material does not take the form of tables, but it comes in messy collections of documents, whether on paper, or scraped off a web site. Over the last few months, I’ve worked on a few such projects, for example with OpenOil on mining SEC…

Why influence mapping matters to journalism

Building Grano started with a desire to map political and economic influences. Developing it further has brought with it a number of questions and re-examination of our motivations behind the tool; actually, why would journalists want software to help map out the connections between people in politics and industry? The question may sound like a no-brainer - understanding connectivities is key to…

What if journalists had story writing tools as powerful as those used by coders?

The last weekend saw the first Al Jazeera Canvas hackday in Doha. I had the opportunity to work with an amazing team of journalists , designers and engineers to tackle the challenge set by the organizers: re-think the way in which context is used in the production and dissemination of news stories. The question we were asking ourselves was this: how can we help reporters to have more relevant…

Oil Rush on Edgar Creek

Over the past month, I’ve worked with OpenOil in an effort to find oil contracts which have been been published as part of filings to the US stock exchange regulator, the SEC. While these contracts are not usually public, companies are required to file full contract documents as part of their annual reports under certain conditions. The OpenOil team, Johnny West, Anton Rühling and Don Hubert, had…

Grano advanced queries, and Linked Data

At its current stage of development, grano has achieved some level of maturity as an influence mapping toolkit. We’ve got a great workflow for importing raw data , and we’re running a number of different sites off the backend. A new web interface, to make the application more accessible to non-technical users, is in the works. So it’s time to look at the next big challenges: building out the way…

Civic Patterns, a language for news and citizen engagement design

This is a cross-post from IJNet , ICFJ’s network of resources for journalists. While many in journalism are searching for ways to harness their readers’ expertise and to use data to tell compelling stories, technologists and NGOs who build civic technologies around the world are asking some of the same questions. Organizations like the UK’s MySociety , US-based Code for America , Code for South…

Opening up Europe's procurement data

What is the next European dataset that investigative journalists need to look at? This question was raised at the DataHarvest conference back in 2012 . The event brings together technologists and reporters from across the EU’s member states. Brigitte , investigative superstar of FarmSubsidy fame and co-host of the conference, had a clear answer: let’s open up TED (Tenders Electronic Daily) , the…

OffenesParlament: Open Data-Projekt sucht neues Team

Vor drei Jahren habe ich mit der Entwicklung des OffenenParlaments begonnen. Ziel ist, einen einfachen Überblick über die Aktivitäten der Gesetzgeber im Bund zu erlauben: welche Gesetze stehen grade zur Debatte? Welche Debatten finden im Plenum statt? Welche Politiker sind an welchen Prozessen beteiligt? Was passiert in Bundestag und Bundesrat? Die Daten dazu wurden von den verschiedenen Webseiten…

Why I'm building Grano, and why you should help

Over the past weeks, I’ve spent a lot of time working on grano , a social network analysis toolkit for news applications. My goal is to make a reusable, adaptable and open source tool that helps journalists and researchers to map relations between people, companies and institutions. OpenInterests.eu is a data application based on grano. While there are many awesome projects working along similar…

OpenInterests.eu: relating lobbying, expert groups and public finance in the European Union

During last weekend’s #EPhack in Brussels, I built a minimalistic frontend for OpenInterests.eu . The site lets everyone explore which people, companies and institutions have political or financial interests in decisions of the European Union institutions. OpenInterests.eu interface prototype from the #hackEP coding session. What is it good for? While it’s still an early prototype, my hope is to…

Charting Social Network Analysis Tools

This post is a cross-post from the Knightlab’s new Untangled site, a set of resources around social network analysis for journalists. Working on different open data projects throughout the last years, I kept coming back to the idea that there should be an intermediate layer for analyzing and contextualizing data. Such a system would help researchers, journalists and other non-technical users to…

Where to learn about data journalism

Whenever we do data journalism training, I mention many different resources where people can learn more about the techniques, tools and community that we discuss during the workshops. The links to these usually end up in the notes section my slides, which isn’t very helpful. So, instead, here’s a list of interesting resources for getting started with data journalism. What is data journalism,…

datawi.re: when the data mountain comes to you

For a few months now, Annabel and I have been working on datawi.re , an effort to create a better way for journalists to keep track of data feeds. If This Then News, a prototype we hacked on during the OpenNews introduction in January. The tool would create a simple way for users to list people, companies and institutions that they want to track across different sources of data. Such sources could…

5 Project Ideas for News Technology

One aspect of the fellowship is to think about the types of tools that news organisations may find useful in their work. For the past few moneths, I’ve been keeping a list of ideas that I feel there might be a need for. None of these are really new concepts, but they are places in which I feel that none of the existing technologies are mature enough to work in a newsroom. Easy choropleth maps.…

Why German elections don't make sense

There’s only a week left on the German federal campaign, so I’ve decided to take a closer look at the mechanics of the actual election by implementing the complete tabulation process in a JavaScript library, btw13.js (see it in action ). Tally of the 2013 German elections, based on btw13.js's calculations of seat allocations. Our electoral system is notoriously complex , with 299 districts…

Twindle: lessons learned from data mining Twitter

Twitter is the new voxpop: a quick way for news organisations to show what the people think, without the actual hassle of talking to any. Over the last few months, I’ve spent some time recording a subset of German Twitter traffic as a way to track public engagement on political issues during the current campaign season. While the resulting application didn’t end up getting launched in time for the…

Update on ReGENESIS: API and Slides

It’s been two weeks since I’ve first blogged about ReGENESIS , and the project has gathered some very positive feedback. I’ve been able to present the site to a number of people and to extend the functions of the site. Here’s a few quick updates: JSON API for statistics One of the key value propositions of ReGENESIS had been the option to get live statistical data and to use it to quickly render…

ReGENESIS: German statistics as raw data

A choropleth map to indicate the availability of high-quality, machine readable statistical data in ReGENESIS. One of the first tasks I was given by Spiegel Online was to make a set of simple maps to display basic statistics about Germany - things like population, unemployment or insolvencies. As Germany’s statistical data are collected in a system called GENESIS , I though that this would be…

Data Management in News Organisations

Over the last few months, I’ve had the pleasure of joining a few discussions regarding the long-term management of data inside Spiegel. The company is already running DIGAS , a massive, hand-curated archive of news content from the past decades, including articles published not just by Spiegel but also by many other German and international news outlets. Can the same thing be created for data? A…

Why civic hackers should hang out in a newsroom

In Spiegel's snack bar, the sixties are going strong. Coming out of a successful hackday, you usually need two things: a good night’s sleep and ten months of time to grow the things you’ve worked on into a real product. While sleep is easy to come by, finding the time and the environment to work on crazy ideas is much harder. This is why, after visiting open government-related hackdays for three…

Notes: Crunching text documents for fun and knowledge

One exciting development at Spiegel is the recent introduction of a weekly data journalism workshop that brings together reporters, fact checkers and designers from both the print and online sections of the organisation. This week’s workshop will focus on dealing with large collections of documents, so I took some time on Monday to experiment with a few different text mining components. My goal…

Privilege in Nepal

As a twentysomething, white male in the tech industry, privilege really shouldn’t come as a surprise to me. Yet during my recent time in Nepal as a trainer for the Data Bootcamp , I’ve been the subject of positive discrimination to a degree which I hadn’t experienced before. Before reaching Nepal I started to notice a change in treatment, as I was moved to the emergency exit row over a made-up…

Exploring Europe through data

Last weekend, Stijn and I visited inkLink13 , where I presented a few ideas about using data to cover the European Union. The meet-up was a first encounter between Budapest’s tech and journalism communities. As both groups shared some of their ideas for the future of news in Hungary, the discussion was given a special weight by the actions of the Orban government, which has shown little regard for…

sqlaload, an ETL wrapper for SQLAlchemy

sqlaload is a small library that I use to handle databases in Python data processing. In many projects, your process starts with very messy data (something you’ve scraped or loaded from a hand-prepared Excel sheet). In subsequent stages, you gradually add cleaned values in new columns or new tables. Managing a full SQL schema for such operations can be a hassle, you really want something close to…

DataFreeze - scripted static data exports

Every hour, Spiegel Online serves more than half a million visitors. To make that work, all content has to be served via a CDN. For data-driven applications that means: no dynamic queries can be served easily, data needs to be static. This doesn’t need to be a showstopper for great content, sites like the UNDP data explorer demonstrate that often, a set of JSON file is enough to power a great…

The statistic that doesn't exist.

German national statistics are a thorough business. If, for example, you wanted to know the amount of horsemeat produced in the country, the answer is right there, in table 41331-0002: 967 horses were slaughtered in December 2012, yielding 256 tons of raw meat. Yet there are some numbers that the government claims cannot be collected. One such indicator is the number of people who live on the…

OpenNews, one week in: a first taste of the world of online news

That was quick. My first week as an OpenNews fellow at Spiegel Online is already coming to an end today. I’ve spent these first days exploring different aspects of the organization, inevitably failing to give a succinct explanation of what kind of a creature I am and what sort of trouble I plan. What I’ve been trying to convey is that I am a technologist who is excited about news. When I applied…

Journoid, data notifications

At the Open Interests hackday in November, a discussion with Martin Stabe from the FT’s interactive desk led me to code up a prototype of Journoid . The idea is to monitor changing on-line datasets for remarkable information, like earth quakes , procurement in a particular industry or a close parliamentary vote. While we’d discussed alerting in the context of OpenSpending before, Martin had a…

Wrangling dirty data with messytables.

One of the largest data collection projects we have done so far has been the consolidation of the UK’s departmental expenditure . Over 370 different government entities have published a total of more than 7000 spreadsheets. Many of those have obviously been hand-crafted or at least manually processed. Our goal was to consolidate the contained information into a single spreadsheet, discarding all…

Data Catalogues are People!

Last week, Matej Kurian published a message on the okfn-labs mailing list, describing the various sources he had discovered for machine-readable excerpts of the EU’s joint procurement system, TED. What struck me about this message was that, apparently, this polite and brilliant policy wonk had turned into something strange: into a data catalogue. While not quite a Kafka-grade transformation, it’s…

dataissues.org - public issue tracking for data defects

On June 21st, the Knight News Challenge Round on Data ended. The day before, Rufus , Ross and I sat down to write out some ideas that we’d been discussing for a while. While we submitted proposals for Grano and DataProtocols , we decided to hold back on this idea for another round. Still, sharing is caring. 1. What do you propose to do? [20 words] We’ll create a web service where data wranglers…

Social network analysis for advocates and journalists.

On June 21st, the Knight News Challenge Round on Data ended. The day before, Rufus , Ross and I sat down to write out some ideas that we’d been discussing for a while. The first idea I want to repost here is a proposal for Grano, which I’ve discussed in this blog before . 1. What do you propose to do? [20 words] We’ll make a powerful tool for journalists and advocates to keep track of actors and…

Six months of open data - Shuttleworth Recap

In November 2011, the Shuttleworth Foundation awarded me a so-called flash grant. Flash grants are an awesome idea: instead of a full fellowship, a smaller sum of money is given to people, with only the condition of reporting about their activities to the world. Since I haven’t been blogging very much about many of the projects in which I’ve been involved over the last few months, here’s a small…

OffenesParlament - was passiert im Bundestag?

Schon seit langem wünsche ich mir eine Seite wie TheyWorkForYou oder OpenCongress für Deutschland. Ein Portal, auf dem man einfachen Zugang zu den Plenarreden und Arbeitsabläufen des Bundestags bekommt. Im Gegensatz zu Abgeordnetenwatch soll der Fokus also nicht auf den Parlamentariern, sondern auf der Sache liegen. Stefan hatte dazu vor einigen Jahren mit dem Bundestagger einen ersten Aufschlag…

Tranzparenz in Deutschland - die Chronik des geschlossenen Haushalts

Seit etwa zwei Jahren versuchen wir nun aus offizieller Quelle Zugriff auf die analysierbaren Zahlen des Bundeshaushalts zu erhalten. Während Staaten wie England und die USA längst einen direkten Zugriff auf die eigenen Kassen (also die einzelnen Empfänger öffentlicher Gelder) ermöglichen, scheitert im Bund bereits der einfache Zugriff auf die Budgetdaten. Hier also die beinahe komische Chronik…

Notizen zu OpenGov

Gestern hat Stefan fuer die OKF-DE an einer Anhoerung des Bundestags zum Thema Partizipation teilgenommen. Im Vorlauf dazu haben wir ein kleines Papier zusammengestellt, um die Fragen der Fraktionen zu beantworten. Um Schwung in die Sache zu bringen hatte ich ein paar Thesen gemacht, die nicht im Papier auftauchen: Der Staat muss sich als Plattform fuer das vernetzte Handeln und Entscheiden seiner…

Exploring the EU Transparency Register

In all, there are four different roles that an actor - such as an organisation, person or company - can have in the transparency register: They can be registered directly, in one of the six categories. They can be listed as a member organization of one of the registered groups. While some groups actually represent a variety of different organisations, many (especially from Germany, Austria and…

Open Data im e-Gov-Gesetz

Fragt man in Open Data-Diskussionen Verwaltungsleute, wie das Thema in Deutschland einen echten Schub bekommen koennte, dann gibt es eine haeufige Antwort: ein Gesetz! Obwohl ich mir natuerlich ein eigenes Open Data-Gesetz wuensche ist damit in naehrer Zukunft wohl kaum zu rechnen. Dafuer gibt es zwei anstehende Gesetze, in denen Open Data zum Thema werden koennte: die Novelle des…

A brief journey into the UK's accounting data

The Whole of Government Accounts: there’s a bold claim in the name of the UK’s central reporting system. The database, released as part of a larger system called COINS, contains key financial documents – including balance sheets and cash flow statements – for all departments, public bodies and local authorities in Britain. As research for an upcoming OpenSpending journalism project with Lisa Evans…

Hakuna your data - data bootcamp in Nairobi

My Name is XXXX, I am a member of the Kenyan parliament for the constituency of XXXX in the 2007-2012 election cycle. During my time in parliament, I have positioned myself against taxes for MPs. Of the Development Funds allocated to my constituency, I have spent 12mn KSH in 2010 and 8mn KSH in 2009. Since 2007, I’ve funded 201 projects, of which 72 (9mn KSH) related to Education, 56 (7.2mn KSH)…