Category Archives: data journalism

Data journalism pt1: Finding data (draft – comments invited)

The following is a draft from a book about online journalism that I’ve been working on. I’d really appreciate any additions or comments you can make – particularly around sources of data and legal considerations

The first stage in data journalism is sourcing the data itself. Often you will be seeking out data based on a particular question or hypothesis (for a good guide to forming a journalistic hypothesis see Mark Hunter’s free ebook Story-Based Inquiry (2010)). On other occasions, it may be that the release or discovery of data itself kicks off your investigation.

There are a range of sources available to the data journalist, both online and offline, public and hidden. Typical sources include:

Continue reading

Telegraph launches powerful election database

The Telegraph have finally launched – in beta – the election database I’ve been waiting for since the expenses scandal broke. And it’s rather lovely.

Starting with the obvious part (skip to the next section for the really interesting bit): the database allows you to search by postcode, candidate or constituency, or to navigate by zooming, moving and clicking on a political map of the UK.

Searches take you to a page on an individual candidate or a constituency. For the former you get a biography, details on their profession and education (for instance, private or state, oxbridge, redbrick or neither), as well as email, website and Twitter page. Not only is there a link to their place in the Telegraph’s ‘Expenses Files’ – but also a link to their allowances page on Parliament.uk. Continue reading

Interview: Nicolas Kayser-Bril, head of datajournalism at Owni.fr

Past OJB contributor Nicolas Kayser-Bril is now in charge of datajournalism at Owni.fr, a recently launched news site that defines itself as an “open think-tank”.

“Acting as curators, selecting and presenting content taken deep in the immense and self-expanding vaults of the internet,” explains Nicolas, “the Owni team links to the best and does the rest.”

I asked Nicolas 2 simple questions on his work at Owni. Here are his responses:

What are you trying to do?

What we do is datajournalism. We want to use the whole power of online and computer technologies to bring journalism to a new height, to a whole new playing field. The definition remains vague because so little has been made until now, but we don’t want to limit ourselves to slideshows, online TV or even database journalism.

Take the video game industry, for instance. In the late 1970’s, a personal computer could be used to play Pong clones or text-based games. Since then, a number of genres have flourished, taking action games to 3D, building an ever-more intelligent AI for strategy games, etc. In the age of the social web, games were quick to use Facebook and even Twitter.

Take the news industry. In the late 1970’s, you could read news articles on your terminal. In the early 2010’s you can, well… read articles online! How innovative is that? (I’m not overlooking the innovations you’ll be quick to think of, but the fact remains that most online news content are articles.)

We want to enhance information with the power of computers and the web. Through software, databases, visualizations, social apps, games, whatever, we want to experiment with news in ways traditional and online media haven’t done yet.

What have you achieved?

We started to get serious about this in February, when I joined the mother company (22mars) full-time. In just a month, we have completed 2 projects

The first one, dubbed Photoshop Busters (see it here), gives users digital forensics tools to assess the authenticity of an image. It was made as a widget for one of our partners, LesInrocks.com.

More importantly, we made a Facebook app, Where do I vote? There, users can find their polling station and their friends’ for the upcoming regional election in France.

It might sound underwhelming, but it required finding and locating the addresses of more than 35,000 polling stations.

On top of convincing a reluctant administration to hand over their files, we set up a large crowdsourcing effort to convert the documents from badly scanned PDFs to computer-readable data. More than 7,000 addresses have been treated that way.

Dozens of other ideas are in the works. Within Owni.fr, we want to keep the ratio of developers/non-developers to 1, so as to be able to go from idea to product very quickly. I code most of my ideas myself, relying on the team for help, ideas and design.

In the coming months, we’ll expand our datajournalism activities to include another designer, a journalist and a statistician. Expect more cool stuff from Owni.fr.

How to make interactive geographical timelines using Google Calendar and Yahoo Pipes

I was recently given a task where my job was to create a calendar holding around 50 events. Each event also needed to be mapped, and have a corresponding blog post.

Mapping calendar entries made me think, if this could be used for other stuff than simply putting events on a map, – which is quite useful in it’s own way. I thought it would be cool if you could create an interactive map-timeline, controlled dynamically by a (shared)calendar.

Yahoo Pipes by default uses Yahoo Maps, which is great when it comes to narratives. As you can see from the map below (If you don’t see it, click here), each entry has a little arrow that let’s you navigate from marker to marker in a specific order. Each marker also has a number indicating it’s place in a sequence. This is nothing more than entries in a Google Calender with time/date stamps, geo info and a description, mapped automatically using Yahoo Pipes.

{“pipe_id”:”ed13a198a2a83050dd4ace10d12eae16″,”_btype”:”map”,”pipe_params”:{“Curl”:”http://kaspersorensen.com/wp-content/uploads/files/icalyahoopipes.ics”}}

Here’s how you do it.

1. Create a Google Calendar

Simply go to your Google Calendar and create or import a new calendar. You can do this from the settings page under calendars.

2. Make it public

You need to make the calendar public, otherwise Yahoo Pipes won’t have access to it. You can do this while you create it, or afterwards by ticking the box ‘Make this calendar public’ from the sharing settings on your specific calendar. To access the settings for a specific calendar, you click the little arrow in the box on the left hand side that contains your calendars (My Calendars).

3. Create events

Now you simply start adding events to your calendar. Specify what happened, where it happened, when and add the description. You don’t have to add the entries chronologically, they will be sorted by date/time automatically.

4. Feed the iCal file to the Pipe

Go to your calendar settings page, not the general Calendar settings, but the settings for your specific calendar. You will see a section called ‘Calendar address’ with three buttons. Click the green ICAL button and copy the link that pops up. Now go to Mapping Google Calendar Events Pipe and paste it into the ‘Calendar iCal URL’ field and hit ‘Run Pipe’. – Your events are now mapped.

5. Embed on your website

To embed the timeline/map on your website, simply select ‘Get as badge’ just above the map. This will allow you to insert it on your blog or website.

I’m sure there are ways to make this more stable. So if you know how to optimize the pipe, please feel free to do so and let me know.

As Google Maps is already a part of Google Calendar, you would think that there was a nifty way to quickly put a whole calendar on a map, but no. And after failing to use what looked like a saviour, I bumped into a post by Tony Hurst on how to display Google Calendar events on a Google Map. Unfortunately it turns out that the XML feed Tony uses, only parses the 25 most recent calendar entries.

Google Calendar releases their event-entries in iCal format which contains all events. And with a little customization of Tony’s pipe, I managed to come up with a way to map all events from a calendar.

I think this could be potentially useful for developing stories, especially if you can collaborate on the calendar. You end up with data that can be used for nearly anything, not just maps. And if locations aren’t relevant for the story, you could simply take your iCal file and make a normal timeline.

Data and the future of journalism panel discussion: Linked Data London

Tonight I had the pleasure of chairing an extremely informative panel discussion on data and the future of journalism at the first London Linked Data Meetup. On the panel were:

What follows is a series of notes from the discussion, which I hope are of some use.

For a primer on Linked Data there is A Skim-Read Introduction to Linked DataLinked Data: The Story So Far PDF) by Tom Heath, Christian Bizer and Berners-Lee; and this TED video by Sir Tim Berners-Lee (who was on the panel before this one).

To set some brief context, I talked about how 2009 was, for me, a key year in data and journalism – largely because it has been a year of crisis in both publishing and government. The seminal point in all of this has been the MPs’ expenses story, which both demonstrated the power of data in journalism, and the need for transparency from government – for example, the government appointment of Sir Tim Berners-Lee, seeking developers to suggest things to do with public data, and the imminent launch of Data.gov.uk around the same issue.

Even before then the New York Times and Guardian both launched APIs at the beginning of the year, MSN Local and the BBC have both been working with Wikipedia and we’ve seen the launch of a number of startups and mashups around data including Timetric, Verifiable, BeVocal, OpenlyLocal, MashTheState, the open source release of Everyblock, and Mapumental.

Q: What are the implications of paywalls for Linked Data?

The general view was that Linked Data – specifically standards like RDF – would allow users and organisations to access information about content even if they couldn’t access the content itself. To give a concrete example, rather than linking to a ‘wall’ that simply requires payment, it would be clearer what the content beyond that wall related to (e.g. key people, organisations, author, etc.)

Leigh Dodds felt that using standards like RDF would allow organisations to more effectively package content in commercially attractive ways, e.g. ‘everything about this organisation’.

Q: What can bloggers do to tap into the potential of Linked Data?

This drew some blank responses, but Leigh Dodds was most forthright, arguing that the onus lay with developers to do things that would make it easier for bloggers to, for example, visualise data. He also pointed out that currently if someone does something with data it is not possible to track that back to the source and that better tools would allow, effectively, an equivalent of pingback for data included in charts (e.g. the person who created the data would know that it had been used, as could others).

Q: Given that the problem for publishing lies in advertising rather than content, how can Linked Data help solve that?

Dan Brickley suggested that OAuth technologies (where you use a single login identity for multiple sites that contains information about your social connections, rather than creating a new ‘identity’ for each) would allow users to specify more specifically how they experience content, for instance: ‘I only want to see article comments by users who are also my Facebook and Twitter friends.’

The same technology would allow for more personalised, and therefore more lucrative, advertising.

John O’Donovan felt the same could be said about content itself – more accurate data about content would allow for more specific selling of advertising.

Martin Belam quoted James Cridland on radio: “[The different operators] agree on technology but compete on content”. The same was true of advertising but the advertising and news industries needed to be more active in defining common standards.

Leigh Dodds pointed out that semantic data was already being used by companies serving advertising.

Other notes

I asked members of the audience who they felt were the heroes and villains of Linked Data in the news industry. The Guardian and BBC came out well – The Daily Mail were named as repeat offenders who would simply refer to “a study” and not say which, nor link to it.

Martin Belam pointed out that The Guardian is increasingly asking itself ‘How will that look through an API’ when producing content, representing a key shift in editorial thinking. If users of the platform are swallowing up significant bandwidth or driving significant traffic then that would probably warrant talking to them about more formal relationships (either customer-provider or partners).

A number of references were made to the problem of provenance – being able to identify where a statement came from. Dan Brickley specifically spoke of the problem with identifying the source of Twitter retweets.

Dan also felt that the problem of journalists not linking would be solved by technology. In conversation previously, he also talked of “subject-based linking” and the impact of SKOS and linked data style identifiers. He saw a problem in that, while new articles might link to older reports on the same issue, older reports were not updated with links to the new updates. Tagging individual articles was problematic in that you then had the equivalent of an overflowing inbox.

(I’ve invited all 4 participants to correct any errors and add anything I’ve missed)

Finally, here’s a bit of video from the very last question addressed in the discussion (filmed with thanks by @countculture):

Linked Data London 090909 from Paul Bradshaw on Vimeo.

Data and the future of journalism: what questions should I ask?

Tomorrow I’m chairing a discussion panel on the Future of Journalism at the first London Linked Data Meetup. On the panel are:

What questions would you like me to ask them about data and the future of journalism?

The Guardian kicks off the local data landgrab

Tonight I’ve been speaking at a Guardian-sponsored event in Birmingham: a special meetup of the Birmingham Social Media Cafe doubling as a sort-of-build-up-to-a-Hack Day.

And I think it’s a very significant event indeed.

For years I’ve lectured newspaper execs on the value of data and why they needed to get their APIs in order.

Now The Guardian is about to prove just why it is so important, and in the process take first-mover advantage in an area the regionals – and maybe even the BBC – assumed was theirs.

This shouldn’t be a surprise to anyone: The Guardian has long led the way in the UK on database journalism, particularly with its Data Blog and this year’s Open Platform. But this initial move into regional data journalism is a wise one indeed: data becomes more relevant the more personal it is, and local data just tends to be more personal.

Reaching out to those with access to that data, and the ability and knowledge to pick through it, makes perfect sense. But it also means treading on regional toes, and it will be interesting to see how (and indeed if) regional newspapers and broadcasters react.

Cobbling together some sort of regional API would be a welcome start – but is not going to be enough alone: The Guardian have spent years building a reputation in technology circles for their understanding of the web. As The Guardian’s Michael Brunton-Spall pointed out tonight, theirs is the only newspaper to offer ‘full fat’ RSS feeds that allow you to read full articles on an RSS reader – not to mention customisable URLs that allow you to build your own feeds based on combinations of tags, authors and categories. And Open Platform is one of the most, well – open news platforms in the world.

So if other news operations want to compete in this arena, they’ll need to make cultural efforts, not just technical ones.

There are few people in those organisations who truly understand why they should want to compete. They may see it in the context of the mutterings about a move by Guardian Media Group (GMG) into hyperlocal media, but that could be a different kettle of fish entirely (a red herring of sorts if you want to mix metaphors).

These early moves on the data side of things are about more than the prospect of launching competing web publications. It means the Guardian (rather than the GMG) is well positioned to provide a platform for a bottom-up network of hyperlocal sites, to become, in short, a Press Association for the 21st century, catering for a grassroots journalism movement filling ever-increasing holes in the regional news map: not just feeding national and international news to local and specialist websites, but pulling data the other way (although that doesn’t mean there isn’t scope to meet GMG hyperlocal plans in the middle). They have competition here from MSN Local and Reuters’ Open Calais, but I’ve not seen evidence of the same cultural efforts from that direction.

It’s very early days, but things move fast in this sphere. A cry is being taken up that all news organisations need to heed: “Raw data now!“.

Add context to news online with a wiki feature

In journalism school you’re told to find the way that best relates a story to your readers. Make it easy to read and understand. But don’t just give the plain facts, also find the context of the story to help the reader fully understand what has happened and what that means.

What better way to do that than having a Wikipedia-like feature on your newspaper’s web site? Since the web is the greatest causer of serendipity, says Telegraph Communities Editor Shane Richmond, reading a story online will often send a reader elsewhere in search of more context wherever they can find it.

Why can’t that search start and end on your web site?

What happens today

Instead of writing this out, I’ll try to explain this with a situation:

While scanning the news on your newspaper’s web site, one story catches your eye. You click through and begin to read. It’s about a new shop opening downtown.

As you read, you begin to remember things about what once stood where the new shop now is. You’re half-way through the story and decide you need to know what was there, so you turn to your search engine of choice and begin hunting for clues.

By now you’ve closed out the window of the story you were reading and are instead looking for context. You don’t return to the web site because once you find the information you were looking for, you have landed on a different news story on a different news web site.

Here’s what the newspaper has lost as a result of the above scenario: Lower site stickiness, fewer page views, fewer uniques (reader could have forwarded the story onto a friend), and a loss of reader interaction through potential story comments. Monetarily, this all translates into lower ad rates that you can charge. That’s where it hurts the most.

How it could be

Now here’s how it could be if a newspaper web site had a wiki-like feature:

The story about the new shop opening downtown intrigues you because, if memory serves, something else used to be there years ago. On the story there’s a link to another page (additional page views!) that shows all of the information about that site that is available in public records.

You find the approximate year you’re looking for, click on it, and you see that before the new shop appeared downtown, many years ago it was a restaurant you visited as a child.

It was owned by a friend of your father’s and it opened when you were six years old. Since you’re still on the newspaper web site (better site stickiness!), you decide to leave a comment on the story about what was once there and why it was relevant to you (reader interaction!). Then you remember that a friend often went there with you, so you email it to them (more uniques!) to see if they too will remember.

Why it matters to readers

For consumers, news is the pursuit of truth and context. Both the news organization and the journalists it employs are obligated to give that to them. The hardest part of this is disseminating public records and putting it online.

The option of crowd-sourcing it, much like Wikipedia does with its records, could work out well. However just the act of putting public records online in a way that makes theme contextually relevant would be a big step forward. It’s time consuming, however the rewards are great.

Newspapers on Twitter – how the Guardian, FT and Times are winning

National newspapers have a total of 1,068,898 followers across their 120 official Twitter accounts – with the Guardian, Times and FT the only three papers in the top 10. That’s according to a massive count of newspaper’s twitter accounts I’ve done (there’s a table of all 120 at that link).

The Guardian’s the clear winner, as its place on the Twitter Suggested User List means that its @GuardianTech account has 831,935 followers – 78% of the total …

@GuardianNews is 2nd with 25,992 followers, @TimesFashion is 3rd with 24,762 and @FinancialTimes 4th with 19,923.

Screenshot of the data

Screenshot of the data

Other findings

  • Glorified RSS Out of 120 accounts, just 16 do something other than running as a glorified RSS feed. The other 114 do no retweeting, no replying to other tweets etc (you can see which are which on the full table).
  • No following. These newspaper accounts don’t do much following. Leaving GuardianTech out of it, there are 236,963 followers, but they follow just 59,797. They’re mostly pumping RSS feeds straight to Twitter, and  see no reason to engage with the community.
  • Rapid drop-off There are only 6 Twitter accounts with more than 10,000 followers. I suspect many of these accounts are invisible to most people as the newspapers aren’t engaging much – no RTing of other people’s tweets means those other people don’t have an obvious way to realise the newspaper accounts exist.
  • Sun and Mirror are laggards The Sun and Mirror have work to do – they don’t seem to have much talent at this so far and have few accounts with any followers. The Mail only seems to have one account but it is the 20th largest in terms of followers.

The full spreadsheet of data is here (and I’ll keep it up to date with any accounts the papers forgot to mention on their own sites)… It’s based on official Twitter accounts – not individual journalists’. I’ve rounded up some other Twitter statistics if you’re interested.

ABCe: please sort out your terrible website (again)

In March, I appealed to the Audit Bureau of Circulations to sort out its terrible ABCe website. It’s had a redesign. Here’s a list of its latest problems (originally published here).

If at any point the ABC wants to pay me a consultancy fee, for all this free advice, just leave me a comment to tell me how to receive my money …

All the URLs have changed but there are no redirects

New ABCe homepage in Google

New ABCe homepage in Google

They’ve had a redesign, but they haven’t redirected the old URLs to new ones. So, for instance, if you click the second link shown in Google for a search on ABCe, you get page not found.

Lesson When relaunching a website, always 301 redirect your old pages to new ones (even if they’re all just to your new home page). That way, external links still work and you keep the SEO benefit of any links.

They haven’t sorted www vs non www

The more observant will have noticed that the title of the first result in that screenshot says ‘To access IIS Help’. The ABC hasn’t realised that abce.org.uk is not the same URL as http://www.abce.org.uk. And if you go to the ABC URLs without www, you get page not found or server errors.

Compare these pages:

and these ones:

Lesson When you set up your website, redirect yourdomain.co.uk/whatever to http://www.yourdomain.co.uk/whatever. And log in to your google webmaster account to set your preferred domain (www or non-www).

They’re running two absolutely identical websites

ABCs new homepage. No, it's ABCe's. No, it's aaaaggghhh

ABCs new homepage. No, it's ABCe's. No, it's aaaaggghhh

You can access the entire website at www.abc.org.uk – or you can see an identical website at www.abce.org.uk. Continue reading