Tag Archives: Guardian

How do you navigate a liveblog? The Guardian’s Second Screen solution

I’ve been using The Guardian’s clever Second Screen webpage-slash-app during much of the Olympics. It is, frankly, a little too clever for its own good, requiring a certain learning curve to understand its full functionality.

But one particular element has really caught my eye: the Twitter activity histogram.

In the diagram below – presented to users before they use Second Screen – this histogram is highlighted in the upper left corner.

Guardian's Second Screen Olympics interactive

What the histogram provides is an instant visual cue to help in hunting down key events.

Continue reading

A case study in online journalism: investigating the Olympic torch relay

Infographic: Where did the Olympic torch relay places go? What we know so far

image by @CarolineBeavon

For the last two months I’ve been involved in an investigation which has used almost every technique in the online journalism toolbox. From its beginnings in data journalism, through collaboration, community management and SEO to ‘passive-aggressive’ newsgathering,  verification and ebook publishing, it’s been a fascinating case study in such a range of ways I’m going to struggle to get them all down.

But I’m going to try.

Data journalism: scraping the Olympic torch relay

The investigation began with the scraping of the official torchbearer website. It’s important to emphasise that this piece of data journalism didn’t take place in isolation – in fact, it was while working with Help Me Investigate the Olympics‘s Jennifer Jones (coordinator for#media2012, the first citizen media network for the Olympic Games) and others that I stumbled across the torchbearer data. So networks and community are important here (more later).

Indeed, it turned out that the site couldn’t be scraped through a ‘normal’ scraper, and it was the community of the Scraperwiki site – specifically Zarino Zappia – who helped solve the problem and get a scraper working. Without both of those sets of relationships – with the citizen media network and with the developer community on Scraperwiki – this might never have got off the ground.

But it was also important to see the potential newsworthiness in that particular part of the site. Human stories were at the heart of the torch relay – not numbers. Local pride and curiosity was here – a key ingredient of any local newspaper. There were the promises made by its organisers – had they been kept?

The hunch proved correct – this dataset would just keep on giving stories.

The scraper grabbed details on around 6,000 torchbearers. I was curious why more weren’t listed – yes, there were supposed to be around 800 invitations to high profile torchbearers including celebrities, who might reasonably be expected to be omitted at least until they carried the torch – but that still left over 1,000.

I’ve written a bit more about the scraping and data analysis process for The Guardian and the Telegraph data blog. In a nutshell, here are some of the processes used:

  • Overview (pivot table): where do most come from? What’s the age distribution?
  • Focus on details in the overview: what’s the most surprising hometown in the top 5 or 10? Who’s oldest and youngest? What about the biggest source outside the UK?
  • Start asking questions of the data based on what we know it should look like – and hunches
  • Don’t get distracted – pick a focus and build around it.

This last point is notable. As I looked for mentions of Olympic sponsors in nomination stories, I started to build up subsets of the data: a dozen people who mentioned BP, two who mentioned ArcelorMittal (the CEO and his son), and so on. Each was interesting in its own way – but where should you invest your efforts?

One story had already caught my eye: it was written in the first person and talked about having been “engaged in the business of sport”. It was hardly inspirational. As it mentioned adidas, I focused on the adidas subset, and found that the same story was used by a further six people – a third of all of those who mentioned the company.

Clearly, all seven people hadn’t written the same story individually, so something was odd here. And that made this more than a ‘rotten apple’ story, but something potentially systemic.

Signals

While the data was interesting in itself, it was important to treat it as a set of signals to potentially more interesting exploration. Seven torchbearers having the same story was one of those signals. Mentions of corporate sponsors was another.

But there were many others too.

That initial scouring of the data had identified a number of people carrying the torch who held executive positions at sponsors and their commercial partners. The GuardianThe Independent and The Daily Mail were among the first to report on the story.

I wondered if the details of any of those corporate torchbearers might have been taken off off the site afterwards. And indeed they had: seven disappeared entirely (many still had a profile if you typed in the URL directly – but could not be found through search or browsing), and a further two had had their stories removed.

Now, every time I scraped details from the site I looked for those who had disappeared since the last scrape, and those that had been added late.

One, for example – who shared a name with a very senior figure at one of the sponsors – appeared just once before disappearing four days later. I wouldn’t have spotted them if they – or someone else – hadn’t been so keen on removing their name.

Another time, I noticed that a new torchbearer had been added to the list with the same story as the 7 adidas torchbearers. He turned out to be the Group Chief Executive of the country’s largest catalogue retailer, providing “continuing evidence that adidas ignored LOCOG guidance not to nominate executives.”

Meanwhile, the number of torchbearers running without any nomination story went from just 2.7% in the first scrape of 6,056 torchbearers, to 7.2% of 6,891 torchbearers in the last week, and 8.1% of all torchbearers – including those who had appeared and then disappeared – who had appeared between the two dates.

Many were celebrities or sportspeople where perhaps someone had taken the decision that they ‘needed no introduction’. But many also turned out to be corporate torchbearers.

By early July the numbers of these ‘mystery torchbearers’ had reached 500 and, having only identified a fifth, we published them through The Guardian datablog.

There were other signals, too, where knowing the way the torch relay operated helped.

For example, logistics meant that overseas torchbearers often carried the torch in the same location. This led to a cluster of Chinese torchbearers in StanstedHungarians in Dorset,Germans in BrightonAmericans in Oxford and Russians in North Wales.

As many corporate torchbearers were also based overseas, this helped narrow the search, with Germany’s corporate torchbearers in particular leading to an article in Der Tagesspiegel.

I also had the idea to total up how many torchbearers appeared each day, to identify days when details on unusually high numbers of torchbearers were missing – thanks to Adrian Short – but it became apparent that variation due to other factors such as weekends and the Jubilee made this worthless.

However, the percentage per day missing stories did help (visualised below by Caroline Beavon), as this also helped identify days when large numbers of overseas torchbearers were carrying the torch. I cross-referenced this with the ‘mystery torchbearer’ spreadsheet to see how many had already been checked, and which days still needed attention.

But the data was just the beginning. In the second part of this case study, I talk about the verification process, SEO and collaboration.

A case study in online journalism: investigating the Olympic torch relay

Infographic: Where did the Olympic torch relay places go? What we know so far

For the last two months I’ve been involved in an investigation which has used almost every technique in the online journalism toolbox. From its beginnings in data journalism, through collaboration, community management and SEO to ‘passive-aggressive’ newsgathering,  verification and ebook publishing, it’s been a fascinating case study in such a range of ways I’m going to struggle to get them all down.

But I’m going to try. Continue reading

Two guest posts on using data journalism techniques to investigate the Olympics

Corporate Olympic torchbearers exchange a 'torch kiss'

Investigating corporate Olympic torchbearers - analysing the data and working collaboratively led to this photo of a 'torch kiss' between two retail bosses

If I’ve been a little quiet on the blog recently, it’s because I’ve been spending a lot of time involved in an investigation into the Olympic torch relay over on Help Me Investigate the Olympics.

I’ve written two guest posts – for The Guardian’s Data Blog and The Telegraph’s new Olympics infographics and data blog – talking about some of the processes involved in that investigation. Here are the key points: Continue reading

Online journalism jobs – from the changing subeditor to the growth of data roles

The Guardian’s Open Door column today describes the changes to the subeditor’s role in a multiplatform age in some detail:

“A subeditor preparing an article for our website will, among other things, be expected to write headlines that are optimised for search engines so the article can be easily seen online, add keywords to make sure it appears in the right places on the website, create packages to direct readers to related articles, embed links, attach pictures, add videos and think about how the article will look when it is accessed on mobile phones and other digital platforms. Continue reading

The future of open journalism: how journalists need to step up their game

Wolf blowing down the pig's house

Illustration by Leonard Leslie Brooke, from Wikimedia Commons

Cross-posted from XCity Magazine

The future of journalism, according to The Guardian’s ‘3 Little Pigs’ film, is “open journalism”. Users are becoming part of every element of news production. The newsroom no longer has walls.

If that is going to happen then journalists need to huff, and puff, and blow down three particular houses of our own: our preconceptions around the sources that we use online; around why people contribute to the news process; and about how we protect our sources. Continue reading

Video: how a local website helped uncover police surveillance of muslim neighbourhoods

Cross-posted from Help Me Investigate

The Stirrer was an independent news website in Birmingham that investigated a number of local issues in collaboration with local people. One investigation in particular – into the employment of CCTV cameras in largely muslim areas of the city without consultation – was picked up by The Guardian’s Paul Lewis, who discovered its roots in anti-terrorism funds.

The coverage led to an investigation into claims of police misleading councillors, and the eventual halting of the scheme.

As part of a series of interviews for Help Me Investigate, founder Adrian Goldberg – who now presents ‘5 live Investigates‘ and a daily show on BBC Radio WM – talks about his experiences of running the site and how the story evolved from a user’s tip-off.

Comparing apples and oranges in data journalism: a case study

A must-read for any data journalist, aspiring or otherwise, is Simon Rogers’ post on The Guardian Datablog where he compares public and private sector pay.

This is a classic apples-and-oranges situation where politicians and government bodies are comparing two things that, really, are very different. Is a private school teacher really comparable to someone teaching in an unpopular school? What is the private sector equivalent of a director of public health or a social worker?

But if these issues are being discussed, journalists must try to shed some light, and Simon Rogers does a great job in unpicking the comparisons. From pay and hours worked, to qualifications and age (big differences in both), and gender and pay inequality (more women in the public sector, more lower- and higher-paid workers in the private sector), Rogers crunches all the numbers: Continue reading

Guardian to act as platform for arts organisations

The Guardian has been talking about being ‘of the web’ rather than ‘on the web’ for some years now, with a “federated” (as some staff call it) approach to publishing which often involves either selling advertising across, or pulling in content from, other sites (disclosure: this is one of them). Its Open Platform is a technical expression of the same idea, allowing others to build things with its content – which can then take advertising with it. And its successful Facebook app shows its ability to adopt any platform that works.

Now it has announced a partnership with arts organisations – and YouTube – that demonstrates a further development of this approach.  Continue reading

The straw man of data journalism’s “scientific” claim

Guardian cover March 10 2012: Half UK's young black men out of work

Over the weekend Fleet Street Blues has had a bee in its bonnet about the “pretence” of data journalism and Saturday’s Guardian front page: “Half UK’s young black men out of work“.

This, says FSB, is a lie that demonstrates the “pretence” that “‘crunching the numbers’ is somehow an an abstract, scientific, mathematical task”. Continue reading