Showing posts with label data analytics. Show all posts
Showing posts with label data analytics. Show all posts

Thursday, December 16, 2010

Tesco Metrics: Every Little Bit of Data Helps

Liam linked to an article in the Guardian earlier this week, which was all about Nudge. One comment in the article was that "while shopping, working, or even deciding on who to share their lives with, individuals are less thoughtful and less calculating than modern-day economists... typically assume." This blog-post zones in on shopping, in particular the data-analysis of consumer purchasing behaviour at Tesco. The Guardian article linked above also suggests that "any critic who points out that that's hardly news to the women...(and) the men at Tesco... is spot on." Indeed, Tesco have been conducting interesting micro-level analysis on individual behaviour for many years now.

An informative article on this topic was written by Jenny Davies in the Sunday Times last year. According to Davies, Tesco gets its data from its loyalty clubcard scheme; this was launched 15 years ago with much fanfare - the advert below may jog memories for some readers. Davies also informs us that around this time last year, Tesco was tracking "the shopping habits of 16 million families across Britain, delivering an extraordinary insight into their lives — not only for itself but for companies such as Coca-Cola, Nestlé and Unilever, which buy the rights to the data." Readers in the Republic of Ireland might also remember that the Tesco Clubcard was launched there on the 13th. Oct 1997. To date almost 800,000 members have joined in the Republic.



Jenny Davies also tells us that: "Each bill detailing every item in a customer’s shopping basket is logged in a data centre in London Docklands and decoded by Dunnhumby, the marketing firm that is in charge of the scheme. It has to process 100 baskets a second — six million transactions a day. This helps Tesco to decide which products should go on to the shelves at what times, and in early trials it increased sales by as much as 12% in some of the supermarkets." According to the Guardian (in this article), the power of the clubcard was demonstrated in 2009, "when Tesco harnessed the card's database to halt the exodus of shoppers to cheaper retailers because (of) the recession, by doubling the points available to shoppers."

In a blog-post on Tesco data from two years ago, Tony Hirst desribes the early analysis conducted by Dunnhumby, and how this has changed over the last 15 years. A couple of months ago, Dunnhumby (and its recently departed co-founders) were profiled in the Guardian. The article says:
According to company lore, there was a 30-second silence after Humby presented the initial trial's results to the Tesco board, until the then chairman, Lord MacLaurin, declared: "What scares me is that you know more about my customers after three months than I know after 30 years."
One question that readers might have is: what's in it for club-card holders? According to Tony Hirst, a good place to get an answer to this question is the book: Scoring Points: How Tesco Continues to Win Customer Loyalty. Hirst describes the "Clubcard customer contract: more data means better segmentation, means more targeted/personalised services, means better profiling. In short, the more you shop with us, the more benefit you will accrue." According to the Marketing Week magazine, "from the day of its launch in February 1995 the Tesco Clubcard was immediately embraced by customers attracted to the 1% discount off their shopping bills. But its long term success has not been built on discounts alone, rather on the personalisation of the shopping experience."

However, perhaps the last word should go to UCD social psychologist Ken McKenzie, writing on his A Head in Business Blog: "I don’t have a loyalty card, and every time I’m in Boots, Tesco or Dunnes, and they ask if I have one, I feel a slight sense that I should justify why I don’t, as it it’s odd to not have one. And according to rational actor theory in Economics, it is odd to not have a loyalty card and avail of discounts. However, there’s a small but growing body of work in the overlapping area between Psychology and Economics that might explain why (some) people might behave like me."

Thursday, November 11, 2010

Google Refine: A power tool for working with messy data

A story on Mashable tells us that "if you live for data, slave over spreadsheets and constantly find yourself sifting through endless rows and columns of facts and figures, Google’s got a lovely new product just for you — and it’s free and open-source, too." The product in question is Google Refine: a tool for cleaning up data, and more. I watched the introductory video on this new piece of kit; it seems to be developed to a very high standard.

However, I would not recommend Refine to graduate students conducting empirical research. As we have mentioned before on the blog, in academic research it is crucial to keep a record of any changes made to a data-set using the syntax-editors available in most software packages. Nonetheless, Google Refine could be extremely useful for undergraduate projects or for data-analysts who only use Excel. One appealing feature for security-conscious users is that the program must be downloaded, which means that the data is never on the web.

For more sophisticated data-handling, the first thing that comes to mind is Scott Long's recent book on "The Workflow of Data Analysis". This is an excellent starting-point for anyone looking to go about best practice in their research. Also, Daniel Hamermesh has a paper on replication in economics, that is arguably a must-read for graduate students beginning a program in empirical economics. An IZA WP version of the Hamermesh article is available here: Replication in Economics. The most recent discussion of data-issues on this blog (including workflow, publication bias, replication, retractions and empirical controversies) is available here.

Tuesday, September 07, 2010

Assorted Links: 7th September 2010

1. The Hewlett Packard Social Computing Lab focuses on 'methods for harvesting the collective intelligence of groups of people in order to realise greater value from the interaction between users and information'.

2. "The agony and the ecstasy: the history and meaning of the Journal Impact Factor" (Garfield and Sher, 2005).

3. Measurement of administrative burden imposed on Irish business by Central Statistics Office inquiries

4. http://www.bundle.com: Track your spending

5. DIT School of Computing launches a new master's programme in data analytics

6. A story on the OECD's Education at a Glance 2010, in today's Irish Independent: "For every €7 spent on a primary pupil, nine is spent on second level and 12 is spent on third level".

7. Another Irish Independent story: the mechanics of random selection in the CAO application process

8. Bloomberg Businessweek: The World's 50 Most Innovative Companies

9. Below is a video from the OECD showing the importance of tax credits for understanding government funding of research and development (R&D). When tax credits are taken into account, Korea spends the most (as a percentage of GDP) on R&D. Canada spends ten times more on R&D tax credits than it does on direct funding for R&D.

Monday, March 01, 2010

The Data Deluge

Martin has discussed the increasing amount of publically available data several times (e.g. here). This week's Economist features a special report on the "Data Deluge".

Tuesday, February 02, 2010

WePredict

Using Twitter for macro-level analysis has been discussed on the blog before:
(i) Sample Selection, Twitter's Public Timeline, TweetScan and Quotably
(ii) Twilert (re-launched this month), Summize Labs, Twitter's acquisition of Summize
(iii) Life Analytics: Sentiment on the United States Economy
(iv) The Google-Index of Social Media, and the apparent superiority of Bing for deciphering real-time breaking trends

So it was interesting to read a recent article in the Irish Times about two students who used Twitter to predict that Joe McElderry would win the recent X-Factor competition... before the results were announced. "Ben McRedmond (17) and Patrick O’Doherty (16), two fifth years from Gonzaga College, Dublin demonstrated the power of their social networking analysis system, We Predict, at the BT Young Scientist and Technology Exhibition. It was developed over more than five months and is based on storing and studying a growing database of 24.5 million tweets which hold clues about what people are thinking."

A quick search led me to the WePredict website. The information there states that: "WePredict is a showcase of several technologies we have built: a data mining application, capable of mining data from multiple social networks; and a complex suite of analysis tools for analyzing this data...WePredict's large database of over 22 million status updates growing at 20 a second makes it the largest user survey ever done... using the collective intelligence of the whole internet, a mere 1.5 billion people, we can predict the outcomes of elections, talent shows or who will be the christmas #1 and analyze the public reaction to new legislation or medical epidemics."

An exciting endeavour such as this one reminds me of Hal Varian's famous quote that "the sexy job in the next ten years will be statisticians." A two-minute YouTube clip of Google's Chief Economist is shown below, discussing this very issue.

Wednesday, March 04, 2009

The Text Tax

It has been suggested by the Green Party that a 1c tax on text messages would raise some much needed funds for the public purse, in the current economic crisis. This has been referred to as as “unfair and injust” by Tommy McCabe, director of the Irish Cellular Industry Association (ICIA) - see story here.

How much revenue do we think this would generate? According to the Irish Examiner, a record two billion text messages were sent by Irish mobile phone users in the final three months of last year. Say we assume that this is a steady state level of texting, and that a 1c tax would not deter anyone to send a text. With these assumptions made, then the 1c text tax would produce 8 billion cents in revenue per annum, or an annual sum of 80 million euro.

This is all well and good, but I would like to know more about how this tax would be collected. What I have been able to find is a news story from 2006 which suggests that European Union lawmakers have already considered tax on e-mails and text messages as a way to fund the 25-member bloc in the future. Also, a text message tax was introduced in Sacramento, California last December (see story here). The city sent out letters to telecommunications companies to instruct them to levy the tax on customers' bills.

This is an interesting development in the economics of information. While I don't yet have any fears about negative consequences for the widespread distribution of information, comminication taxes could be undesirable if they prevent useful information exchange. Especially in the so-called Information Economy. On a related note, it was announced yesterday that UCD won SFI strategic research cluster funding of €3.56 million - which will be focused on "data analytics".