The comparison of Google Trends data (on the British Election) with polls, bookmakers' odds and prediction markets shows once again that we need to be careful when interpreting search volume data. While the (search-data related) innovations in unemployment forecasting may not be earth-shattering, what else can we learn from trends in seach queries? There needs to be a focus on what the analyst expects when typing in "zombie" or "inflation" or "dole" into a trend-analyser. Hopefully we can all agree that there is no such thing as a zombie. Unemployed people on the other hand, are a very real human problem, and growing in large numbers.
For example, do unemployed individuals looking for information about welfare payments type in "unemployment" or "dole"? Or something else? One experiment (view here) is to type in "unemployment", "dole" and "jobs" into Google Trends, separated by commas. A few observations can be made:
(i) The search volume for "jobs" is relatively stable over the last 6 years
(ii) News reference volume for "jobs" has exploded over the last 2 years, much more so than for "unemployment"
(iii) There is only enough search activity related to "unemployment" for it to register half-way during 2008
(iv) There is only enough search activity related to "dole" for it to register at the start of 2009
(v) There is a fall-off in search volume for "jobs" at the end of every calendar year
While much of this mirrors what we already know about recent economic activity, I had expected "jobs" to have a much higher search volume over the last year. We of course have to be very careful about drawing conclusions, but the stylised facts about search volume suggest that there were more people searching for jobs in 2004 and 2005 than there were in 2008 and 2009. We know that there were more people in need of a job in 2008 and 2009, so what is the explanation? Perhaps job-search is more intense during boom-times. In recessions, maybe people are less likely to search for a job (which they simply believe isn't there). This could of course be incorrect, but now there is an open question.
Other challenging questions about search data are currently at play in the commercial arena; it may be no coincidence that Google Trends was opened up to the public (including academics) in 2006, just as these questions were coming more to the fore. At present, Google, Yahoo and Bing are strongly focused on distinguishing between "interest" and "intent" in search data. There are obvious commercial implications, but solving this problem about interest versus intent would also help academic researchers. When somebody searches for "jobs" do they just want to *see what's out there* (maybe in boom times) or do they *desperately intend* to obtain employment (maybe in recessions)?
Maybe additional keywords would help in solving this interesting puzzle. If you search for "XBOX Price", Google can assume to some extent that you intend to buy an XBOX. Here is an article from last year about Google executives stating that "understanding people, health, communication, education and knowledge" is the next frontier of search. Here is a link to Yahoo!'s "Mindset" research project on 'Intent-driven Search'. Recently, there was an article in the Economist about about Qi Lu: the man behind Bing. According to him, the focus is firmly on "understanding user intent".
It's clear that understanding more about search is the big challenge: for the search-engine based advertising business, and for social scientists. And here is the main reason why search data is (or should be) so interesting for academics: we don't *ask* people for the information they provide in search queries. It's a simple statement, but it has merit. No matter how well-designed surveys are, there will always be things in the ether, trends in society, that will potentially appear in search data first.
Showing posts with label Google Trends. Show all posts
Showing posts with label Google Trends. Show all posts
Thursday, May 06, 2010
Google Trends, Polls, Betting, Prediction Markets and the British Election
Posted by
Anonymous
Those interested in Google Trends may find it worthwhile to bear the first chart (directly below) in mind, as the results of the British Election come through. The red line represents "Gordon Brown", the orange line "Nick Clegg", and the blue line "David Cameron".

A different picture is illustrated in the second chart (directly below). The red line represents "Labour Party", the orange line "Liberal Democrats", and the blue line "Conservative Party".

Of course, this type of information may have no predictive value. The UK Polling Report (an independent survey and polling news website) shows that the majority of polls paint a closer picture, but with a distinct lead for the Conservative Party. Could the higher search (and news) reference volume for "Gordon Brown" be due to an incumbency effect? Could the higher search (and news) reference volume for "Liberal Democrats" be due to a novelty effect? Of course, we can't attempt an answer until later in the week. Many punters have already placed their bets though: bookmakers estimate that up to £40m will have been bet on this election, smashing previous records.
Finally, the Intrade prediction market indicates that the Conservative Party have roughly a 90% chance of winning, as shown in the third chart (directly below).

The Inkling Prediction Market indicates here that there is a 66% chance of a hung parliament, and a 32% chance of a win by the Conservative Party.
Hal Varian, Chief Economist at Google, discusses prediction markets here in the NYT, from a few years ago. The emphasis in the article is on the "Pentagon-sponsored futures market in terrorism indicators (that) was announced and squashed in all of two days."
A different picture is illustrated in the second chart (directly below). The red line represents "Labour Party", the orange line "Liberal Democrats", and the blue line "Conservative Party".
Of course, this type of information may have no predictive value. The UK Polling Report (an independent survey and polling news website) shows that the majority of polls paint a closer picture, but with a distinct lead for the Conservative Party. Could the higher search (and news) reference volume for "Gordon Brown" be due to an incumbency effect? Could the higher search (and news) reference volume for "Liberal Democrats" be due to a novelty effect? Of course, we can't attempt an answer until later in the week. Many punters have already placed their bets though: bookmakers estimate that up to £40m will have been bet on this election, smashing previous records.
Finally, the Intrade prediction market indicates that the Conservative Party have roughly a 90% chance of winning, as shown in the third chart (directly below).
The Inkling Prediction Market indicates here that there is a 66% chance of a hung parliament, and a 32% chance of a win by the Conservative Party.
Hal Varian, Chief Economist at Google, discusses prediction markets here in the NYT, from a few years ago. The emphasis in the article is on the "Pentagon-sponsored futures market in terrorism indicators (that) was announced and squashed in all of two days."
Thursday, March 18, 2010
Using Google Trends to measure zombie attacks
Posted by
Liam Delaney
Marginal Revolution points to one of the more unusual pieces to be written by Andrew Gelman. Cowen particularly likes the line: "We originally wrote this article in Word, but then we converted it to Latex to make it look more like science."
I am gathering from this that Professor Gelman is sceptical about the use of Google Trends and related net harvested data. We have posted a lot of links to work that is ongoing in this area. Over time, the material has looked increasingly promising but we certainly have not seen anything yet on this blog that would knock your socks off in terms of the potential of these types of data e.g. the use of google trends to forecast unemployment is mildly interesting but doesn't look much better than using simple consumer sentiment data (though it may be cheaper of course, which is obviously a consideration). If anyone wants to defend the use of this type of data for social science from Gelman's ridicule, please feel free to suggest something to post. For now, I am going to have a nice comfortable lie-down on the fence, as I have seen nothing that makes me want to change from using good survey data but I would be surprised if people do not come up with some strong applications soon.
link here
I am gathering from this that Professor Gelman is sceptical about the use of Google Trends and related net harvested data. We have posted a lot of links to work that is ongoing in this area. Over time, the material has looked increasingly promising but we certainly have not seen anything yet on this blog that would knock your socks off in terms of the potential of these types of data e.g. the use of google trends to forecast unemployment is mildly interesting but doesn't look much better than using simple consumer sentiment data (though it may be cheaper of course, which is obviously a consideration). If anyone wants to defend the use of this type of data for social science from Gelman's ridicule, please feel free to suggest something to post. For now, I am going to have a nice comfortable lie-down on the fence, as I have seen nothing that makes me want to change from using good survey data but I would be surprised if people do not come up with some strong applications soon.
link here
Thursday, January 21, 2010
Unemployment Data and Google: From Forecasts to History Class
Posted by
Anonymous
A belated thanks to Michael Breen for pointing out a recent article on VoxEU.org about predicting unemployment using Google Trends. The article is by D'Amuri and Marcucci: "The predictive power of Google data: New evidence on US unemployment". I have followed the literature on Google-search data and unemployment ( previously, here), but wasn't aware of the work by D'Amuri and Marcucci until now.
The research mentioned on the blog before (by Varian and Choi) was conducted to predict U.S. unemployment insurance claims using Google Trend data based around keywords such as "unemployment" and "social insurance". D'Amuri and Marcucci differ in their approach by using the "Google Index" – the incidence of Google job-search related queries over total queries – proved to have predictive power in forecasting unemployment in Germany and Israel (see Askitas and Zimmermann 2009 and Suhoy 2009).
Both approaches improve the predictive power of unemployment forecasting, but each has some limitation. D'Amuri and Marcucci mention that the Google Index could be partly driven by on-the-job search, rather than unemployed job search activities. It can be argued that this is only a problem to the extent that on-the-job search happens during a recession. D'Amuri and Marcucci also consider that not everyone has access to the internet, and therefore that people using the internet for job search are not randomly selected among job-seekers. It follows that people using the internet for unemployment benefit information are not randomly selected among the newly unemployed. These are illustrations of the sample selection problem in econometric analysis (as distinct from self-selection; Heckman provides an overview here).
So while Google Trends is a useful tool in unemployment (and other) predictions, there are reasons to be cautious; in particular, in relation to sample selection. Another limitation that has been remarked upon is that Google Trends only provides data from 2004 onwards. This is why I was intrigued when I typed "unemployment" into Google earlier today, investiagted the "options" at the top of the page, and then clicked "timeline" instead of "standard view". What I got was a picture along the lines of the one below, except that the chart was for unemployment from 1900-2010, instead of "Book of Revelation". I was unable to take a screen-grab of the unemployment chart; but you can follow a link to it here.

It's possible to click on any decade in the chart, any year and any month. Associated news stories appear in each of these categories. This is a powerful tool for finding out what was beeing reported in the media at the time any major news story was being covered. All of this is powered by Google News Timeline: a web application that organises information chronologically. "It allows users to view news and other data sources on a zoomable, graphical timeline. You can navigate through time by dragging the timeline, setting the "granularity" to weeks, months, years, or decades, or just including a time period in your query." However, using the News Timeline application directly is somewhat different to using the "timeline view" in web search. The latter provides charts such as the one shown above.
The unemployment chart shows two spikes: one at the start of the 1930's, and one in 2009. But what does this mean? According to the Google Blog, "the graph across the top of the page summarizes how dates in your results are spread through time, with higher bars representing a larger number of unique dates." Where does this historical data come from? From Google's "News Archive Search" service. News Archive Search produces the same results as the "timeline view" in web search. Search results include content from a number of sources, including both partner content digitized by Google through their News Archive Partner Program and online archival materials. More information about News Archive Search is available here.
There are some parallels to be drawn between News Archive Search (particularly the associated graphical illustrations) and the "news reference volume" feature in Google Trends. However, it is important to note that Google's "timeline" news graphs are based on monthly data-points; not daily data-points such as those used in (Trends) news reference volume. According to Google, Archive Search works as follows: "Articles related to a single story within a given time period are grouped together to allow users to see a broad perspective on the topics they are searching."
The research mentioned on the blog before (by Varian and Choi) was conducted to predict U.S. unemployment insurance claims using Google Trend data based around keywords such as "unemployment" and "social insurance". D'Amuri and Marcucci differ in their approach by using the "Google Index" – the incidence of Google job-search related queries over total queries – proved to have predictive power in forecasting unemployment in Germany and Israel (see Askitas and Zimmermann 2009 and Suhoy 2009).
Both approaches improve the predictive power of unemployment forecasting, but each has some limitation. D'Amuri and Marcucci mention that the Google Index could be partly driven by on-the-job search, rather than unemployed job search activities. It can be argued that this is only a problem to the extent that on-the-job search happens during a recession. D'Amuri and Marcucci also consider that not everyone has access to the internet, and therefore that people using the internet for job search are not randomly selected among job-seekers. It follows that people using the internet for unemployment benefit information are not randomly selected among the newly unemployed. These are illustrations of the sample selection problem in econometric analysis (as distinct from self-selection; Heckman provides an overview here).
So while Google Trends is a useful tool in unemployment (and other) predictions, there are reasons to be cautious; in particular, in relation to sample selection. Another limitation that has been remarked upon is that Google Trends only provides data from 2004 onwards. This is why I was intrigued when I typed "unemployment" into Google earlier today, investiagted the "options" at the top of the page, and then clicked "timeline" instead of "standard view". What I got was a picture along the lines of the one below, except that the chart was for unemployment from 1900-2010, instead of "Book of Revelation". I was unable to take a screen-grab of the unemployment chart; but you can follow a link to it here.
It's possible to click on any decade in the chart, any year and any month. Associated news stories appear in each of these categories. This is a powerful tool for finding out what was beeing reported in the media at the time any major news story was being covered. All of this is powered by Google News Timeline: a web application that organises information chronologically. "It allows users to view news and other data sources on a zoomable, graphical timeline. You can navigate through time by dragging the timeline, setting the "granularity" to weeks, months, years, or decades, or just including a time period in your query." However, using the News Timeline application directly is somewhat different to using the "timeline view" in web search. The latter provides charts such as the one shown above.
The unemployment chart shows two spikes: one at the start of the 1930's, and one in 2009. But what does this mean? According to the Google Blog, "the graph across the top of the page summarizes how dates in your results are spread through time, with higher bars representing a larger number of unique dates." Where does this historical data come from? From Google's "News Archive Search" service. News Archive Search produces the same results as the "timeline view" in web search. Search results include content from a number of sources, including both partner content digitized by Google through their News Archive Partner Program and online archival materials. More information about News Archive Search is available here.
There are some parallels to be drawn between News Archive Search (particularly the associated graphical illustrations) and the "news reference volume" feature in Google Trends. However, it is important to note that Google's "timeline" news graphs are based on monthly data-points; not daily data-points such as those used in (Trends) news reference volume. According to Google, Archive Search works as follows: "Articles related to a single story within a given time period are grouped together to allow users to see a broad perspective on the topics they are searching."
Tuesday, July 28, 2009
Google Trends and Irish Unemployment: Searching For The "Dole"?
Posted by
Anonymous
Liam recently mentioned some new research by Google that forecasts unemployment. The Google research notes that one of the strongest leading indicators of economic activity is the number of people who file for unemployment benefits. Google forecast initial claims (for unemployment benefit) using the past values of the time series, and then add Google Trends variables to see how much they improve the forecast. They find a 15.74% reduction in mean absolute error.
The relevant data-source in Ireland is the Live Register. In the comments section of this earlier post, I discuss the various definitions of claims for unemployment benefit in Ireland. Unfortunately Google Trends only provides information on search volume in Ireland for "job seekers benefit" since near the end of the first quarter in this year. There is data on search volume in Ireland for "job seekers allowance" since the start of this year. For the latter term there is spike at the end of the first quarter this year, which may be related to the supplementary budget of April 7th.
There is a longer time series for search volume in Ireland related to the term "dole". This data is available since the end of 2007; there appears to be an upward trend in place since last summer, as shown below.
The relevant data-source in Ireland is the Live Register. In the comments section of this earlier post, I discuss the various definitions of claims for unemployment benefit in Ireland. Unfortunately Google Trends only provides information on search volume in Ireland for "job seekers benefit" since near the end of the first quarter in this year. There is data on search volume in Ireland for "job seekers allowance" since the start of this year. For the latter term there is spike at the end of the first quarter this year, which may be related to the supplementary budget of April 7th.
There is a longer time series for search volume in Ireland related to the term "dole". This data is available since the end of 2007; there appears to be an upward trend in place since last summer, as shown below.
Sunday, June 07, 2009
Inflation: bah humbug
Posted by
Kevin Denny
More on Google Trends: Searching for the Jolly Green Giant?
Posted by
Anonymous
The Economist Blog reported in April that Google Trends showed a decline in news reference volume for all three of the terms “Great Depression,” “credit crisis,” and “layoffs.” They asked whether this suggests that the worst is behind us, and whether we can expect a pending economic recovery?
Experimenting with these keywords now shows that the decline in news reference (and search) volume has continued. However, it is possible that media operators could provide less coverage of a recession, even though it continues (the question may be: what sells newspapers?). Also, people may not search for information about a recession, even though it is happening around them (sticking one's head in the sand?).
Another approach is to experiment with keywords like "recovery" and "economic recovery". The latter has a spike in search (and news reference) volume on February 10th, when the U.S. Senate passed President Obama's economic recovery plan. The former has been trending upward in news reference volume since the start of the year.
It's difficult to say anything about so-called "green shoots" using this data; at the very least we know that Obama harnessed a lot of attention at the start of February with his plan for economic revival.
Experimenting with these keywords now shows that the decline in news reference (and search) volume has continued. However, it is possible that media operators could provide less coverage of a recession, even though it continues (the question may be: what sells newspapers?). Also, people may not search for information about a recession, even though it is happening around them (sticking one's head in the sand?).
Another approach is to experiment with keywords like "recovery" and "economic recovery". The latter has a spike in search (and news reference) volume on February 10th, when the U.S. Senate passed President Obama's economic recovery plan. The former has been trending upward in news reference volume since the start of the year.
It's difficult to say anything about so-called "green shoots" using this data; at the very least we know that Obama harnessed a lot of attention at the start of February with his plan for economic revival.
Monday, April 20, 2009
How To Improve Econometric Analysis Using Data from Google Trends - They Can Predict The Flu
Posted by
Anonymous
In the current edition of the Economist, there is an article on how data from Google Trends can help predict economic statistics before they become available. For example, using data on searches for trucks and SUVs to predict the monthly sales of motor vehicles reduces the average error by up to 18% compared with the predictions from a model that did not incorporate the search data. These findings are from a new economics paper written Hal Varian, the Chief Economist at Google, with Hyunyoung Choi, also at Google. (There is a link to the Google working paper here on the Google Research Blog).
The authors argue that fluctuations in the frequency with which people search for certain words or phrases online can improve the accuracy of the econometric models used to predict, for example, retail-sales figures or house sales. "Actual numbers for such things are usually available only with a lag. But Google’s search data are updated every day, so they can in theory capture shifts in consumer behaviour before official numbers are released."
These data are available through a site called Google Trends; this software has been discussed on the blog quite a few times: here in relation to predicting economic sentiment from search engine behaviour.
I mentioned Gord Hotchkiss from searchengineland.com, who asked in the middle of 2008 "what if our mood turns to anxiety about the future? We still search, but we search for different things. We search for information needed to help us weather the storm. Or, we search out of a desperate desire need to know just how bad things are." To illustrate, Hotchkiss presents the following Google Trend graph which shows the relative search volume and news coverage volume of "house plans" (blue line) and "foreclosures" (red line) in America over the last few years:

The Varian and Choi paper discusses how for some things, like retail sales, the categories into which Google classifies its search-trend data correspond closely to what people may want to predict, such as the sales of a particular brand of car. For others, like sales of houses, things are less clear. It appears that searches for estate agents work better than those for home financing.
Some experimentation that I have done with with the Trends software has convinced me that the selection of the keyword is a crucial consideration when trying to analyse search volume. For example, the use of "Bush", "George Bush" and "George Bush Jr" produces very different results. So how can this issue be addressed? The answer may be to find the most popular keywords related to a core question, and to aggregate these for analysis. I have yet to find an aggregation function for keywords in Google Trends, but I have discovered a website that provides information about the most popular keywords used in web searches: www.Sitepsych.com
A list of the top 200 search terms that people use, week by week or month by month, is available for free from Sitepsych. A casual inspection of the top 200 list over a 90 day period, quickly tells you that the most popular things that people are looking for on the web are sex, music, games, dogs, golf, the weather and map-directions. Sex and music dominate.
Getting back to the Google Trends software, I noted before that Google lets users get their hands dirty with the secondary data. In fact, Varian and Choi write on the Google Research Blog that they want forecasting wannabes to download some Google Trends data and try to relate it to other economic time series. If you find an interesting pattern, they invite you to post your findings on a website and send a link to econ-forecast@google.com. They will report on the most interesting results in a later blog post.
I'm thinking of putting together something on when the recession entered the public consciousness, with particular reference to Ireland. Was this a slow-burning process or where there shocks? I suspect it was largely the former but with a preliminary shock in August 2007, a subsequent shock in August 2008 and a critical threshold in November 2008. Did it come through media reference first or through search volume? Again, I suspect that it was largely the former but that there was convergence over time. If the temporal evolution is distinct, can I show that one affected the other? This seems tricky. Should I expect non-stationarity in both series? I definitely think so.
For a list of links to all the software mentioned above, and a discussion of how online search statistics may help drive Irish economic recovery, see this post from earlier on the blog: Web-based Technology and the Recovery - What Do Irish Consumers Want?
Finally, below is a video from Google.org which shows that certain search terms are good indicators of flu activity. Google Flu Trends uses aggregated Google search data to estimate flu activity up to two weeks faster than traditional flu surveillance systems. There was an article published about this in Nature during February: Detecting influenza epidemics using search engine query data.
The authors argue that fluctuations in the frequency with which people search for certain words or phrases online can improve the accuracy of the econometric models used to predict, for example, retail-sales figures or house sales. "Actual numbers for such things are usually available only with a lag. But Google’s search data are updated every day, so they can in theory capture shifts in consumer behaviour before official numbers are released."
These data are available through a site called Google Trends; this software has been discussed on the blog quite a few times: here in relation to predicting economic sentiment from search engine behaviour.
I mentioned Gord Hotchkiss from searchengineland.com, who asked in the middle of 2008 "what if our mood turns to anxiety about the future? We still search, but we search for different things. We search for information needed to help us weather the storm. Or, we search out of a desperate desire need to know just how bad things are." To illustrate, Hotchkiss presents the following Google Trend graph which shows the relative search volume and news coverage volume of "house plans" (blue line) and "foreclosures" (red line) in America over the last few years:
The Varian and Choi paper discusses how for some things, like retail sales, the categories into which Google classifies its search-trend data correspond closely to what people may want to predict, such as the sales of a particular brand of car. For others, like sales of houses, things are less clear. It appears that searches for estate agents work better than those for home financing.
Some experimentation that I have done with with the Trends software has convinced me that the selection of the keyword is a crucial consideration when trying to analyse search volume. For example, the use of "Bush", "George Bush" and "George Bush Jr" produces very different results. So how can this issue be addressed? The answer may be to find the most popular keywords related to a core question, and to aggregate these for analysis. I have yet to find an aggregation function for keywords in Google Trends, but I have discovered a website that provides information about the most popular keywords used in web searches: www.Sitepsych.com
A list of the top 200 search terms that people use, week by week or month by month, is available for free from Sitepsych. A casual inspection of the top 200 list over a 90 day period, quickly tells you that the most popular things that people are looking for on the web are sex, music, games, dogs, golf, the weather and map-directions. Sex and music dominate.
Getting back to the Google Trends software, I noted before that Google lets users get their hands dirty with the secondary data. In fact, Varian and Choi write on the Google Research Blog that they want forecasting wannabes to download some Google Trends data and try to relate it to other economic time series. If you find an interesting pattern, they invite you to post your findings on a website and send a link to econ-forecast@google.com. They will report on the most interesting results in a later blog post.
I'm thinking of putting together something on when the recession entered the public consciousness, with particular reference to Ireland. Was this a slow-burning process or where there shocks? I suspect it was largely the former but with a preliminary shock in August 2007, a subsequent shock in August 2008 and a critical threshold in November 2008. Did it come through media reference first or through search volume? Again, I suspect that it was largely the former but that there was convergence over time. If the temporal evolution is distinct, can I show that one affected the other? This seems tricky. Should I expect non-stationarity in both series? I definitely think so.
For a list of links to all the software mentioned above, and a discussion of how online search statistics may help drive Irish economic recovery, see this post from earlier on the blog: Web-based Technology and the Recovery - What Do Irish Consumers Want?
Finally, below is a video from Google.org which shows that certain search terms are good indicators of flu activity. Google Flu Trends uses aggregated Google search data to estimate flu activity up to two weeks faster than traditional flu surveillance systems. There was an article published about this in Nature during February: Detecting influenza epidemics using search engine query data.
Subscribe to:
Posts (Atom)
