Showing posts with label technology and economics. Show all posts
Showing posts with label technology and economics. Show all posts
Wednesday, December 01, 2010
Economics Research at Microsoft
Posted by
Anonymous
I've blogged before about economics research at Google and Yahoo. Another major technology company with a research area devoted to economics is Microsoft. Within this stream of research, topics of particular interest include game theory, research on the family, financial services for low-SES groups, household well-being, socio-economic mobility, social networking and trust/reputation. Microsoft staffers associated with economic research are shown here.
Monday, April 20, 2009
How To Improve Econometric Analysis Using Data from Google Trends - They Can Predict The Flu
Posted by
Anonymous
In the current edition of the Economist, there is an article on how data from Google Trends can help predict economic statistics before they become available. For example, using data on searches for trucks and SUVs to predict the monthly sales of motor vehicles reduces the average error by up to 18% compared with the predictions from a model that did not incorporate the search data. These findings are from a new economics paper written Hal Varian, the Chief Economist at Google, with Hyunyoung Choi, also at Google. (There is a link to the Google working paper here on the Google Research Blog).
The authors argue that fluctuations in the frequency with which people search for certain words or phrases online can improve the accuracy of the econometric models used to predict, for example, retail-sales figures or house sales. "Actual numbers for such things are usually available only with a lag. But Google’s search data are updated every day, so they can in theory capture shifts in consumer behaviour before official numbers are released."
These data are available through a site called Google Trends; this software has been discussed on the blog quite a few times: here in relation to predicting economic sentiment from search engine behaviour.
I mentioned Gord Hotchkiss from searchengineland.com, who asked in the middle of 2008 "what if our mood turns to anxiety about the future? We still search, but we search for different things. We search for information needed to help us weather the storm. Or, we search out of a desperate desire need to know just how bad things are." To illustrate, Hotchkiss presents the following Google Trend graph which shows the relative search volume and news coverage volume of "house plans" (blue line) and "foreclosures" (red line) in America over the last few years:

The Varian and Choi paper discusses how for some things, like retail sales, the categories into which Google classifies its search-trend data correspond closely to what people may want to predict, such as the sales of a particular brand of car. For others, like sales of houses, things are less clear. It appears that searches for estate agents work better than those for home financing.
Some experimentation that I have done with with the Trends software has convinced me that the selection of the keyword is a crucial consideration when trying to analyse search volume. For example, the use of "Bush", "George Bush" and "George Bush Jr" produces very different results. So how can this issue be addressed? The answer may be to find the most popular keywords related to a core question, and to aggregate these for analysis. I have yet to find an aggregation function for keywords in Google Trends, but I have discovered a website that provides information about the most popular keywords used in web searches: www.Sitepsych.com
A list of the top 200 search terms that people use, week by week or month by month, is available for free from Sitepsych. A casual inspection of the top 200 list over a 90 day period, quickly tells you that the most popular things that people are looking for on the web are sex, music, games, dogs, golf, the weather and map-directions. Sex and music dominate.
Getting back to the Google Trends software, I noted before that Google lets users get their hands dirty with the secondary data. In fact, Varian and Choi write on the Google Research Blog that they want forecasting wannabes to download some Google Trends data and try to relate it to other economic time series. If you find an interesting pattern, they invite you to post your findings on a website and send a link to econ-forecast@google.com. They will report on the most interesting results in a later blog post.
I'm thinking of putting together something on when the recession entered the public consciousness, with particular reference to Ireland. Was this a slow-burning process or where there shocks? I suspect it was largely the former but with a preliminary shock in August 2007, a subsequent shock in August 2008 and a critical threshold in November 2008. Did it come through media reference first or through search volume? Again, I suspect that it was largely the former but that there was convergence over time. If the temporal evolution is distinct, can I show that one affected the other? This seems tricky. Should I expect non-stationarity in both series? I definitely think so.
For a list of links to all the software mentioned above, and a discussion of how online search statistics may help drive Irish economic recovery, see this post from earlier on the blog: Web-based Technology and the Recovery - What Do Irish Consumers Want?
Finally, below is a video from Google.org which shows that certain search terms are good indicators of flu activity. Google Flu Trends uses aggregated Google search data to estimate flu activity up to two weeks faster than traditional flu surveillance systems. There was an article published about this in Nature during February: Detecting influenza epidemics using search engine query data.
The authors argue that fluctuations in the frequency with which people search for certain words or phrases online can improve the accuracy of the econometric models used to predict, for example, retail-sales figures or house sales. "Actual numbers for such things are usually available only with a lag. But Google’s search data are updated every day, so they can in theory capture shifts in consumer behaviour before official numbers are released."
These data are available through a site called Google Trends; this software has been discussed on the blog quite a few times: here in relation to predicting economic sentiment from search engine behaviour.
I mentioned Gord Hotchkiss from searchengineland.com, who asked in the middle of 2008 "what if our mood turns to anxiety about the future? We still search, but we search for different things. We search for information needed to help us weather the storm. Or, we search out of a desperate desire need to know just how bad things are." To illustrate, Hotchkiss presents the following Google Trend graph which shows the relative search volume and news coverage volume of "house plans" (blue line) and "foreclosures" (red line) in America over the last few years:
The Varian and Choi paper discusses how for some things, like retail sales, the categories into which Google classifies its search-trend data correspond closely to what people may want to predict, such as the sales of a particular brand of car. For others, like sales of houses, things are less clear. It appears that searches for estate agents work better than those for home financing.
Some experimentation that I have done with with the Trends software has convinced me that the selection of the keyword is a crucial consideration when trying to analyse search volume. For example, the use of "Bush", "George Bush" and "George Bush Jr" produces very different results. So how can this issue be addressed? The answer may be to find the most popular keywords related to a core question, and to aggregate these for analysis. I have yet to find an aggregation function for keywords in Google Trends, but I have discovered a website that provides information about the most popular keywords used in web searches: www.Sitepsych.com
A list of the top 200 search terms that people use, week by week or month by month, is available for free from Sitepsych. A casual inspection of the top 200 list over a 90 day period, quickly tells you that the most popular things that people are looking for on the web are sex, music, games, dogs, golf, the weather and map-directions. Sex and music dominate.
Getting back to the Google Trends software, I noted before that Google lets users get their hands dirty with the secondary data. In fact, Varian and Choi write on the Google Research Blog that they want forecasting wannabes to download some Google Trends data and try to relate it to other economic time series. If you find an interesting pattern, they invite you to post your findings on a website and send a link to econ-forecast@google.com. They will report on the most interesting results in a later blog post.
I'm thinking of putting together something on when the recession entered the public consciousness, with particular reference to Ireland. Was this a slow-burning process or where there shocks? I suspect it was largely the former but with a preliminary shock in August 2007, a subsequent shock in August 2008 and a critical threshold in November 2008. Did it come through media reference first or through search volume? Again, I suspect that it was largely the former but that there was convergence over time. If the temporal evolution is distinct, can I show that one affected the other? This seems tricky. Should I expect non-stationarity in both series? I definitely think so.
For a list of links to all the software mentioned above, and a discussion of how online search statistics may help drive Irish economic recovery, see this post from earlier on the blog: Web-based Technology and the Recovery - What Do Irish Consumers Want?
Finally, below is a video from Google.org which shows that certain search terms are good indicators of flu activity. Google Flu Trends uses aggregated Google search data to estimate flu activity up to two weeks faster than traditional flu surveillance systems. There was an article published about this in Nature during February: Detecting influenza epidemics using search engine query data.
Thursday, January 22, 2009
Web-based Technology and the Recovery - What Do Irish Consumers Want?
Posted by
Anonymous
In keeping with the recent theme of Irish economic recovery, I have been wondering how existing (and not-yet-developed) web-based technologies can be geared towards enterprise and the enhancement of existing business activities. In particular I have been thinking about technologies which summarise trends in online search and media covearge - often these are provided for free, in the spirit of open-source software.
Now more than ever, there is a need for market research to find out what Irish people will buy during a recession (Fergal mentioned cosmetics recently). Also, what is it that Irish people really want - what are the new product developments that could be rolled out to re-invigorate consumer spending?
According to Genevieve Carbery in last week's Irish Times, social networking site Bebo was the most searched-for term in Ireland in 2008. But rival Facebook was the fastest-growing search term last year, while Polish social networking site Nasza Klasa topped the rising searches list in Limerick and Galway.
These statistics may change during 2009, but it is possible to keep track of developments using Google’s 'Insight for Search' website (www.google.com/ insights/search/), which provides information on the most popular search terms in different geographical categories and time frames. For the moment, it seems that there is a definite demand for "ambient awareness".
We have discussed the Google Trends software (similar to Search for Insights) quite enthusiastically on this blog before, again here, and also Quantcast Web User Demographics and Alexa.com. The Google Trends software has been enhanced to let users get their hands dirty with the secondary data, which potentially could be extremely useful. Writing here on the Google Research Blog, Heej Hwang from the Google Trends team describes how the Trends data can be downloaded to a .csv file (a common format to import/export data), which can be opened in most spreadsheet applications (or easily converted to do so).
The technology developed by Summize Labs is also worth considering. (Summize was recently acquired by Twitter, which I just discovered here). One used to be able to enter a topic in the Summize Labs search engine to find up-to-the-second "tweets" about that topic, then automatically analyze the attitudes expressed in the "tweets". (As an example, the last time I looked, the overall sentiment on Obama was "swell"). This could prove to be a very powerful tool for political scientists, marketers and all kinds of researchers. Twitter will be adding search and its related features to its core offering in the very near future.
We've also mentioned before that Google's chief economist is predicting that the brightest graduates in economics will seek their fortune in marketing in the years ahead. This will be due to the Internet giving companies the information-rich environment once available only in financial markets (see post here).
In June, searches for the word “recession” peaked at 30 times the 2007 average. There were almost three times as many searches for “recession” as there were for “Celtic Tiger” in 2008. Maybe there will be more searches for "recovery" in 2009?
Now more than ever, there is a need for market research to find out what Irish people will buy during a recession (Fergal mentioned cosmetics recently). Also, what is it that Irish people really want - what are the new product developments that could be rolled out to re-invigorate consumer spending?
According to Genevieve Carbery in last week's Irish Times, social networking site Bebo was the most searched-for term in Ireland in 2008. But rival Facebook was the fastest-growing search term last year, while Polish social networking site Nasza Klasa topped the rising searches list in Limerick and Galway.
These statistics may change during 2009, but it is possible to keep track of developments using Google’s 'Insight for Search' website (www.google.com/ insights/search/), which provides information on the most popular search terms in different geographical categories and time frames. For the moment, it seems that there is a definite demand for "ambient awareness".
We have discussed the Google Trends software (similar to Search for Insights) quite enthusiastically on this blog before, again here, and also Quantcast Web User Demographics and Alexa.com. The Google Trends software has been enhanced to let users get their hands dirty with the secondary data, which potentially could be extremely useful. Writing here on the Google Research Blog, Heej Hwang from the Google Trends team describes how the Trends data can be downloaded to a .csv file (a common format to import/export data), which can be opened in most spreadsheet applications (or easily converted to do so).
The technology developed by Summize Labs is also worth considering. (Summize was recently acquired by Twitter, which I just discovered here). One used to be able to enter a topic in the Summize Labs search engine to find up-to-the-second "tweets" about that topic, then automatically analyze the attitudes expressed in the "tweets". (As an example, the last time I looked, the overall sentiment on Obama was "swell"). This could prove to be a very powerful tool for political scientists, marketers and all kinds of researchers. Twitter will be adding search and its related features to its core offering in the very near future.
We've also mentioned before that Google's chief economist is predicting that the brightest graduates in economics will seek their fortune in marketing in the years ahead. This will be due to the Internet giving companies the information-rich environment once available only in financial markets (see post here).
In June, searches for the word “recession” peaked at 30 times the 2007 average. There were almost three times as many searches for “recession” as there were for “Celtic Tiger” in 2008. Maybe there will be more searches for "recovery" in 2009?
Tuesday, December 02, 2008
Where Would You Live in London?
Posted by
Anonymous
Would it depend on the price? Or the distance to work? Or both? You would probably find these travel-time maps very useful. They were developed by Chris Lightfoot at MySociety.Org after the UK Department of Transport approached MySociety about experimenting with novel ways of re-using public sector data. One particular set of intearctive maps allows users to set both the maximum time they’re willing to commute, and the median house price they’re willing or able to pay. Slide the sliders on the last link to see constrained minimisation at work --- with the BBC Television Centre and Olympic Stadium as the focal points.
On the main page, the same can be done with the Department of Transport close to the centre of London. Try setting a maximum travel time here of one hour, and a maximum price of £500,000. You'll see that Chelsea/Kensington is blacked out, as is Hampstead Heath. (House prices are based on house sales recorded in the Land Registry for a large random sample of London postcodes, inflation adjusted to be the price as at December 2006. Journey times to work are for a week day in 2007. They were generated by screen scraping the Transport for London and Transport Direct journey planner websites.)
MySociety makes open source software, so you can get the source code for the scripts that made these maps, and MySociety will provide copies of the OpenStreetMap base mapping. (Other data requires permission from the owners). MySociety is a non-profit with a community of volunteers and (paid) open source coders. It runs most of the best-known democracy and transparency websites in the UK. One of its initiatives that blog readers may find interesting is PledgeBank. This allows people to set up a campaign or a committed behaviour where they say "I’ll do something, but only IF other people will too."
On the main page, the same can be done with the Department of Transport close to the centre of London. Try setting a maximum travel time here of one hour, and a maximum price of £500,000. You'll see that Chelsea/Kensington is blacked out, as is Hampstead Heath. (House prices are based on house sales recorded in the Land Registry for a large random sample of London postcodes, inflation adjusted to be the price as at December 2006. Journey times to work are for a week day in 2007. They were generated by screen scraping the Transport for London and Transport Direct journey planner websites.)
MySociety makes open source software, so you can get the source code for the scripts that made these maps, and MySociety will provide copies of the OpenStreetMap base mapping. (Other data requires permission from the owners). MySociety is a non-profit with a community of volunteers and (paid) open source coders. It runs most of the best-known democracy and transparency websites in the UK. One of its initiatives that blog readers may find interesting is PledgeBank. This allows people to set up a campaign or a committed behaviour where they say "I’ll do something, but only IF other people will too."
Monday, November 24, 2008
Early to Bed, Early to Rise... Depends on the TV Schedule in Your Time Zone
Posted by
Anonymous
The Chicago Journals website provides a useful summary (here) of a recent paper in the Journal of Labour Economics: Hamermesh, Daniel S., Caitlin Knowles Myers, and Mark L. Pocock (2008): “Cues for Timing and Coordination: Latitude, Letterman, and Longitude.”
...Daylight Saving Time has its roots in the Standard Time Act of 1918... Last year, Daylight Saving was extended by four weeks. Although the prime-time television schedule is a “relic of the technology of radio transmission” — it was created when signals could not be broadcast across the country — it remains a powerful cue. Reflecting on his own weekday television watching schedule, Hamermesh recalled, “I lived twenty years in the Eastern Time Zone, I used to stay up until 11:45 p.m. to watch the monologue on the Tonight Show. Living in Texas, I typically turn out the lights at 10:45 p.m., when the monologue is done."
...The authors... use...the American Time Use Survey (ATUS), which enabled them to observe how Americans split their time between their three most time-consuming activities: work, sleep, and television watching. After merging ATUS with sunrise and sunset data, the authors found that while natural daylight patterns have some effect on people’s life patterns, the demands of global business—market openings, etc. — and regular television schedules, demarcate the boundaries of most Americans’ lives.
Subscribe to:
Posts (Atom)