Showing posts with label news. Show all posts
Showing posts with label news. Show all posts

Sunday, December 11, 2011

Using the Internet to Predict the Future

I blogged recently about the predictive power (or not) of Twitter: from marketing to finance; and X-Factor to elections. While there may be skepticism about the predictive power of social networks; (perhaps more so for finance than for marketing/elections/popular culture); even when using an ostensibly professional network such as Twitter; there is no doubt that the medium produces a lot of user-generated information. However, Twitter-users are a select sample: a point emphasised in a recent study by Yahoo! Research. Nonetheless, the production, flow, and consumption of information will undoubtedly be an interesting area of economic research to follow in the future. Indeed, such information is not confined to networks such as Twitter; the whole internet is a veritable goldmine of potentially predictive data.

This possibility has been tapped into before by researchers using Google Trends. I posted on the old Geary blog about a comparison of Google Trends data with polls, bookmakers' odds and prediction markets (based on the British Election): which showed that it is important to be careful when interpreting search data. In particular, I noted before that it is important to distinguish between "interest" and "intent". It is clear that understanding more about search is a big challenge: for the search-engine based advertising business, and for social scientists. Search data is (or should be) interesting to academics; principally because we don't ask people for the information that they provide in search queries. No matter how well-designed surveys are, there will always be things in the ether, trends in society, that will potentially appear in search data first.

Of course, anyone interested in predicting the future should be poring over search data, social-network data, and whatever else they can find on the internet. That is exactly what a company called Recorded Future does. A recent article in the New York Times says:
"A company called Recorded Future looks at 100,000 Web pages an hour, scanning across 50,000 sources that include everything from Securities and Exchange Commission filings to Twitter comments. The idea is to look for statements about the future, like notice of an annual meeting or predictions about when a product might be released, look at past developments and then create a temporal index that suggests trends... its clients have included government agencies and banks. Its products include a $9,000-a-month service for hedge funds that plugs Recorded Future’s insights into their trading networks... (it) also started offering a Web-based version of its product on a subscription basis for $149 a month."
According to the NYT article: "two... key competitors in the Web-based predictions business (are) Palantir Technologies and Quid. Aside from those companies, the open-source statistical programming language known as R is being used as a cheap way to make statistical inference in our data-drenched world; a company called Revolution Analytics sells a commercial version to financial companies and manufacturers, among others."

Wired Magazine ran a piece on Recorded Future last month; saying:
"They aren't traders... but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future."
As the promos say: "unleash all that mankind knows about the future". It's what they used to call "the wisdom of crowds".

Postscript: The Salfordian reports that: "the Living Earth Simulator Project (LES) aims to ‘simulate everything’ on the planet, using anything from tweets to government statistics to map out social trends and predict the next economic crisis... The European Commission has... put the Living Earth Simulator at the top of its shortlist for £900m in funding." Also, DCU PhD graduate Adam Bermingham won this year's Irish Software Association award for a student project with the greatest commercial potential. Adam researched, designed and implemented a real-time sentiment monitoring system, SentiSense, to determine how people value different types of opinion when they are monitoring real-time social media content.

Sunday, November 20, 2011

Whatever Will Tweet Will Be?

A story on China Daily USA came to my attention this evening: about Zhong Lin, a graduate of Tsinghua University; and Zhao Siqi, from Hong Kong University. Both are engineers at Rice University, who designed a computer program that analyses tweets in real time. They hope to use it to predict the winner of the next US presidential election. To date, Lin and Siqi have been working on a project called SportSense, which examines tweets posted by NFL fans to infer what is happening in a game, and how excited the fans are. "It does so in real-time and provides visualized results for live games."

"SportSense is part of a larger project that aims to utilize people as sensors to infer what is happening in the physical world and what people feel about it." SportSense is not the first project to harness the power of Twitter, of course. I posted on the old Geary blog about a Dublin-based start-up called "WePredict". Kevin posted on this blog a few months ago about a study in Science which shows that work, sleep and the amount of daylight people are exposed to all affect mood. There has also been work (by InboxQ) on where the highest concentration of tweeters with the most knowledge about a specific topic are located. WiseWindow is a marketing firm that uses social-media activity to forecast demand for products.

Another group to keep an eye on is Derwent Capital Markets. They use Twitter sentiment to manage their hedge fund. There's been a lot of demand for the fund, according to this article. The strategy is based on an academic study by Johan Bollen (Indiana University), Huina Mao (Indiana University), and Xiao-Jun Zeng (University of Manchester) that established the connection between emotion-related words appearing in Twitter posts and subsequent movements in the Dow Jones Industrial Average. Here's the original research paper.

This article is critical of the approach taken by Derwent. It makes a number of points, one of which is: "Beyond the difficulty of assigning sentiment to tweets, there's a much bigger issue at play. If you look at the patterns of tweets what you find is that most are reactive rather than proactive... Twitter sentiment is likely to be a lagging indicator, at least in the real-time world of algo trading."

Nonetheless, this is an area which has garnered a lot of interest. This BBC story mentions a PhD student at Munich who has done similar work on predicting the stock market (and elections) with Twitter sentiment. The Economist had a piece on the topic in their second Technology Quarterly for this year. That article raised a number of interesting issues; such as the role of meaning in Twitter updates:
Humans excel at extracting meaning and sentiment from even the tiniest snippets of text, a task that stumps machines. To a computer, a tweet that reads “Feeling joyful after my trip to the dentist. Yeah, really” says that the author has been to the dentist and is now happy. Researchers have recently made strides in teaching machines to recognise such sarcasm, as well as double meanings or cultural references. In February Watson, a supercomputer devised by IBM, trounced two human champions at “Jeopardy!”, an American quiz show renowned for the way its clues are laden with ambiguity, irony, riddles and puns. But, for the most part, processing natural language remains a challenge.
While there may be skepticism about the predictive power of Twitter (especially for use in the domains of marketing and finance), there is no doubt that the medium produces a lot of user-generated information. While I am not (yet) a Twitter user, I have been keen to tap into it as a source of information for some time now. Recently, I found the means to do so: inagist is a Twitter-based news-service; probably as useful to both Twitter users and non-users alike. Even better is TweetMinster (due to its automatic updating): it's essentially a twitter-feed about current affairs (London-orientated). I'm following their live feed on breaking news.

TweetMinster tracks "the content most shared between expert users on Twitter and (we) use that data to discover and organise content for our news platform... this is what politicians, civil servants, activists, academics, business analysts and journalists think is the most important news of the day... We also feature live feeds of relevant twitter posts by the expert networks we track, so that you can follow the breaking news stories, big events and trending topics live – even if you’re not on Twitter..."

Addendum: It turns out that Twitter is only part of the story. I just read about a company called "Recorded Future" and blogged about them here: Using the Internet to Predict the Future.

Thursday, January 21, 2010

Unemployment Data and Google: From Forecasts to History Class

A belated thanks to Michael Breen for pointing out a recent article on VoxEU.org about predicting unemployment using Google Trends. The article is by D'Amuri and Marcucci: "The predictive power of Google data: New evidence on US unemployment". I have followed the literature on Google-search data and unemployment ( previously, here), but wasn't aware of the work by D'Amuri and Marcucci until now.

The research mentioned on the blog before (by Varian and Choi) was conducted to predict U.S. unemployment insurance claims using Google Trend data based around keywords such as "unemployment" and "social insurance". D'Amuri and Marcucci differ in their approach by using the "Google Index" – the incidence of Google job-search related queries over total queries – proved to have predictive power in forecasting unemployment in Germany and Israel (see Askitas and Zimmermann 2009 and Suhoy 2009).

Both approaches improve the predictive power of unemployment forecasting, but each has some limitation. D'Amuri and Marcucci mention that the Google Index could be partly driven by on-the-job search, rather than unemployed job search activities. It can be argued that this is only a problem to the extent that on-the-job search happens during a recession. D'Amuri and Marcucci also consider that not everyone has access to the internet, and therefore that people using the internet for job search are not randomly selected among job-seekers. It follows that people using the internet for unemployment benefit information are not randomly selected among the newly unemployed. These are illustrations of the sample selection problem in econometric analysis (as distinct from self-selection; Heckman provides an overview here).

So while Google Trends is a useful tool in unemployment (and other) predictions, there are reasons to be cautious; in particular, in relation to sample selection. Another limitation that has been remarked upon is that Google Trends only provides data from 2004 onwards. This is why I was intrigued when I typed "unemployment" into Google earlier today, investiagted the "options" at the top of the page, and then clicked "timeline" instead of "standard view". What I got was a picture along the lines of the one below, except that the chart was for unemployment from 1900-2010, instead of "Book of Revelation". I was unable to take a screen-grab of the unemployment chart; but you can follow a link to it here.


It's possible to click on any decade in the chart, any year and any month. Associated news stories appear in each of these categories. This is a powerful tool for finding out what was beeing reported in the media at the time any major news story was being covered. All of this is powered by Google News Timeline: a web application that organises information chronologically. "It allows users to view news and other data sources on a zoomable, graphical timeline. You can navigate through time by dragging the timeline, setting the "granularity" to weeks, months, years, or decades, or just including a time period in your query." However, using the News Timeline application directly is somewhat different to using the "timeline view" in web search. The latter provides charts such as the one shown above.

The unemployment chart shows two spikes: one at the start of the 1930's, and one in 2009. But what does this mean? According to the Google Blog, "the graph across the top of the page summarizes how dates in your results are spread through time, with higher bars representing a larger number of unique dates." Where does this historical data come from? From Google's "News Archive Search" service. News Archive Search produces the same results as the "timeline view" in web search. Search results include content from a number of sources, including both partner content digitized by Google through their News Archive Partner Program and online archival materials. More information about News Archive Search is available here.

There are some parallels to be drawn between News Archive Search (particularly the associated graphical illustrations) and the "news reference volume" feature in Google Trends. However, it is important to note that Google's "timeline" news graphs are based on monthly data-points; not daily data-points such as those used in (Trends) news reference volume. According to Google, Archive Search works as follows: "Articles related to a single story within a given time period are grouped together to allow users to see a broad perspective on the topics they are searching."

Tuesday, December 08, 2009

One Step Closer to Perfect Information

"MySpace and Facebook have both signed deals with Google to allow publicly available status updates to be indexed in real-time by the search giant...Up till now Google had only signed a similar deal with Twitter and was missing agreements with Facebook and MySpace...This means that when somebody searches for a particular topic on Google they will receive real-time updates from a variety of social media sites, as well as the usual list of search results." Full story here.