Showing posts with label information. Show all posts
Showing posts with label information. Show all posts

Sunday, December 11, 2011

Using the Internet to Predict the Future

I blogged recently about the predictive power (or not) of Twitter: from marketing to finance; and X-Factor to elections. While there may be skepticism about the predictive power of social networks; (perhaps more so for finance than for marketing/elections/popular culture); even when using an ostensibly professional network such as Twitter; there is no doubt that the medium produces a lot of user-generated information. However, Twitter-users are a select sample: a point emphasised in a recent study by Yahoo! Research. Nonetheless, the production, flow, and consumption of information will undoubtedly be an interesting area of economic research to follow in the future. Indeed, such information is not confined to networks such as Twitter; the whole internet is a veritable goldmine of potentially predictive data.

This possibility has been tapped into before by researchers using Google Trends. I posted on the old Geary blog about a comparison of Google Trends data with polls, bookmakers' odds and prediction markets (based on the British Election): which showed that it is important to be careful when interpreting search data. In particular, I noted before that it is important to distinguish between "interest" and "intent". It is clear that understanding more about search is a big challenge: for the search-engine based advertising business, and for social scientists. Search data is (or should be) interesting to academics; principally because we don't ask people for the information that they provide in search queries. No matter how well-designed surveys are, there will always be things in the ether, trends in society, that will potentially appear in search data first.

Of course, anyone interested in predicting the future should be poring over search data, social-network data, and whatever else they can find on the internet. That is exactly what a company called Recorded Future does. A recent article in the New York Times says:
"A company called Recorded Future looks at 100,000 Web pages an hour, scanning across 50,000 sources that include everything from Securities and Exchange Commission filings to Twitter comments. The idea is to look for statements about the future, like notice of an annual meeting or predictions about when a product might be released, look at past developments and then create a temporal index that suggests trends... its clients have included government agencies and banks. Its products include a $9,000-a-month service for hedge funds that plugs Recorded Future’s insights into their trading networks... (it) also started offering a Web-based version of its product on a subscription basis for $149 a month."
According to the NYT article: "two... key competitors in the Web-based predictions business (are) Palantir Technologies and Quid. Aside from those companies, the open-source statistical programming language known as R is being used as a cheap way to make statistical inference in our data-drenched world; a company called Revolution Analytics sells a commercial version to financial companies and manufacturers, among others."

Wired Magazine ran a piece on Recorded Future last month; saying:
"They aren't traders... but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future."
As the promos say: "unleash all that mankind knows about the future". It's what they used to call "the wisdom of crowds".

Postscript: The Salfordian reports that: "the Living Earth Simulator Project (LES) aims to ‘simulate everything’ on the planet, using anything from tweets to government statistics to map out social trends and predict the next economic crisis... The European Commission has... put the Living Earth Simulator at the top of its shortlist for £900m in funding." Also, DCU PhD graduate Adam Bermingham won this year's Irish Software Association award for a student project with the greatest commercial potential. Adam researched, designed and implemented a real-time sentiment monitoring system, SentiSense, to determine how people value different types of opinion when they are monitoring real-time social media content.

Sunday, November 20, 2011

Whatever Will Tweet Will Be?

A story on China Daily USA came to my attention this evening: about Zhong Lin, a graduate of Tsinghua University; and Zhao Siqi, from Hong Kong University. Both are engineers at Rice University, who designed a computer program that analyses tweets in real time. They hope to use it to predict the winner of the next US presidential election. To date, Lin and Siqi have been working on a project called SportSense, which examines tweets posted by NFL fans to infer what is happening in a game, and how excited the fans are. "It does so in real-time and provides visualized results for live games."

"SportSense is part of a larger project that aims to utilize people as sensors to infer what is happening in the physical world and what people feel about it." SportSense is not the first project to harness the power of Twitter, of course. I posted on the old Geary blog about a Dublin-based start-up called "WePredict". Kevin posted on this blog a few months ago about a study in Science which shows that work, sleep and the amount of daylight people are exposed to all affect mood. There has also been work (by InboxQ) on where the highest concentration of tweeters with the most knowledge about a specific topic are located. WiseWindow is a marketing firm that uses social-media activity to forecast demand for products.

Another group to keep an eye on is Derwent Capital Markets. They use Twitter sentiment to manage their hedge fund. There's been a lot of demand for the fund, according to this article. The strategy is based on an academic study by Johan Bollen (Indiana University), Huina Mao (Indiana University), and Xiao-Jun Zeng (University of Manchester) that established the connection between emotion-related words appearing in Twitter posts and subsequent movements in the Dow Jones Industrial Average. Here's the original research paper.

This article is critical of the approach taken by Derwent. It makes a number of points, one of which is: "Beyond the difficulty of assigning sentiment to tweets, there's a much bigger issue at play. If you look at the patterns of tweets what you find is that most are reactive rather than proactive... Twitter sentiment is likely to be a lagging indicator, at least in the real-time world of algo trading."

Nonetheless, this is an area which has garnered a lot of interest. This BBC story mentions a PhD student at Munich who has done similar work on predicting the stock market (and elections) with Twitter sentiment. The Economist had a piece on the topic in their second Technology Quarterly for this year. That article raised a number of interesting issues; such as the role of meaning in Twitter updates:
Humans excel at extracting meaning and sentiment from even the tiniest snippets of text, a task that stumps machines. To a computer, a tweet that reads “Feeling joyful after my trip to the dentist. Yeah, really” says that the author has been to the dentist and is now happy. Researchers have recently made strides in teaching machines to recognise such sarcasm, as well as double meanings or cultural references. In February Watson, a supercomputer devised by IBM, trounced two human champions at “Jeopardy!”, an American quiz show renowned for the way its clues are laden with ambiguity, irony, riddles and puns. But, for the most part, processing natural language remains a challenge.
While there may be skepticism about the predictive power of Twitter (especially for use in the domains of marketing and finance), there is no doubt that the medium produces a lot of user-generated information. While I am not (yet) a Twitter user, I have been keen to tap into it as a source of information for some time now. Recently, I found the means to do so: inagist is a Twitter-based news-service; probably as useful to both Twitter users and non-users alike. Even better is TweetMinster (due to its automatic updating): it's essentially a twitter-feed about current affairs (London-orientated). I'm following their live feed on breaking news.

TweetMinster tracks "the content most shared between expert users on Twitter and (we) use that data to discover and organise content for our news platform... this is what politicians, civil servants, activists, academics, business analysts and journalists think is the most important news of the day... We also feature live feeds of relevant twitter posts by the expert networks we track, so that you can follow the breaking news stories, big events and trending topics live – even if you’re not on Twitter..."

Addendum: It turns out that Twitter is only part of the story. I just read about a company called "Recorded Future" and blogged about them here: Using the Internet to Predict the Future.

Thursday, July 01, 2010

School quality information & school choice

The main purpose of providing information on school characteristics is that parents & students can make informed decisions about school choice. It is, after all, one of the most important decisions that parents make for their kids. All the more curious that the law in Ireland prohibits the publication of such information. But I digress.
The paper below looks at whether information on school quality does influence decisions about school choice. The effects seem rather small with students only prepared to travel 200 metres more to attend an above average school. Perhaps, there is very little variation in school quality or perhaps the Dutch are pathologically lazy.

Ranking the Schools: How Quality Information Affects School Choice in the Netherlands

Koning, Pierre &van der Wiel, Karen

This paper analyzes whether information on high school quality published by a national newspaper affects school choice in the Netherlands. For this purpose, we use both school level and individual student level data. First, we study the causal effect of quality scores on the influx of new high school students using a longitudinal school dataset. We find that negative (positive) school quality scores decrease (increase) the number of students choosing a school after the year of publication. The positive effects are particularly large for the academic school track. An academic school track receiving the most positive score sees its inflow of students rise by 15 to 20 students. Second, we study individual school choice behaviour to address the relative importance of the quality scores, as well as potential differences in the quality response between socio-economic groups.

Although the probability of attending a school is affected by its quality score, it is mainly driven by the travelling distance.

Students are only willing to travel about 200 meters more in order to attend a well-performing rather than an average school. In contrast to equity concerns that are often raised, we cannot find differences in information responses between socio-economic groups.