Showing posts with label internet economics. Show all posts
Showing posts with label internet economics. Show all posts

Sunday, December 11, 2011

Using the Internet to Predict the Future

I blogged recently about the predictive power (or not) of Twitter: from marketing to finance; and X-Factor to elections. While there may be skepticism about the predictive power of social networks; (perhaps more so for finance than for marketing/elections/popular culture); even when using an ostensibly professional network such as Twitter; there is no doubt that the medium produces a lot of user-generated information. However, Twitter-users are a select sample: a point emphasised in a recent study by Yahoo! Research. Nonetheless, the production, flow, and consumption of information will undoubtedly be an interesting area of economic research to follow in the future. Indeed, such information is not confined to networks such as Twitter; the whole internet is a veritable goldmine of potentially predictive data.

This possibility has been tapped into before by researchers using Google Trends. I posted on the old Geary blog about a comparison of Google Trends data with polls, bookmakers' odds and prediction markets (based on the British Election): which showed that it is important to be careful when interpreting search data. In particular, I noted before that it is important to distinguish between "interest" and "intent". It is clear that understanding more about search is a big challenge: for the search-engine based advertising business, and for social scientists. Search data is (or should be) interesting to academics; principally because we don't ask people for the information that they provide in search queries. No matter how well-designed surveys are, there will always be things in the ether, trends in society, that will potentially appear in search data first.

Of course, anyone interested in predicting the future should be poring over search data, social-network data, and whatever else they can find on the internet. That is exactly what a company called Recorded Future does. A recent article in the New York Times says:
"A company called Recorded Future looks at 100,000 Web pages an hour, scanning across 50,000 sources that include everything from Securities and Exchange Commission filings to Twitter comments. The idea is to look for statements about the future, like notice of an annual meeting or predictions about when a product might be released, look at past developments and then create a temporal index that suggests trends... its clients have included government agencies and banks. Its products include a $9,000-a-month service for hedge funds that plugs Recorded Future’s insights into their trading networks... (it) also started offering a Web-based version of its product on a subscription basis for $149 a month."
According to the NYT article: "two... key competitors in the Web-based predictions business (are) Palantir Technologies and Quid. Aside from those companies, the open-source statistical programming language known as R is being used as a cheap way to make statistical inference in our data-drenched world; a company called Revolution Analytics sells a commercial version to financial companies and manufacturers, among others."

Wired Magazine ran a piece on Recorded Future last month; saying:
"They aren't traders... but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future."
As the promos say: "unleash all that mankind knows about the future". It's what they used to call "the wisdom of crowds".

Postscript: The Salfordian reports that: "the Living Earth Simulator Project (LES) aims to ‘simulate everything’ on the planet, using anything from tweets to government statistics to map out social trends and predict the next economic crisis... The European Commission has... put the Living Earth Simulator at the top of its shortlist for £900m in funding." Also, DCU PhD graduate Adam Bermingham won this year's Irish Software Association award for a student project with the greatest commercial potential. Adam researched, designed and implemented a real-time sentiment monitoring system, SentiSense, to determine how people value different types of opinion when they are monitoring real-time social media content.

Friday, December 10, 2010

The Psychology of Facebook

Of the 20 most popular websites in the world, 13 relate to corporations that have European headquarters in Ireland: Ebay, Google, Yahoo and Facebook. The last three of these corporations have been the source of much commentary on this blog. Facebook (FB) has 1.5 millions users in Ireland and it is competing strongly with Google for web-user time-allocation. The chart below (from comScore/Citi) shows FB's increasing share of web-user time-allocation since 2006. FB has an explicit interest in internet economics and it has been mentioned many times before on this blog: from ambient awareness to the App Economy, FB search data to Mulley Com's study of FB eye-tracking, the choice architecture of FB, the FB Global Happinness Index, and of course, the debate on whether FB-use hurts students' grades.


Earlier this week in the Irish Times, Eoin Burke Kennedy wrote an interesting article about the psychology of FB. "Extroverted people tend to have more friends on Facebook but reveal less about themselves while introverts disclose more personal details but to a smaller group." Extraversion is one of the Big 5 personality traits: often discussed on this blog in the context of non-cognitive ability. Here and here, for example. The Big 5 personality traits have also been discussed on this blog before - in relation to social networking and web-users' choice of email address. Readers who find these topics interesting may want to read about the developing field of cyberpsychology - this subject encompasses all the psychological phenomena that are associated with or affected by emerging technology.

The content of the Irish Times article (mentioned above) is based on work by Dublin Business School psychologist Dr Ciarán McMahon. "McMahon has conducted an extensive review of the psychological literature on Facebook to better understand what makes it tick". McMahon runs a blog, PsychBook Research, which links to one of his recent presentations: "Facebook and psychology: What we know so far". While the Geary Blog has often mentioned the career opportunities for economists in corporations such as Google, Yahoo, and Microsoft; and opportunities for applied economists in private sector firms such as Netflix, SeatGeek, Yapta, Inon and Nielsen; opportunities for psychologists seem likely in Facebook, and in other Web 2.0 commerce.

Wednesday, November 03, 2010

What Can Search Predict?

I have discussed the economics of internet search in detail before. Last month, Yahoo! Labs published a news item on how what people are searching for today can be predictive of what they will do in the near future. Below is an excerpt.
"This week research scientists at Yahoo! Labs published a paper in the Proceedings of the National Academy of Sciences that examines the possibility of using web search data to predict consumer behavior. Their results have captured the public imagination and the attention of more than a few media outlets, including Technology Review, ARS Technica, Reuters and the BBC. Today, study co-author Sharad Goel shares his thoughts on the team’s conclusions."
This is an interesting advancement. As has been discussed on this blog before, recent work has demonstrated that Web search volume can “predict the present,” meaning that it can be used to accurately track outcomes such as unemployment levels, auto and home sales, and disease prevalence in near real time. As the PNAS paper describes, the new research by Yahoo! Labs shows "that what consumers are searching for online can also predict their collective future behavior days or even weeks in advance".

Thursday, May 06, 2010

The Economics of Internet Search

The comparison of Google Trends data (on the British Election) with polls, bookmakers' odds and prediction markets shows once again that we need to be careful when interpreting search volume data. While the (search-data related) innovations in unemployment forecasting may not be earth-shattering, what else can we learn from trends in seach queries? There needs to be a focus on what the analyst expects when typing in "zombie" or "inflation" or "dole" into a trend-analyser. Hopefully we can all agree that there is no such thing as a zombie. Unemployed people on the other hand, are a very real human problem, and growing in large numbers.

For example, do unemployed individuals looking for information about welfare payments type in "unemployment" or "dole"? Or something else? One experiment (view here) is to type in "unemployment", "dole" and "jobs" into Google Trends, separated by commas. A few observations can be made:

(i) The search volume for "jobs" is relatively stable over the last 6 years
(ii) News reference volume for "jobs" has exploded over the last 2 years, much more so than for "unemployment"
(iii) There is only enough search activity related to "unemployment" for it to register half-way during 2008
(iv) There is only enough search activity related to "dole" for it to register at the start of 2009
(v) There is a fall-off in search volume for "jobs" at the end of every calendar year

While much of this mirrors what we already know about recent economic activity, I had expected "jobs" to have a much higher search volume over the last year. We of course have to be very careful about drawing conclusions, but the stylised facts about search volume suggest that there were more people searching for jobs in 2004 and 2005 than there were in 2008 and 2009. We know that there were more people in need of a job in 2008 and 2009, so what is the explanation? Perhaps job-search is more intense during boom-times. In recessions, maybe people are less likely to search for a job (which they simply believe isn't there). This could of course be incorrect, but now there is an open question.

Other challenging questions about search data are currently at play in the commercial arena; it may be no coincidence that Google Trends was opened up to the public (including academics) in 2006, just as these questions were coming more to the fore. At present, Google, Yahoo and Bing are strongly focused on distinguishing between "interest" and "intent" in search data. There are obvious commercial implications, but solving this problem about interest versus intent would also help academic researchers. When somebody searches for "jobs" do they just want to *see what's out there* (maybe in boom times) or do they *desperately intend* to obtain employment (maybe in recessions)?

Maybe additional keywords would help in solving this interesting puzzle. If you search for "XBOX Price", Google can assume to some extent that you intend to buy an XBOX. Here is an article from last year about Google executives stating that "understanding people, health, communication, education and knowledge" is the next frontier of search. Here is a link to Yahoo!'s "Mindset" research project on 'Intent-driven Search'. Recently, there was an article in the Economist about about Qi Lu: the man behind Bing. According to him, the focus is firmly on "understanding user intent".

It's clear that understanding more about search is the big challenge: for the search-engine based advertising business, and for social scientists. And here is the main reason why search data is (or should be) so interesting for academics: we don't *ask* people for the information they provide in search queries. It's a simple statement, but it has merit. No matter how well-designed surveys are, there will always be things in the ether, trends in society, that will potentially appear in search data first.

Sunday, March 14, 2010

A Few Links

1. MySpace has allowed a large quantity of bulk user data to be put up for sale: user playlists, mood updates, mobile updates, photos, vents, reviews, blog posts, names and zipcodes. Friend lists are not included.

2. InfoChimps is a bulk data marketplace with more than 5000 data sets in its catalog so far. The vast majority are free.

3. A Mulley Communications/NCI study shows that people do not pay much attention to more than three results on a Google search result page. Also that the ads on the right hand side of the results page are barely looked at.

4. Here is a video of heatmaps being generated based on eye movements.

5. It's worth looking at data-series on capital expenditure: in this case in the UK. A pick-up in this series should lead improvements in employment figures.

6. Ellerdale, still in alpha testing, tracks data sources from around the web, primarily Twitter, and examines what topics are being discussed. It then organizes these conversations into categories like "people," "sports," "politics," "music," "television," and more.

7. The most rented movie in Chicago last year? "The Curious Case of Benjamin Button." This app on the NTY website allows for examination of Netflix rental patterns, neighborhood by neighborhood, in a dozen cities.

8. Netflix have announced that they canceling plans for a second Netflix Prize contest, one that would have involved the release of more information than the first.

9. Privacy concerns were an issue in the Netflix decision. Among the first to draw attention to the issue was University of Colorado law professor Paul Ohm, who said: "Researchers have known for more than a decade that gender plus ZIP code plus birthdate uniquely identifies a significant percentage of Americans (87% according to Latanya Sweeney's famous study)."

10. Here's a recent paper by Ohm on re-identification of individual data.

11. The deadline for Facebook's Ph.D. Fellowship Program has passed, but it's interesting to note that they have a particular focus on Internet Economics.