Showing posts with label analyticsx. Show all posts
Showing posts with label analyticsx. Show all posts

Thursday, December 23, 2010

Crowdsourcing Taxonomy

I had an interesting debate on twitter with Panos Ipeirotis and Anand Kulkarni about what counts as a crowdsourcing application.

Wikipeidia defines crowdsourcing as
"Crowdsourcing is the act of outsourcing tasks, traditionally performed by an employee or contractor, to an undefined, large group of people or community (a crowd), through an open call."
And they are even so nice as to give a list of crowdsourcing projects


These seem to fall into a few areas
1. Microwork. MTurk, Freelancer.com, 99designs. You go on mturk you put in a task like "tell me if this sentence is in English" and an amount "1 cent per task" and then you let the people on mturk do the classifications for you.

I think this sort of work will be huge. There are talented people sitting around willing to do something interesting for cash in their spare time just waiting to be harnessed.

Many projects exist to crowdsource advertising ideas. for example IdeaBounty. These advertising ideas seem similar to other microwork projects.

2. Games with a purpose. Galaxy finder, Foldit etc. There is a task like "count the number of spiral galaxies" you put this out there as a 'fun' task and people do it for you free, gratis and for nothing.

This is similar to citizen science projects but explicitly involves the use of fun and game mechanics to get people to do the work.

3. Market mechanisms could be thought of as a form of collective intelligence. In a market the competing forces of supply and demand result on a value being placed on an item. So for example on ebay many people have decided the value of collectible baseball cards. Or on intrade many have decided the probability Obama is to be reelected.

Is market formation a form of crowdsourcing? I think prediction markets are great and should be used more but their placing accurate values on the probability of an event is a byproduct of their action rather than the intention of those operating on it. People would rather if the odds were wrong so they could make money being right.

For some reason I don't think a prediction market is a crowdsourcing application and none are listed on the wikipedia list.

As an aside I read a brilliant blog post two days ago (that I now cannot find) that pointed out that ebay is essentially a storage device. You sell things they go into the ebay and when you want them back you go on the ebay and find the product again and buy it. Your storage costs are usually just the postage and packing charges. This is similar to the argument that there are two ways to produce cars. One is to build factories and hire people and such the other is to send boat loads of cows to Tokyo where magically you get boats filled with cars back in return.

4. Microlending
Kiva and other sites allow you to lend money to business. Say a guy want to plough my field. 10 people give you 5 dollars each so you can buy a plough. The field gets ploughed. They get 6 dollars back after harvest. This again does not seem like crowdsourcing of jobs as you are lending capital not labour.

Artists using a street performer protocol to pay to get an album made is something similar. Kickstarter is one site that does this

5. Wikies, Forums, social search engines, message boards
People on Forums, give their advise and knowledge almost always for free. It is pretty amazing when you think about it. The fact that some very clever and skilled individual is willing to take their time to tell me the arguments to a Java function are wrong is amazing. The wikipedia page does not class these as crowdsourcing but again I cannot see why not.

6. Competitions, kaggle, Netflix, DARPA, Goldcorp and InnoCentive regularly have prizes where they put out a task such as 'improve on our recomendations' or 'make a car that drives itself' and people try to do the best they can at it. I think prizes as a way to encourage innovation are also going to be massively important in the future.

7. Response to events. The guardian set up a webpage to get people to read through all the mp expenses reports to find things that looked excessive and wikileaks have used crowdsourcing to examine the millions of leaked documents and highlight the most interesting.

8. Crime detection. Crowdsourcing has been used for illegal immigrant spotting in Texas and other similar uses.

9. Missing persons. After hurricane Katrina and Steve Fossett's loss crowdsourcing was used to locate missing people.

10. Politics. Oxfam Novib, Moveon.org and other such sites try to use crowdsourcing to create a community based political movement. If these are counted as crowdsourcing ventures I do not see why hobby sites that try promote brewing or knitting are not.

11. Art. Improv everywhere, mechanical olympics, Flash mobs are all listed as crowdsourcing projects.

12. Citizen science. Digitized versions of old weather or other science reports are made available to people to increase our knowledge of past events. Even the open data movement must have some basis in crowdsourcing. If people are not expected to analyse the data there is not much point making it availible to them.

13 Cartography. Open street map, Waze and other projects hope to use peoples GPS systems to build up a map of the environment. This can include real time traffic data.


Reading through wikipedias list of crowdsourcing projects these are the types of projects that occured to me. Most classes are probably wrong and need to be split or combined. Maybe the hierarchy needs to be different. Please comment or write a rebuttal to if you are interested in what projects count as crowdsourcing.

Wednesday, January 13, 2010

Analytics X Prize Outside Bets

Black swans are not predictable but are there fairly rare events that could coour this year in Philadelphia that would skew the homicide rate in an area?

I know of nothing that would cause a sudden drop in the number of homicides in one particular area. Except a mass evacuation. There are a few things that could show up as a sudden rise though.

Terrorist attack. Philadelphia is not a well known terrorist target so any attack there would have to be pretty unpredictable. Terrorist attacks are also very rare so I do not think worth considering in a model.

Going Postal: These seem common enough in America. They have their own list here. I doubt you can predict where they will occur though. Might be worth considering.

Religious nutjobs:The solar temples and the kool aid drinkers in general tend to head off into the sticks before topping each other. So I think there is not much chance of a mass homicide by a cult in
Philadelphia.

Riots:Cities kick off on a fairly regular basis, The LA riots in 1992 resulted in 53 deaths for example. Philadelphia has had riots in the past. Riots in America seem to be mainly caused by racial tension. It might be possible to if not predict them localise where they are most likely to occur. Then submit a higher homicide count for that area in one of the analytics x prize submissions as an outside bet.

Prisoneers
: They love a good ruck. In general prison populations have a higher homicide rate. So looking at changes in prisoner ecosystem in Philadelphia might be worth a look

Natural Disasters:After natural disaster people generally think the world turns into something from a Romero film.
The evidence for this is not that strong for example tales of post Katrina anarchy seem to be overblown. Also natural disasters could reduce the homicide rate as people leave an area after one. Philadelphia is unlikely to undergo a natural disaster though.

There are rare events that still occur often enough to make some sort of prediction on. I do not think any of these is worth including in a model with the possible exception of a riot. But predicting that would need more information about riots and Philadelphia than I have at the moment.

Tuesday, January 12, 2010

Survey the people of Philadelphia

In order to tell if a zip code in Philadelphia is getting more dangerous maybe we should ask people who live there.

So I created a survey here to ask them here. If you are or know someone living in Philadelphia if you could fill that out it would be great.

The idea is to find areas people think are changing in safety and see this turns out to help predict homicides. If it does this could be used to focus police resource in future to help reduce homicides.





I will release any data that is input as I am not the best person to do analysis on them. People going to the effort to submit a survey deserve to get the most out they possible. I will anonymise any data that does contain personal info before releasing it. There shouldn't be any info like this but I will check through the data in case.

I think ideally such a survey would let user place pins in a map in areas they think are improving/disimproving. Any thoughts on the survey? Or the idea of asking people for their local predictions in general?

Monday, January 11, 2010

Analytics X Prize

There is a competition here to try and predict what proportion of murders in Philadelphia will occur in each of the cities 47 zip codes. Many people who are interested in these sorts of puzzles have started submitting predictions.

So how would you go about predicting the murder proportion in each zip code?
Well if nothing changes in Philadelphia you would expect each zip code to keep exactly the same proportion of murders, well with some random variation you could not predict. So my first guess is a repeat of exactly what happened last year.

But in the real world things do change. Say the population changes if every person had the same chance of being murdered then the proportion of murders in a zip code would change proportionate to the change in population. If this was the case the prediction problem would become to find out what changes in population will take place over the year.



The dataset I am using is here and some errors in it need to be removed. Each dot is a zip code. It looks like number of murders does roughly follow population but it is not nearly an exact match. So changes in population are important but they are far from the only thing we need to predict.

How expensive the house in an area are or the average income or the number of people per house might help indicate the murder rate. Here I am looking at number of (murders/population)*10000.





So it looks to me a bit like areas with crowded houses could be more likely to have murders.






House cost looks like it is not connected to murder rate. This could be because zip code is too rough grained for this to be a good judge. Maybe the average cost of a house in a block would be a better measure of risk. Philadelphia has even been broken down into 60ft squares here



Does household income look like it is related to murder rate?

So if the graph is a random scattering of dots then it looks like the independant variable on the x-axis has no relation to homicide rate the dependant variable on the y-axis. If the dots form a line (well not just a line but that is another story) then the homicide rate may be related to that independant variable. It really is not this simple but that's the basics.



As Siah pointed out here young black males seem to be murdered out of proportion. The graph above does seem to suggest that predicting changes in ethnicity of a zip code may improve predictions. Age is another important variable and I do not have data on that so that might be the next thing to get.

There are interesting posts already on this puzzle
"Evaluating Spatial Predictions" and "Second Pass at Analytics X Prize" and "Homicides as non homogeneous poisson processes" are very informative.