Saturday, February 28, 2015

What Colour are Books?

What colour are famous books?

Colours Used I counted up the occurrence of the
colours = ["red","orange","yellow","green","blue","purple","pink","brown","black","gray","white", "grey"]
in Ulysses by James Joyce. I'll post the word count code soon

red 113, orange 12, yellow 50, green 98, blue 82, purple 17, pink 21, brown 59, black 146, gray 2, white 163, grey 68

Turned this count into a barchart with r package ggplot2 graphing package

library(ggplot2)
df <- data.frame(colours = factor(c("pink","red","orange","yellow","green","blue","purple", "brown", "black", "white", "grey"), levels=c("pink","red","orange","yellow","green","blue","purple","brown", "black", "white", "grey")),
                 total_counts = c(21.0, 113.0,12.0, 50.0, 98.0, 82.0, 17.0, 59.0, 146.0,163.0,70.0))
colrs = factor(c("pink","red","orange","yellow","green","blue","purple", "brown", "black", "white", "grey"))

bp <- ggplot(data=df, aes(x=colours, y=total_counts)) + geom_bar(stat="identity",fill=colrs)+guides(fill=FALSE)
bp + theme(axis.title.x = element_blank(), axis.title.y = element_blank())+ ggtitle("Ulysses Color Counts")
bp 

There is a huge element of unweaving the rainbow in just counting the times a colour is mentioned in a book. The program distills “The sea, the snotgreen sea, the scrotumtightening sea.” into a single number. Still I think the ability to quickly look at the colour palette of a book is interesting.

The same picture made from the colours in Anna Karenina by Leo Tolstoy, Translated by Constance Garnett


Translations
Translations produce really funny graphs with this method. According to Jenks@GreekMythComix the ancient Greeks did not really use colours in the same abstract way we did. Things were not 'orange' so much as 'the colour of an orange'. The counts in the Alexander Pope translation of the Iliad are
red 36, yellow 11, green 16, blue 9, purple 43, brown 4, black 69, gray 1, white 25, grey 6

Because colours are not really mentioned in the original Iliad these sorts of graphs could be a quick way to compare translations. Google book trends does not seem to show increased use of these colours overtime.

Sunday, February 22, 2015

2014 Weather Visualizations

There is a great tutorial by Brad Boehmke here on how to build a visualization of temperature in one year compared to a dataset. The infographic is based on one by Tufte

Met Eireann have datasets going back to 1985 on their website here. Some basic data munging on the Met Eireann set for Dublin Airport and I followed the rstats code from the tutorial above to build the graphs below. Wexford would be more interesting for Sun and Kerry for Rain and Wind but those datasets would not download for me.

The first is a comparison of the temperature in 2014 compared to the same date in other years.

Next I looked at average wind speed

And finally the number of hours of sun

These visualizations doesn't look like 2014 was a particularly unusual year for Irish weather. With 30 years of past data if weather was random (which it isn't) at random around 12 days would break the high and low mark for most of these measures. Only the number of sunny days beat this metric. The data met.ie gives contains every day since 1985

maxtp: - Maximum Air Temperature (C)

mintp: - Minimum Air Temperature (C)

rain: - Precipitation Amount (mm)

wdsp: - Mean Wind Speed (knot)

hm: - Highest ten minute mean wind speed (knot)

ddhm: - Mean Wind Direction over 10 minutes at time of highest 10 minute mean (degree)

hg: - Highest Gust (knot)

sun: - Sunshine duration (hours)

dos: - Dept of Snow (cm)

g_rad - Global Radiation (j/cm sq.)

i: - Indicator

Gust might be an interesting one given the storms we had winter 2014. I put big versions of these pictures here, here and here.

Wednesday, February 18, 2015

When were Wodehouse's stories set?

They seem to be sometime before the first world war. But I have never figured out when Wodehouse's books take place. From Something New by P.G. Wodehouse "Whoever carries this job through gets one thousand pounds.” Ashe started. “One thousand pounds–five thousand dollars!” “Five thousand.” Looking at historical exchange rates at www.measuringworth.com the rate stayed close to 1901s $4.87 up until the book was published in 1915. Because exchange rates did not change much at the time they do not help work out when a book was set.

Monday, February 09, 2015

Ancient Death Counts from Poems

What killed you in an ancient battle? Could we look at ancient epics for clues as to what killed people in fights at the time?

Pinker's better Angels of our Nature talks about how archeologists look at bones to see evidence of violent injuries that lead to death. The book talks about examinations of ancient bones unearthed in peat bogs and on long-forgotten battlefields. This bone examination will not tell us about injuries to people that do not cut bones.

The epic poems include the Iliad, Beowulf and the Táin. They were passed down from Bards who memorised them and travelled from place to place reciting them. Some recent research suggests that these epics may have some basis in history. The social network described for the characters usually resembles one real people would have. The social network between characters in Homer’s Odyssey is remarkably similar to real social networks today. That suggests the story is based, at least in part, on real events, say researchers. 'They discovered that while the networks associated with Beowulf and the Iliad had many of the properties of real social networks, the network associated with Tain was less realistic. That led them to conclude that the societies described in the Iliad and Beowulf are probably based on real ones, whereas the Tain appears more artificial.'

There is a site that examines and lists the deaths in the Iliad here. I extracted from there counts for each mentioned body part killed or wounded someone*.

head 21 
jaw 2 
cheek 1 
ear 1 
eye 1 
mouth 1 
nose 1 
skull 1  

neck 12 
throat 3  
  
collar 1 
chest 17 
shoulder 7 
collar bone 2 
nipple 1 
ribs 1 1 of these wound
 
arm 4 3 of these wound
hand 1 1 of these wound
  
back 11 
buttock 2 
  
gut 10 1 of these wound
stomach 5 
liver 3 
 
side 6 
  
thigh 2 1 of these wound
hip 1 
knee 1 
leg 1 
foot 1 1 of these wound
 
groin 2 
testicles 1

I totalled these by body region

Head  29
Neck  15
Upper Body 29
Arm  5
Back  13
Lower Body 18
Side  6
Leg  6
Groin  3 
Using Color brewer to pick out colours I made bins of 5
25-30 RGB 153,0,13
20-25 RGB 203,24,29
15-20 RGB 239,59,44
10-15 RGB 251,106,74
5-10  RGB 252,146,114
1-5   RGB 252,187,161
0     RGB 0,0,0
And I made this into this weird picture. I got the drawing from here. And the idea from Greek myth comix.

Any translation will have disagreements so the original source or as close as we can get to it should be used. Ian Johnson's is the basis for these counts.

Upper body counts for 73 of the deaths: arm, back, legs and lower body count for only 51. But gut, liver and stomach (and maybe buttock) do account for 18 deaths which seems like modern archeology could miss. For example many bog bodies seem to have been ritually killed which may have involved more beheading then the standard violent death.

It would be interesting to do a similar count with the other epic poems to see if liver injury is as common in them or whether that relates to Greek culture.

Anyway please comment what you think about this sort of quantitative analysis of stories that are meant to be entertainment. Can they tell us anything about the ancient world?

*Alcmaon's death I left out as no specific part is named.

Wednesday, February 04, 2015

Irish Alcohol Consumption in 2020

Drink blitz sees bottle of wine rise to €9 minimum 'Irish people still drink an annual 11.6 litres of pure alcohol per capita, 20pc lower than at the turn of the last decade. The aim is to bring down Ireland's consumption of alcohol to the OECD average of 9.1 litres in five years' time.'

What would Irish alcohol consumption be if current trends continue? Knowing this the effectiveness of new measures can be estimated.

The OECD figures are here. I put them in a .csv here.The WHO figures for alcohol consumption are here I loaded the data in R Package

datavar <- read.csv("OECDAlco.csv")

attach(datavar)

plot(Date,Value,

     main="Ireland Alcohol Consumption")
Which looks like this

Looking at that graph alcohol consumption rose from the first year we have data for 1960 until about 2000 and then started dropping. So if the trend since 2000 continued what would alcohol consumption be in 2020?

'Irish people still drink an annual 11.6 litres' I would like to see the source for this figure. We drank 11.6 litres in 2012 according to the OECD. I cannot find OECD figures for 2014. In 2004 we drank 13.6L the claimed 20pc reduction of this is 10.9L, not 11.6L. Whereas the 14.3L we drank in 2002 with a 20pc reduction would now be 11.4. This means it really looks to me like the Independent were measuring alcohol usage up to 2012.

Taking the data since 2000 until 2012.

newdata <- datavar[ which(datavar$Date > 1999), ]

detach(datavar)

attach(newdata)

plot(Date,Value,

     main="Ireland Alcohol Consumption")

cor(Date,Value)

The correlation between year and alcohol consumption since 2000 is [1] -0.9274126. It look like there is a close relationship between the year and the amount of alcohol consumed in that time. Picking 2000, near the peak of alcohol consumption, as the starting date for analysis is arguable. But 2002 was the start of this visible trend in reduced alcohol consumption.

Now I ran a linear regression to predict based on this data alcohol consumption in 2015 and 2020.

> linearModelVar <- lm(Value ~ Date, newdata)
> linearModelVar$coefficients[[2]]*2015+linearModelVar$coefficients[[1]]
[1] 10.42143
> linearModelVar$coefficients[[2]]*2020+linearModelVar$coefficients[[1]]
[1] 9.023077
> 
This means based on data from 2000-2012 we would expect people to drink 10.4 litres this year. Reducing to drinking 9 litres in 2020. So with current trends Irish alcohol consumption will be lower than 'the aim is to bring down Ireland's consumption of alcohol to the OECD average of 9.1 litres in five years'.

There could be something else that is going to alter the trend. One obvious one would be a glut of young adults. People in their 20 drink more than older people. If there are a higher proportion of youths about then the alcohol consumption will rise all else being equal. So will there be a higher proportion of people in their 20s in 5 years time?

The population pyramids projections for Ireland are here. Looking at these there seems to have been a higher proportion of young adults in 2010 than there will be in 2020 which would imply lower alcohol consumption

it would be interesting to see the data and the model that the prediction of Irish alcohol consumption are based on. And to see how minimum alcohol pricing changes the results of these models. But without seeing those models it looks like the Government strategy is promising current trends to continue in response to a new law.

Sunday, November 16, 2014

The number of new drugs is declining

Why Are So Few Blockbuster Drugs Invented Today?
since 1950, the number of new drugs approved has fallen by half roughly every nine years, meaning a total decline by a factor of 80. They called this Eroom’s Law, because it resembled an inversion of Moore’s Law

Graph from In the Pipeline (more here)

Why has the worm not been emulated?


"The way the prophets of the twentieth century went to work was this. They took something or other that was certainly going on in their time and then said that it would go on more and more until something extraordinary happend." G. K. Chesterton

What if you could emulate the brain the way we emulate computer worms?

C elegans is a 1mm long worm with 302 neurons, 3 Nobel prizes and has survived a space shuttle crash. It is one of the simplest animals and has been studied in massive detail. Like the fruitfly or ecoli anything these lab animals do that you cant explain you wont be able to explain in people either.

Whole brain emulation is a prediction that we will be able to simulate the brain in enough detail to create artificial intelligence very like us.

Robin Hanson on econtalk talked about the possible results of whole brain emulation.
"This scenario, which we've called whole brain emulation--taking a whole brain and emulating it on a computer--requires three technologies. One is scanning--you have to be able to scan something in sufficient detail; have to see exactly which parts are where and what they are made of. Two, you have to have models of these cells, a model of the cell input signature and then what comes out of it as a mapping--doesn't have to be exactly right, just has to be close enough. Three, you need a really big computer. A lot of cells, a lot of interactions."
Hanson blogs about brain emulation here. There are interesting fights about whether whole brain emulation is a reasonable prediction or just "the rapture for nerds".

Many biologists seem to think computer people are completely misunderstanding how complicated biological systems are and computer sciencey whole brain Emulation types say that biologists do not understand abstractions because they deal with this complexity all the time.
'[Robin] Hanson’s fundamental mistake is to treat the brain like a human-designed system we could conceivably reverse-engineer rather than a natural system we can only simulate. We may have relatively good models for the operation of nerves, but these models are simplifications, and therefore they will differ in subtle ways from the operation of actual nerves. And these subtle micro-level inaccuracies will snowball into large-scale errors when we try to simulate an entire brain, in precisely the same way that small micro-level imperfections in weather models accumulate to make accurate long-range forecasting inaccurate.' is an example of the biologists argument against brain emulation.

'We should expect brain emulation to be feasible because brains function to process signals, and the decoupling of signal dimensions from other system dimensions is central to achieving the function of a signal processor.

"We can do trend extrapolation and say: Where are we now; if trends continue how long would it take? The computing technology has a nice solid trend; we can project that pretty confidently into the future. The problem is we don't really know how detailed we're going to need to go into these cells. The scanning technology, we have decent trends. This is a vastly smaller industry; small demand. That technology actually looks likely to be ready first. We've actually done a scanning of a whole mouse brain at a decent resolution. A thousandth smaller than a human brain. What does that mean--scanning of a brain? They slice a layer, do a two-dimensional scan of that layer at a fine resolution, go across each cell, and then they slice another layer and do the same thing again. Let me ask again, sort of naive question: If you could take a person's brain out of their head while they were still alive, are you going to be able to get access to my memories in this process? my creativity? All these things we think of as more than a physical process, but of course as you say, it's just chemicals interacting. Is it imaginable that we would be able to reconstruct my memories? To the extent we are confident that your memories and personality are encoded in these cells and where they are and how they talk to each other, so we get that right, we get it all right. That's all you are. Let me say it differently. Looking at it isn't enough. Scanning means noticing the chemical densities. There's thousands of kinds of cells in your brain, and each cell sort of behaves a bit differently. What we need is to know when a cell gets a signal from the outside, electrical or chemical signal, how does that change a cell and what kind of signal does it send out. So, we need to have a model of each of those cell types. We have, actually, models of a wide range of cell types. Doesn't seem that hard to model these cells. We just have a lot of cells to go through and not that much motivation to do it all in a rush. We have actually pretty good models of some particular cells. We have a cell on a dish, we send a signal in, model on the computer, do the same things."

Both sides here. The brain is really complicated squishy stuff and the simcity looks like a real city if you squint sides here could be right. I want a comparison of the predictions of these two theories now and not in 2040 though.

If we had Hanson's 1,2,3 met for an organism and we were not emulating it that would seem to be a problem for the theory.

One is scanning--you have to be able to scan something in sufficient detail;

The c elegans connectome was mapped in 1986
Two, you have to have models of these cells, a model of the cell input signature and then what comes out of it as a mapping--doesn't have to be exactly right, just has to be close enough. There are not many types of neurons in c elegans so we should have a fairly good model of when they will fire.


Three, you need a really big computer. A lot of cells, a lot of interactions."
How big a computer would you need to model all these cells and interactions?

In When will computer hardware match the human brain?
Hans Moravec (1997) gives some nice graphs of how much processing you get for $1000



Kurzweil gives similar figures here

This puts the amount of processing available to a C. elegans at about 1990 levels for $1000. So in 1986 that processing power would have easily been available to university researchers. Maybe that graph is optimistic but if it is out by 25 years for something as simple as c elegans that means predictions of whole brain emulation by 2050 also on the graph will be out as well.

For the last 25 years we have had the power to emulate the whole brain of C. Elegans. Why haven't we?

1. We have not actually because our neuron firing models has not been accurate enough
2. No one cares about emulation fo a worm. A lot of people care about this worm the numbers of neuroscience papers on it confirm this.
3. They are just a bit delayed. There is a thread here on less wrong about this
an open source project openworm*
4. It is hard to get output from a worm 'Our first goal is to combine the neuronal model with this physical model in order to go beyond the biophysical realism that has already been done in previous studies. The physical model will then serve as the "read out" to make sure that the neurons are doing appropriate things.' Pixar, special effects companies and computer game programmers must have fairly good worm emulation programs. If there is a big problem making an animal bodies simulation surely one of them could easily enough make a good model of a tiny bag of gunk?

'Whole Brain Emulation A Roadmap' acknowledges the gap that exists in our emulation of the animal and suggests alternatives ' While the C. elegans nervous system has been completely mapped (White, Southgate et al., 1986), we still lack detailed
electrophysiology, likely because of the difficulty of investigating the small neurons. Animals
with larger neurons may prove less restrictive for functional and scanning investigation but
may lack sizable research communities'

Why 25 years after having a good map and enough computation to run the calculation have we not emulated C Elegans? If it is the modelling of the cells
'I have talked several times to one of the chief scientists who collected the original connectome data and has been continuing to collect more electron micrographs (David Hall, in charge of www.wormatlas.org). He has said that the physiological data on neuron and synapse function in C. elegans is really limited and suggests that no one spend time simulating the worm using the existing datasets because of this. I.e. we may know the connectivity but we don't know even the sign of many synapses.'

The openworm project is really cool. and it might be a good way to get some evidence into the whole brain emulation debate now.

'The problem is we don't really know how detailed we're going to need to go into these cells. '

If Ken Hayworth is right and it is just that 'He has said that the physiological data on neuron and synapse function in C. elegans is really limited' is this because the biologists are right and step 2 the cell models will not be as easy to build as supporters of whole brain emulation claim?
These is a project here to emulate C Elegans. And a good paper here on the problem involved Dynamics of the model of the C Elegans neural network. Just in time to make me look more stupid is A Worm's Mind In A Lego Body. It is not full emulation yet but it is at lest on the path there. * An article from Popular Science on Whole Brain Emulation here that made me resurrect this post I drafted two years ago. Since then the openworm project has moved on massively.

Drones and Ecology

Drones have become wildly popular recently. They seem to be on the path of military, geeks, specific industries ->everything that successful tech seems to go through. It seems likely that large numbers of small deliveries will take place by drone in ten years time. One thing that damaged bird sized tech in the past was hawks. Jon Bentley described in 'More Programming Pearls'
The computers at the two facilities were linked by microwave, but printing the drawings at the test base would have required a printer that was very expensive at the time. The team therefore drew the pictures at the main plant, photographed them, and sent 35mm film to the test station by carrier pigeon, where it was enlarged and printed photographically. The pigeon's 45-minute flight took half the time of the car, and cost only a few dollars per day. During the 16 months of the project the pigeons transmitted several hundred rolls of film, and only two were lost (hawks inhabit the area; no classified data was carried). Because of the low price of modern printers, a current solution to the problem would probably use the microwave link.
Hawks have a habit of attacking small things we send through the skies. There is this great piece on the effect of Amazon drones on ecology. The Dark Extropian Report: The Evolution of Amazon’s “Prime Air” Drone Delivery Service
Hawks and other birds of prey taking issue with these noisy (for now) airborne intruders into their territory. Everyone was worried about people below shooting them down, but it turns out there may be another threat that can’t be so easily policed; outlaw avians....A delivery drone that takes its shape and forms that outline in the sky will not be attacked by a lesser predator, even if it’s not already wired into their genetic memory

The article suggests creating drones that look like very big hawks to discourage natural hawks from attacking them.

This effect doe not just apply to birds of prey though. Prey species hide when they see the outline of a bird of prey. And doing this increases their anxiety enough to reduce feeding and decrease numbers drastically over time. The drones that look like birds of prey will not have to prey. Just being in the sky with the right silhouette will drastically reduce the number of vermin.

According to this article

Changes to ecology have unpredictable effect on the environment. Less pigeons would seem an improvement to the urban environment but they do eat bread and other foodstuffs. If numbers are reduced enough to prevent this bad things could happen.

tldr: 1. there will be lots of drones 2. They will look like birds of prey 3. They will have a big effect on rats, pigeons and other prey species.

Sunday, November 02, 2014

Water Protest Maps

Liam Hogan and Joan Byrne made a great google map here of the water protests in Ireland on November 1st. They kindly let me use their data to create some graphs.

The code and data is here. I will update is as data improves. And maybe to improve the labels on the graphs.

First the number of protestors per county

Number of protestors as a percentage of population of each county

An interesting way of looking at this would be in terms of blocks of 27,640 people. this is the average number of people per TD in Ireland

Dublin has the most variation on the number of protestors. The number in the main O'Connell Street protest was the largest and the one with the most variation in estimated numbers. Other than this most people seem to agree with the numbers Liam and Joan have on their map.

Thursday, October 30, 2014

Apple were never about home computers

According to Wikipedia Steve Jobs' first TV interview was in 1980 in Ireland with Pat Kenny. He talks about the apple being sent up on the space shuttle.

Pat:'Do you actually believe that every home, in the short term at least, will have a personal computer?'

Steve:'We base our theory on the fact that we make personal computers that can be used irrespective of location. The home just happens to be one of the locations apples can be used in'

This shows a fair amount of ambition given it was 4 years before the Macintosh was released.

Friday, August 22, 2014

Making Ebooks More Physical

For this invention will produce forgetfulness in the minds of those who learn to use it, because they will not practice their memory. Their trust in writing, produced by external characters which are no part of themselves
Socrates to Plato in Phaedrus
"A new study which found that readers using a Kindle were "significantly" worse than paperback readers at recalling when events occurred in a mystery story is part of major new Europe-wide research looking at the impact of digitisation on the reading experience." ... "The Kindle readers performed significantly worse on the plot reconstruction measure, ie, when they were asked to place 14 events in the correct order."

The researchers suggest that "the haptic and tactile feedback of a Kindle does not provide the same support for mental reconstruction of a story as a print pocket book does".

"When you read on paper you can sense with your fingers a pile of pages on the left growing, and shrinking on the right," said Mangen. "You have the tactile sense of progress, in addition to the visual ... [The differences for Kindle readers] might have something to do with the fact that the fixity of a text on paper, and this very gradual unfolding of paper as you progress through a story, is some kind of sensory offload, supporting the visual sense of progress when you're reading. Perhaps this somehow aids the reader, providing more fixity and solidity to the reader's sense of unfolding and progress of the text, and hence the story."

From the Guardian If haptic and tactile feedback to give a sense of progress in a book is so important how could this be added to ebook readers? One possibility is a weight that moves from one side of the ebook reader to another as you progress. This would mirror the feeling of a book starting heavier on the right and ending heavier on the left. Assuming you are reading English and not Manga or other back to front based moving system. This could be accomplished with by moving a marble as the book progresses.

Another option is to change the weight based on how far through the book you are. By adding Pez dispenser to the back of the ebook reader that dispensed sweets at points during the book a physical change in the book would result.

Once you are doling out sweets anyway you could flavour them based on the book in question. That would mean you could supply a flavour for the text at certain points which would provide the sort of physical feedback books currently do not. What flavours would go with which books and where? I have never been able to get access to the kindle SDK. But this seems like the sort of hardware project you could try at a hack weekend if you had access to the kindle api.

Thursday, June 26, 2014

The Great Stagnation, Football Ball Edition

'Since 2002, panel numbers have roughly halved every four years: 32 in 2002, 14 in 2006 and eight in 2010. Thus, by the 2022 World Cup, players should be kicking a single-panel ball around the pitch.' claimed Ken Bray here. I will call this Brays law 'The number of panels on the World Cup football ball will half every tournament'. This is not quite as epoch defining as Moore's law but still cool I think.

This year according to projections the soccer ball in the world cup should have at most 4 panels. Instead the ball has 6. 50% more panels then you would expect if progress continued at the rate Bray predicted. The Great Stagnation is the belief that things are not improving as quickly as they used to and is used to explain why we still have homeless people but not flying cars.

The number of panels each world cup ball has is found on each balls individual Wikipedia page.

2014 the Adidas Brazuca: 'The ball has been made of six polyurethane panels'

2010 the Adidas Jabulani: 'The ball was constructed consisting of eight (down from 14 in the 2006 World Cup) thermally bonded, three-dimensional panels'

2006 the Adidas Teamgeist:'The Teamgeist ball differs from previous balls in having just 14 curved panels rather than the 32 that have been standard since 1970. Like the 32 panel Roteiro which preceded it'

Fewer panels mean the ball is smoother and should fly more true. The ball flying true involves the interaction of several variables other than the panel number though. The 2010 ball was notorious for wobbly flight for example. For this reason just reducing the number of panels at the expense of the quality of the ball is a bad idea and may explain why Bray's law has failed. Still not being able to make a ball work with fewer panels indicates a technological innovation slow down to me. The aerodynamics of soccer balls and why there is a race to fewer panels is described in Brays article 'A fly walks round a football'.

Wednesday, March 12, 2014

Lets All Move to No Insurance Land

Ryanair have finally set up their own country. In order not to buy insurance you have to select that option from the country of residence option. Because the UX of select insurance clearly is to choose from a country this option makes complete sense.

If you are going to have a new country of 'don't insure me' it is clearly not in some sort of non country section of the drop down but resides just after Denmark.

I've talked before about where Ryanair if people did what they asked they would disappear. But I still like them, I just like pointing out when some company acts oddly.

There is a level of hiding extra charges from people. Setting up your own country to get an extra few quid out of people really shows commitment.

Sunday, January 05, 2014

Goodreads Recruitment Hack

I entered in some of the node.js books I have been reading into my goodreads list. And they mentioned that they were recruiting.
The programmer who told recruiters about github meet the same fate as the police officer at the end of the wicker man. But at the risk of the same thing happening to me this is a really clever idea. If you are looking for people who know about an area checking if they read the books is one way. Only goodreads can advertise on their site but anyone can look up book reviews. The other surprising thing is with 800k followers goodreads twitter mentions must be like looking at the digital rain from the Matrix. But they noticed that I mentioned their recruitment idea and replied.

Thursday, November 14, 2013

Wheat Map of the US

I thought it would be cool to make a map of the US counties by how much wheat they grew. I took the code from this article and from the Visualize Data book by Nathan Yau

I got some wheat data from here the US department of Agriculture. The map of the US comes from here

Then I cleaned up the data by taking only the columns for state, county and total wheat production. This dataset includes a county 888 and 999 but that seems to be a combination of all the states counties so I stripped those out. Also there are more than 50 states in these county datasets which seems to be standard. There is always messing with numbers being seen as strings with these sorts of manipulations so some casting is needed.

The svg is 1.9 mbs and google drive does not want to store or convert it at the moment but if anyone wants it I can send it to them. This quality of file means zooming in on an individual state, like Kansas, is fine.

The code to create this picture is here.

JDLong on twitter pointed out where to get data for countries. I got the grains from here and a look at the 'head psd_grains_pulses.csv' shows the file layout

I think I want Country_code and value for the commodity wheat in every country in the most recent year value. The country code is 2 characters (iso 3166-1 alpha 2) and the map I have from wikipedia is that format you can get it here

The code to produce colors for each country based on this data is here. Again this is based on the "Visualize This" book from Yau. This css code to set the color of each country gets pasted into the style section of the BlankMap-World6.svg file. I should read all the documentation describing the values before doing any analysis like this. But I am only doing this to make pretty pictures in Python so I am making assumptions to work quickly.

extra: I made a stacked area graph of what crops have been grown when here with the code here.

Sunday, December 30, 2012

World Cup 2010 Heatmap

I am reading Visualize This by Yau at the moment. It is full of really pretty visualization ideas and examples. One it has is creating a heatmap of NBA players. To practice this visualization I have made one of World Cup 2010 players. The dataset I got from the Gardian Data blog 'World Cup 2010 statistics: every match and every player in data'. The data only has 5 qualities quantified but that is good enough to practice making heatmaps.
The R Package code I used is below
library(RColorBrewer)
#save the guardian data to world.csv and load it
players2<-read.csv('World.csv', sep=',', header=TRUE)
players2[1:3,]
#players with the same name (like Torres) meant I had to merge surnames and countries
players2$Name <-paste(players2[,1], players2[,2])
rownames(players2) <- players2$Name
###I removed one player by hand
###I now do not need these columns
players2$Position <- NULL
players2$Player.Surname <- NULL
players2$Team <- NULL
players3 <-players2[order(players2$Total.Passes, decreasing=TRUE),]
### or to order by time played
###players3 <-players2[order(players2$Time.Played, decreasing=TRUE),]
players3 <- players3[,1:5]
players4<-players3[1:50,]
players_matrix <-data.matrix(players4)
###change names of columns to make graph readable
colnames(players_matrix )[1] <- "played"
colnames(players_matrix )[2] <- "shots"
colnames(players_matrix )[3] <- "passes"
colnames(players_matrix )[4] <- "tackles"
colnames(players_matrix )[5] <- "saves"
players_heatmap <- heatmap(players_matrix, Rowv=NA, Colv=NA, col = brewer.pal(9, 'Blues'), scale='column', margins=c(5,10), main="World Cup 2010")
dev.print(file="SoccerPassed.jpeg", device=jpeg, width=600)       
#players_heatmap <- heatmap(players_matrix, Rowv=NA, Colv=NA, col = brewer.pal(9, 'Greens'), scale='column', margins=c(5,10), main="World Cup 2010")
#dev.print(file="SoccerPlayed.jpeg", device=jpeg, width=600) 
dev.off()
Nothing very fancy here. Just showing that with a good data source and some online tutorials it is easy enough to knock up a picture in a fairly short time.

Monday, December 24, 2012

The Price Of Guinness

When money's tight and hard to get 
And your horse has also ran, 
When all you have is a heap of debt - 
A PINT OF PLAIN IS YOUR ONLY MAN.
Myles Na Gopaleen

How much has Guinness increased in price over time? Below is a graph of the price changes. The data is taken from a combination of the Guinness price index and CSO data

The R package code for this graph is below.

pint<-read.csv('pintindex.csv', sep=',', header=TRUE)
plot(pint$Year, pint$Euros, type="s", main="Price Pint of Guinness in Euros", xlab="Year", ylab="Price Euros", bty="n", lwd=2)
dev.print(file="Guinness.jpeg", device=jpeg, width=600)       
dev.off() 
Paul in the comments asked a good question. How does this compare to earnings?
        price   Earnings/Price     Earnings per Week (Euro)
 2008   4.22    167.31                 706.03
 2009   4.34    161.69                 701.73
 2010   4.2     165.02                 693.08
 2011   4.15    165.81                 688.11
 2012   4.23    163.56                 691.87
Here the earnings are average weekly earnings which is the modern and slightly different value to average industrial wage which the Pint Index used. It shows that even with a price drop in Guinness the total purchasing power of pints with wages decreased. This is based on gross wages increases in tax probably made the situation based on net wages worse.

Pintindex.csv is

Year,  Euros 1969, 0.2 1973, 0.24 1976, 0.48 1979, 0.7 1983, 1.37 1984, 1.48 1985, 1.52 1986, 1.64 1987, 1.73 1988, 1.8 1989, 1.87 1990, 1.93 1991, 2.02 1992, 2.15 1993, 2.24 1994, 2.34 1995, 2.42 1996, 2.5 1997, 2.52 1998, 2.65 1999, 2.74 2000, 2.88 2001, 3.01 2002, 3.24 2003, 3.41 2004, 3.54 2005, 3.63 2006, 3.74 2007, 4.03 2008, 4.22 2009, 4.34 2010, 4.2 2011, 4.15 2012, 4.23

Wednesday, December 19, 2012

Cystic Fibrosis Improved Screening

In the first post I claimed that like Tay-Sachs in Israel Cystic Fibrosis could be drastically reduced with some relatively inexpensive genetic testing. In the second further analysis suggested that such genetic screening of the Irish population would pay for itself several times over. In this post I want to see if some form of targeted screening could be shown to be as cost effective as currently implemented screening.

Currently there is free screening for people who has relatives with CF and their partners. I assume they include second cousin as a relative. Based on this paper and some consanguinity calculations I calculate that an Irish couple with one of their second cousins has CF have about twice the chance of having a child with CF as the general population. This means you can be tested for free currently if you have about a 1 in 700 chance of having a child with cystic fibrosis whereas the general population with a 1 in 1444 chance. If a test can be focused the test so that it is twice as good as random screening that should be enough by current standards to be rolled out.

How could a non random screening be made this focused?

1. Geographic area. Some areas of the country might be more likely to have CF carriers than others. Targeting screening in these areas might make it twice as effective. The Cystic Fibrosis Registry of Ireland annual report 2010 gives numbers for Irish counties. 4 counties do not have their numbers listed but I have estimated these based on their population.

This map is based on the figures of people with CF found in the registry. This could be a biased sample or people could have moved. A better measure would be babies born with CF in each county.

Number of people with CF in each county might be useful for deciding how to allocate some treatment resources. What % of people have CF is more interesting for screening though. To work this out we first need the numbers found in each county.

The number of people with CF in the registry per ten thousand people is

I can send anyone who wants them full sized versions of these maps or the r package code I used to generate them. The code I used is below

library(RColorBrewer)
library(sp)
con <- url("http://gadm.org/data/rda/IRL_adm1.RData")
close(con)
people<-read.csv('cases.csv', sep=',', header=TRUE)
pops = cut(people$cases,breaks=c(0,2,10,20,30,40,50,70,150,300))
myPalette<-brewer.pal(9,"Purples")
spplot(gadm, "pops", col.regions=myPalette, main="Cystic Fibrosis Cases Per County",
       lwd=.4, col="black")
dev.print(file="CFIrl.jpeg", device=jpeg, width=600)
dev.off()
population<-read.csv('countypopths.csv', sep=',', header=TRUE)
pops = cut(population$population,breaks=c(0,20,40,60,70,80,100,160,400,1300))

myPalette<-brewer.pal(9,"Greens")
spplot(gadm, "pops", col.regions=myPalette, main="Population in thousands",
       lwd=.4, col="black")
dev.print(file="PopIrl.jpeg", device=jpeg, width=600)       
dev.off()

gadm$cfpop <- people$cases/(population$population/10)
cfpop = cut(gadm$cfpop,breaks=c(0,0.5,1,1.5,2,2.5,3,3.5))
gadm$cfpop <- as.factor(cfpop)

myPalette<-brewer.pal(7,"Blues")
spplot(gadm, "cfpop", col.regions=myPalette, main="CF/Population Irish Counties",
       lwd=.4, col="black")
dev.print(file="CFperPopIrl.jpeg", device=jpeg, width=600)       
dev.off() 
If this result was replicated in a more complete analysis just picking the darker counties could get you the two times amplification needed to have a test as strong as the currently paid for ones.

2. Pick certain ethnic minorities. Some groups have higher levels of CF than the average population. For example travellers have higher levels of some disorders. 'disorders, including Phenylketonuria and Cystic fibrosis, that are found in virtually all Irish communities and probably are no more common among Travellers than in the general Irish population. The second are disorders, including Galactosaemia, Glutaric Acidaemia Type I, Hurler’s Syndrome, Fanconi’s Anaemia and Type II/III Osteogenesis Imperfecta, that are found at much higher frequencies in the Traveller community than the general Irish population'. 'There is no proactive screening of the Traveller population no more than there is proactive screening of the non-traveller Irish population'. I do not think deliberate screening of one ethnic group, unless that group themselves organise it, is a good idea. Singling out one ethnic group for screening risks stigmatising its members and reminds many of the horror of eugenics.

3. Certain disorders seem to cluster with CF. 'In 1936, Guido Fanconi published a paper describing a connection between celiac disease, cystic fibrosis of the pancreas, and bronchiectasis'. Ireland also has the highest rate of celiac disease in the world (about 1 in 100). If CF and celiac disease or some other observable characteristic are also correlated in Ireland testing people with celiac disease in their family could also provide amplification of a test.

4. Screening parents undergoing IVF. HARI was the first clinic in Ireland to offer IVF and it currently receives up to 800 enquiries a year specifically about the procedure. It carries out over 1,350 cycles of IVF treatment annually and over 3,500 babies have been born as a result. The Merrion Clinic carries out up to 500 cycles of IVF per year, while last year, SIMS carried out 1,063 cycles." IVf is roughly 33% effective per cycle so this means about 1000 children are born through IVF from these three Irish clinics here each year. Screening of these parents would prevent roughly one CF case per year. Screening people who use IVF does not prevent many cases. It can be used by people who know they are CF carriers to avoid having a child with CF though.

Concerns about the privacy and security of a general genetic screening program of the Irish population should not be ignored. Cathal Garvey on twitter pointed out that this screening would require 'With explicit informed consent & ensuing destruction of samples, Just wary of prior shenanigans of HSE bloodspot program. i.e. it's already fashionable among governments to abuse screening programs to create 'law' enforcement databases. Without clear guarantees against that, must weigh the costs of mass DNA false incriminations vs. gains of ntnl screening prog!' I agree that any genetic screening program for Ireland would have to ensure privacy for the individual.

Screening the general population for carriers of serious genetic disorders would save money and suffering. If the level of savings are not sufficient for general screening focusing on certain locations or relatives of people who suffer from disorders that co-occur with CF could amplify the returns sufficiently to be as useful as current screenings.

Thursday, December 13, 2012

Gluten Levels of 73 Beers

I often hear it asked what the gluten content of various beers are. Particularly in relation to celiacs who want to avoid gluten. This post is just a direct google translate of a Swedish research paper. PDF's can be hard to search as can Swedish documents for English speaking users. I am just putting up this translation to aid people searching for this research on beer gluten levels. The appendix here is from the Swedish National Food Agency (NFA). This is from a regularly cited report "Gluteninnehåll i de öl som analyserats vid Livsmedelsverket". Gluten content in beer. SLV. 2009 which is difficult to find online. This commonly linked to location linked to but it is dead.
Gluten content of the beers analyzed at the NFA
A total of 73 analyzed beer. For 12 of these low gluten content of 50 mg gluten per liter or higher.
A further 11 beer contained between 41 and 50 mg of gluten per liter. The list is sorted alphabetically by
manufacturers.
One should be aware that the consumption of beer can lead to increased intake of gluten, even if concentrations gluten in beer
is on a par with those found in foods that are appropriate for gluten intolerance.
Consumption of 0.5-1 liters of beer can in some cases make a significant contribution to the daily intake of gluten,
as for an adult celiac disease should be below 50 mg per day gluten.
The table sometimes describes the gluten level as ep = Not detected, which means less than 10 mg per liter gluten
Manufacturer Alcohol Strength Color Names ppm gluten (Mg / l)
AB Åbro Brewery, Sweden 3.5 light Åbro Original ep
AB Åbro Brewery, Sweden 3.5 Light 18:56 ep
AB Åbro Brewery 5.2 light Andersson Beer 47
AB Åbro Brewery 5.2 light Småland 41
Arthur Guinness Son & Co., Dublin, Ireland 3.5 dark Guinness Draft 48
Arthur Guinness Son & Co., Dublin, Ireland 5 dark Guinness Extra Stout 62
Brau Union Österreich AG 2.8 light Zipfer 23
Carlsberg, Denmark 2.8 light Carlsberg Beer 15
Carlsberg, Denmark 3.5 light Carlsberg Beer 21
Carlsberg, Denmark 3.5 Dark Carnegie Porter 20
Carlsberg, Denmark 4.1 light-Saxon gluten ep
Cerveceria Modelo, Mexico 4.6 Light Corona Extra ep
Cerveceria Cuauhtemoc Moctezuma, Mexico 4.5 light Sun ep
Erdinger Weissbräu, Germany 5.3 light Erdinger Weissbier 1188
Erdinger Weissbräu, Germany 5.6 Dark Erdinger Weissbier obscure 1224
Eriksberg 5.6 dark Christmas beer 33
Falcon Breweries, Sweden 2.8 light-Falcon 28
Falcon Breweries, Sweden 3.5 between Falcon Ale 22
Falcon Breweries, Sweden 3.5 light Falcon Extra brew 24
Falcon Breweries, Sweden 3.5 light Falcon Pilz 67
Falcon Breweries, Sweden 5.2 between Bavarian Falcon 55
Falken Falkenberg 3.5 Dark Beer July 49
Grolsche Bierbrowerij, Holland 3.5 light Grolsch Premium Stock 15
Harboes Brewery, Denmark 2.2 Light The Cheerful Dane 25
Harboes Brewery, Denmark 2.8 light Dansk Pilsner premium beer 42
Harboes Brewery, Denmark 3.5 light Dansk Pilsner premium beer 34
Harboes Brewery AB Denmark 3.5 lighting Christmas beer 31
Harboes Brewery AB Denmark 7.3 light Bjørne brewer 49
Hartwall PLC, Tornio, Finland 3.5 light Lapin Kulta ep
Hartwall PLC, Tornio, Finland 5.2 light Lapin Kulta Premium stock ep
Heinecken Brouwerijen Holland Heineken Light 3.5 45
Hofbräu, Germany 6.3 light Hofbräu October-fest bier 26
Inbev UK Limited 3.5 dark Murphys Irish Strout 43
Jämtland Brewery Ltd 6.5 dark Christmas beer e.p.
Kopparberg Brewery 5.3 light Fagerhult Exports III 93
Kra'sne'Březno 4.8 dark Zlatopramen 47
Kronenbourg Strasbourg, France 5.0 light Kronenbourg 1664 97
Krönleins Brewery AB Halmstad 5.3 dark Christmas beer exports 33
Löwenbräu, Germany 6.1 light Lowenbrau October-fest bier 21
Mariestad Brewery Ltd [Spendrups] 2.8 light Mariestads 40
Mariestad Brewery Ltd [Spendrups] 3.5 light Mariestads e.p.
Mariestad Brewery Ltd 3.5 between Julebrygd 60
Pivovar Nova Paka, Czech republic 2.8 light BrouCzech ep
Pivovary Staropramen 3.5 light Staropramen 21
Pripps Sweden 2.2 light Pripps Light beer 17
Pripps Sweden 3.5 Pripps Blue Light Special Stock 32
Pripps Sweden 3.5 light Pripps Blue Pure 28
Pripps (Carlsberg) 5.0 dark Christmas beer 33
Pripps (Carlsberg) 5.2 light Pripps Blue 66
Shepherd Neame Whitstable Kent 3.5 between Bishops Finger ep
Singha Corp. Thailand 5 Light Singha Premium stock beer 17
Source Castle Brewery 3.5 Uppsala dark Christmas beer 23
Source Castle Brewery Ltd 3.5 Light White Weissbier 67
Source Castle Brewery Ltd 3.5 from Vienna ep
Source Castle Brewery Ltd 9.0 dark Imperial Stout 50
Spendrups Brewery Ltd 2.0 dark Gammeldags Moderate Drinking ep
Spendrups Brewery Ltd 2.1 light Spendrups Premium Stock 31
Spendrups Brewery Ltd 2.8 light Norrland Gold 21
Spendrups Brewery Ltd 3.5 light Norrland Gold 35
Spendrups Brewery Ltd 3.5 light Spendrups Premium Gold ep
Spendrups Brewery Ltd 3.5 light Spendrup Bright Brew 28
Spendrups Brewery Ltd 3.5 light Odin Pilsner 46
Spendrups Brewery Ltd 5.0 light Spendrups Premium Stock 53
Spendrups Brewery Ltd 5.2 dark Christmas beer 24
Spendrups Brewery Ltd 5.3 light Mariestads Exports 45
Spendrups Brewery Ltd 5.3 light Norrland Gold 38
Spendrups Brewery Ltd 5.3 dark Norrland July ep
Spendrups Brewery Ltd 5.9 light Spendrups Premium Gold 35
Spendrups Brewery Ltd 7.0 dark julbock 34
Starobrno Brewery Czech 3.5 light Starobrno Premium Stock 21
St Peters Brewery *, UK 4.2 Light St. Petersburg G-free (gluten-free) ep
Tuborg Copenhagen, Denmark 3.5 light Tuborg Beer Premium Gold 28
Zeunerts AB, Sollefteå 5.1 dark Christmas beer 37
* According to the ingredients list on the brew sorghum.
e.p. = Not detected, which means less than 10 mg per liter gluten
My favorite beer blog is by the beer nut and this links to his gluten free section.

Wednesday, December 12, 2012

Cystic Fibrosis Carrier Screening

In my last post Cystic Fibrosis Screening I described how Tay-Sachs had been nearly eradicated in Israel and America and did a rough calculation as to why it would be cost effective to run a similar program to screen for Cystic Fibrosis in Ireland.

In this post I am going to take a closer look at the figures involved to give more evidence that such a screening program is justified.

The Cystic Fibrosis approve of genetic carrier screening for those related to people with CF and their partners. Genetic Carrier Testing For Cystic Fibrosis. 'Carrier testing is limited to adults over the age of 16 where there is a family history of CF, or where a family member has been found to be a carrier of a CF mutation' says the lab that does the testing.

[in the UK] 'A disadvantage of cascade testing is that it will not identify the majority of carrier couples since more than 80% of affected infants are born in families without a prior history of the disease'. Testing relatives though useful only covers a small fraction of potential CF cases.

This screening of relatives is paid for out of public funds

"How much does the test cost? GP fees will apply for arranging the blood test but molecular genetic testing at NCMG and any genetic counselling you may have is a public service and therefore free of charge".

This means that CF Ireland and the health service are involved in and support CF carrier screening. This means some of the moral objections to public screening that might have existed are not present.

What would population wide screening cost? The cost of genetic screening has fallen amazingly fast. for example here is the cost of sequencing an entire genome compared to Moores law.

23andMe a private company has recently announced it will for $99 dollars. This test checks for over 200 genetic markers including some forms of cystic fibrosis. This is further confirmation that genetic screening is getting much cheaper fast and that its current cost is quite low at less than a tenth of the cost of a night in hospital.

The list price of sending a sample from every 16 year old to 23andMe each year would at present be $7.5 million. There are reasons you might not want a private company to do this but it gives a baseline cost. This $7.5 million is the lifetime cost of under 8 CF patients 'However, the lifetime medical cost of the care of a CF child in today’s dollars was estimated to be slightly >$1,000,000'. To be economical, at US prices, this screening would have to prevent 8 of the roughly 40 CF sufferers born a year. Other genetic disorders are also screened for this $99 cost including many of those listed here. None of these are as common as CF in Ireland but these other disorders should be included in a full cost benefit analysis of genetic screening for the Irish population.

This $100 dollar cost is slightly deceptive as once someone finds out they are a CF carrier there are several options available to them. These vary in cost. They can decide (or matchmakers can ensure) not to have children with another carrier. If they do decide to have children with another carrier they can use IVF techniques to ensure an embryo without CF is implanted. Many of the cost analysis of CF screening (like the Rowley et al paper) include the possibility of screening a fetus for CF and terminating the pregnancy if found. This option is not legal in Ireland. They can ignore their and their partners screening results and have a baby as normal with all the risks that entails.

These costs and the probabilities on each have been worked out for the US in the 1998 paper Prenatal screening for cystic fibrosis carriers: an economic evaluation. 'the marginal cost for prenatal CF carrier screening is estimated to be $8,290 per quality-adjusted life-year. This value compares favorably with that of many accepted medical services. The cost of prenatal CF carrier screening could fall to equal the averted costs of CF patient care if the cost of carrier testing were to fall to $100'. This QALY cost figure is used by health care economists to decide which treatments and screenings meet a cost benefit analysis. The wikipedia page on QALY describes the measue well. According to this paper screening in the US, where CF is about four times rarer, is cost effective for the general population at current screening prices.

The paper 'Economic evaluation of cystic fibrosis screening: A review of the literature' has further figures on the cost of screening. This paper is from 2008 and the figures it quotes can be from years before then. As an example of how much screening costs have dropped in that time 23andMe screening cost $999 in 2007 and is now $99 and screens for more genetic markers.

In the UK £30,000 per QALY is generally considered cost effective.

In Ireland what is the cost for a QALY? 'In Ireland, there is no fixed and generally agreed cost effectiveness threshold below which health care technologies would be considered by policy makers to be cost effective'. ’Pee-in-a-pot’ screening in third level institution/college settings may be considered cost effective if a cost effectiveness threshold in the region of €45,000 per QALY gained is used. This €45,000 per QALY gained seems to be a generally accepted figure.

There are more costs to screening than can be supped up in a € per QALY figure. Any screening will induce worry for example. Prostate, breast cancer and other screenings all also induce extra human costs not measurable in QALY though. These common screenings also have to meet these cost per QALY standards.

Given the US analysis at $8,290, screening costs having dropped drastically since then and the high rate of CF gene in the Irish population this implies to me full CF screening of the Irish population would be very cost effective.