Showing posts with label hadoop. Show all posts
Showing posts with label hadoop. Show all posts

Friday, September 30, 2011

Java, Oracle and Hadoop

Was reading the summary of a /. article ( http://developers.slashdot.org/story/11/09/30/1855204/oracle-proud-self-reliant-increasingly-isolated )and it got be wondering again - what is the fate of Hadoop and the JDK.

Kind of wonder what it would take for a new implementation of hadoop ( or similar implementation ) but in a different language. What if Oracle were to basically reach a point where there was the public domain Java that goes basically stagnant and a commercial version for-pay suddenly exploding the cost of your cluster beyond servers and some switches.

Sure there are some other frameworks ( MongoDB, Cassandra, probably many others ) but in the same idea as the marriage of HDFS and the map/reduce framework itself.

Would OpenJDK even be an option? I've not tried using it for anything. Then again, I just administer hadoop clusters I don't program for them ;)

Wednesday, July 20, 2011

speculative execution

Running some benchmarks of hadoop using teragen/terasort. One of the recommendations I was given was to disable speculative execution. Noticed something rather strange when I forced it to disabled in the config.

Runtime with speculative execution: 18.5 minutes
Runtime without speculative execution: 1 hour

Seems that 2-3 map tasks are taking longer than the rest.

Question now is: why. Each map task is responsible for generating the same % of data - why would speculative execution make the job run quicker. Does this point to hardware differences ( if so, the slow tasks are on different machines - I have not noticed a pattern yet ), configuration problems elsewhere, or just random bad luck.