Analytics/Archive/Infrastructure/status

Last update on: 2012-05-monthly

2012-05-22
First 10 Cisco boxes are available. Just puppetized a generic Java installation.

2012-06-06
Hadoop is up and running (under CDH3). http://analytics1001.wikimedia.org:50070/dfshealth.jsp 

2012-06-08
Testing and benchmarking different hadoop parameters. Using TestDSFIO and Terasort benchmarks. Learning!

2012-05-monthly
Cluster planning continues smoothly: Dave Schoonover has begun writing architecture and dataflow documentation for the cluster. At the Berlin Hackathon, Diederik van Liere and Dave Schoonover gathered community input on analytics plans, and gave a few ad hoc presentations about the upcoming changes to the data-processing workflow.

Cluster setup began in earnest mid-month when ops delivered 10 machines from a 2011 Cisco hardware grant. Andrew Otto and Dave Schoonover set up the systems, user environments, and software dependencies to begin testing Cassandra, Hadoop, and Hbase to evaluate which best meets the storage needs of the batch and stream processing systems of the cluster.