%0 Conference Paper %B 2010 Asia Pacific Software Engineering Conference (APSEC2010) %D 2010 %T Qualitas Corpus: A Curated Collection of Java Code for Empirical Studies %A Tempero, Ewan %A Anslow, Craig %A Dietrich, Jens %A Han, Ted %A Li, Jing %A Lumpe, Markus %A Melton, Hayden %A Noble, James %K curated code corpus %K Empirical studies %K experimental infrastructure %U http://qualitascorpus.com/docs/citation.html %X In order to increase our ability to use measurement to support software development practise we need to do more analysis of code. However, empirical studies of code are expensive and their results are difficult to compare. We describe the Qualitas Corpus, a large curated collection of open source Java systems. The corpus reduces the cost of performing large empirical studies of code and supports comparison of measurements of the same artifacts. We discuss its design, organisation, and issues associated with its development.