2009•Unpublished venueRequires access

Statistical machine learning for internet-scale software repositories

Pierre Baldi, Erik J. Linstead

Open publisher page 0 citations

Abstract

Large repositories of source code available over the Internet create new challenges and opportunities for statistical machine learning. In this dissertation we first develop Sourcerer, an infrastructure for the automated parsing and storage of open source software on an Internet-scale. We gather 4,632 Java projects from Source-Forge and Apache totaling over 38 million lines of code. Simple statistical analyses of the data first reveal robust power-law behavior for package, method call, and lexical containment distributions. We then develop and apply unsupervised, probabilistic topic models to automatically discover the topics embedded in the code and extract topic-word, document-topic, and author-topic distributions. In addition to serving as a convenient summary for program function and developer activities, these and other related distributions provide a statistical and information-theoretic basis for quantifying and analyzing tangling and scattering in the context of Aspect-Oriented Programming. We propose a new theory of aspects that can be summarized as follows: aspects are latent topics with high entropy. This theory is validated for software in the large using the Sourcerer software repository. The theory is also validated for software in the small, with 5 case studies of individual projects. From this study, we show two dozen topics that emerge as general-purpose aspects over the entire data set. We then apply our method to the problem of studying the evolution of software concerns over multiple project versions. We present results for two large, open source Java projects. In addition to detecting the emergence of topics on the release timeline which represent integration points for key source code functionality, our techniques can also be used to pinpoint refactoring events in the underlying software design, as well as to identify general programming concepts whose prevalence is dependent only of the size of the code base to be analyzed. Finally, we turn to the problem of searching Internet-scale software repositories. By combining software textual content with structural information captured by our approach, we are able to significantly improve software retrieval performance, increasing the area-under-curve (AUC) metric significantly compared to previous approaches based on text alone.

About this research paper

What this paper is about

Large repositories of source code available over the Internet create new challenges and opportunities for statistical machine learning. In this dissertation we first develop Sourcerer, an infrastructure for the automated parsing and storage of open source software on an Internet-scale. We gather 4,632 Java projects from Source-Forge and Apache totaling over 38 million lines of code. Simple statistical analyses of the data first reveal robust power-law behavior for package, method call, and lexical containment distributions. We then develop and apply unsupervised, probabilistic topic models to automatically discover the topics embedded in the code and extract topic-word, document-topic, and author-topic distributions. In addition to serving as a convenient summary for program function and developer activities, these and other related distributions provide a statistical and information-theoretic basis for quantifying and analyzing tangling and scattering in the context of Aspect-Oriented Programming. We propose a new theory of aspects that can be summarized as follows: aspects are latent topics with high entropy. This theory is validated for software in the large using the Sourcerer software repository. The theory is also validated for software in the small, with 5 case studies of individual projects. From this study, we show two dozen topics that emerge as general-purpose aspects over the entire data set. We then apply our method to the problem of studying the evolution of software concerns over multiple project versions. We present results for two large, open source Java projects. In addition to detecting the emergence of topics on the release timeline which represent integration points for key source code functionality, our techniques can also be used to pinpoint refactoring events in the underlying software design, as well as to identify general programming concepts whose prevalence is dependent only of the size of the code base to be analyzed. Finally, we turn to the problem of searching Internet-scale software repositories. By combining software textual content with structural information captured by our approach, we are able to significantly improve software retrieval performance, increasing the area-under-curve (AUC) metric significantly compared to previous approaches based on text alone.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Large repositories of source code available over the Internet create new challenges and opportunities for statistical machine learning. In this dissertation we first develop Sourcerer, an infrastructure for the automated parsing and storage of open source software on an Internet-scale. We gather 4,632 Java projects from Source-Forge and Apache totaling over 38 million lines of code. Simple statistical analyses of the data first reveal robust power-law behavior for package, method call, and lexical containment distributions. We then develop and apply unsupervised, probabilistic topic models to automatically discover the topics embedded in the code and extract topic-word, document-topic, and author-topic distributions. In addition to serving as a convenient summary for program function and developer activities, these and other related distributions provide a statistical and information-theoretic basis for quantifying and analyzing tangling and scattering in the context of Aspect-Oriented Programming. We propose a new theory of aspects that can be summarized as follows: aspects are latent topics with high entropy. This theory is validated for software in the large using the Sourcerer software repository. The theory is also validated for software in the small, with 5 case studies of individual projects. From this study, we show two dozen topics that emerge as general-purpose aspects over the entire data set. We then apply our method to the problem of studying the evolution of software concerns over multiple project versions. We present results for two large, open source Java projects. In addition to detecting the emergence of topics on the release timeline which represent integration points for key source code functionality, our techniques can also be used to pinpoint refactoring events in the underlying software design, as well as to identify general programming concepts whose prevalence is dependent only of the size of the code base to be analyzed. Finally, we turn to the problem of searching Internet-scale software repositories. By combining software textual content with structural information captured by our approach, we are able to significantly improve software retrieval performance, increasing the area-under-curve (AUC) metric significantly compared to previous approaches based on text alone.

Key concepts: Computer science, Source code, Java, Software, The Internet, Topic model, Context (archaeology), Data science

Related papers

Back to paper searchBrowse research topicsOriginal source
Statistical machine learning for internet-scale software repositories — Research Paper | ScholarLens