2017•UvA-DARE (University of Amsterdam)Open access

Reverse engineering source code: Empirical studies of limitations and opportunities

Davy Landman

Open full text 3 citations

Abstract

The goal of software renovation is to modernize software. One way to achieve this is to first reverse engineer the essential concepts and abstractions used in the software from source code and then use these during renovation. Scaling reverse engineering to large software systems requires automated analysis. Automation often comes at the cost of over-approximation or under-approximation. This thesis explores the limits of and opportunities for these approximations via three research questions. First, we have explored the limits of domain model recovery by manually recovering domain models. Comparing these models to a manually constructed reference domain model we found that most domain information could be recovered -- with high quality -- from the source code. Second, we have explored using both Cyclomatic Complexity (CC) and Source Lines of Code (SLOC) for automating reverse engineering. Almost all of the literature claims a strong linear correlation between these two metrics. This is often interpreted as indication that CC and SLOC are redundant to each other. In two large corpora we did not observe a strong correlation. We interpret this as a lack of evidence for CC being redundant to SLOC. Finally, we have explored the limits of statically analyzing Java’s Reflection API. Analyzing a representative corpus revealed that 78% of all projects use Reflection. After identifying the common assumptions and limitations of relevant static analysis tools we found them widely challenged in the corpus. We propose new opportunities for static analysis tools.

Open-access reader

About this research paper

What this paper is about

The goal of software renovation is to modernize software. One way to achieve this is to first reverse engineer the essential concepts and abstractions used in the software from source code and then use these during renovation. Scaling reverse engineering to large software systems requires automated analysis. Automation often comes at the cost of over-approximation or under-approximation. This thesis explores the limits of and opportunities for these approximations via three research questions. First, we have explored the limits of domain model recovery by manually recovering domain models. Comparing these models to a manually constructed reference domain model we found that most domain information could be recovered -- with high quality -- from the source code. Second, we have explored using both Cyclomatic Complexity (CC) and Source Lines of Code (SLOC) for automating reverse engineering. Almost all of the literature claims a strong linear correlation between these two metrics. This is often interpreted as indication that CC and SLOC are redundant to each other. In two large corpora we did not observe a strong correlation. We interpret this as a lack of evidence for CC being redundant to SLOC. Finally, we have explored the limits of statically analyzing Java’s Reflection API. Analyzing a representative corpus revealed that 78% of all projects use Reflection. After identifying the common assumptions and limitations of relevant static analysis tools we found them widely challenged in the corpus. We propose new opportunities for static analysis tools.

Why it matters

OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The goal of software renovation is to modernize software. One way to achieve this is to first reverse engineer the essential concepts and abstractions used in the software from source code and then use these during renovation. Scaling reverse engineering to large software systems requires automated analysis. Automation often comes at the cost of over-approximation or under-approximation. This thesis explores the limits of and opportunities for these approximations via three research questions. First, we have explored the limits of domain model recovery by manually recovering domain models. Comparing these models to a manually constructed reference domain model we found that most domain information could be recovered -- with high quality -- from the source code. Second, we have explored using both Cyclomatic Complexity (CC) and Source Lines of Code (SLOC) for automating reverse engineering. Almost all of the literature claims a strong linear correlation between these two metrics. This is often interpreted as indication that CC and SLOC are redundant to each other. In two large corpora we did not observe a strong correlation. We interpret this as a lack of evidence for CC being redundant to SLOC. Finally, we have explored the limits of statically analyzing Java’s Reflection API. Analyzing a representative corpus revealed that 78% of all projects use Reflection. After identifying the common assumptions and limitations of relevant static analysis tools we found them widely challenged in the corpus. We propose new opportunities for static analysis tools.

Key concepts: Reverse engineering, Computer science, Source code, Source lines of code, Cyclomatic complexity, Static program analysis, Domain engineering, Software engineering

Related papers

Back to paper searchBrowse research topicsOriginal source
Reverse engineering source code: Empirical studies of limitations and opportunities — Research Paper | ScholarLens