Design and evaluation of a compiler algorithm for prefetching
Todd C. Mowry, Monica S. Lam, Anoop Gupta
Abstract
Open-access reader
Todd C. Mowry, Monica S. Lam, Anoop Gupta
Abstract
Open-access reader
Software-controlled data prefetching is a promising technique for improving the performance of the memory subsystem to match today's high-performance processors.While prefctching is useful in hiding the latency, issuing prefetches incurs an instruction overhead and can increase the load on the memory subsystem.As a resu 1~ care must be taken to ensure that such overheads do not exceed the benefits.This paper proposes a compiler algorithm to insert prefetch instructions into code that operates on dense matrices.Our algorithm identiEes those references that are likely to be cache misses, and issues prefetches only for them.We have implemented our algorithm in the SUfF (Stanford University Intermediate Form) optimizing compiler.By generating fully functional code, we have been able to measure not only the improvements in cache miss rates, but also the oversdl performance of a simulated system.We show that our algorithm significantly improves the execution speed of our benchmark programs-some of the programs improve by as much as a factor of two.When compared to an algorithm that indiscriminately prefetches alf array accesses, our algorithm can eliminate many of the unnecessary prefetches without any significant decrease in the coverage of the cache misses.
OpenAlex reports 768 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Software-controlled data prefetching is a promising technique for improving the performance of the memory subsystem to match today's high-performance processors.While prefctching is useful in hiding the latency, issuing prefetches incurs an instruction overhead and can increase the load on the memory subsystem.As a resu 1~ care must be taken to ensure that such overheads do not exceed the benefits.This paper proposes a compiler algorithm to insert prefetch instructions into code that operates on dense matrices.Our algorithm identiEes those references that are likely to be cache misses, and issues prefetches only for them.We have implemented our algorithm in the SUfF (Stanford University Intermediate Form) optimizing compiler.By generating fully functional code, we have been able to measure not only the improvements in cache miss rates, but also the oversdl performance of a simulated system.We show that our algorithm significantly improves the execution speed of our benchmark programs-some of the programs improve by as much as a factor of two.When compared to an algorithm that indiscriminately prefetches alf array accesses, our algorithm can eliminate many of the unnecessary prefetches without any significant decrease in the coverage of the cache misses.
Key concepts: Computer science, Compiler, Citation, Programming language, World Wide Web