Cooperative prefetching: compiler and hardware support for effective instruction prefetching in modern processors
Chi-Keung Luk, Todd C. Mowry
Abstract
Chi-Keung Luk, Todd C. Mowry
Abstract
Instruction cache miss latency is becoming an increasingly importantperformance bottleneck, especially for commercial applications. Although instruction prefetching is an attractive technique for tolerating this latency, we find that existing prefetching schemes are insufficient for modern superscalar processors since they fail to issue prefetches early enough (particularly for non-sequential accesses). To overcome these limitations, we propose a new instruction prefetching technique whereby the hardware and software cooperate to hide the latency as follows. The hardware performs aggressive sequential prefetching combined with a novel prefetch filtering mechanism to allow it to get far ahead without polluting the cache. To hide the latency of non-sequential accesses, we propose and implement a novel compiler algorithm which automatically inserts instructionprefetch instructions into the executable to prefetch the targets of control transfers far enough in advance. Our experimental res...
OpenAlex reports 61 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Instruction cache miss latency is becoming an increasingly importantperformance bottleneck, especially for commercial applications. Although instruction prefetching is an attractive technique for tolerating this latency, we find that existing prefetching schemes are insufficient for modern superscalar processors since they fail to issue prefetches early enough (particularly for non-sequential accesses). To overcome these limitations, we propose a new instruction prefetching technique whereby the hardware and software cooperate to hide the latency as follows. The hardware performs aggressive sequential prefetching combined with a novel prefetch filtering mechanism to allow it to get far ahead without polluting the cache. To hide the latency of non-sequential accesses, we propose and implement a novel compiler algorithm which automatically inserts instructionprefetch instructions into the executable to prefetch the targets of control transfers far enough in advance. Our experimental res...
Key concepts: Instruction prefetch, Computer science, Parallel computing, Compiler, Cache, Latency (audio), Bottleneck, CAS latency