A Method to Improve the Throughput of the Instruction Fetch Unit in SMT VLIW Processors
Shuming Chen
Abstract
Shuming Chen
Abstract
In a simultaneous multithreaded processor, improving the throughput of the instruction fetch unit usually means that there is more drastic cache competition between threads, but this competition limits the throughput reversely. Based on the characteristics of the current VLIW architectures,this paper presents an instruction fetch scheme that improves the throughput of the fetch unit and the whole processor.By canceling the invalid addresses in the instruction fetching pipeline, it decreases those conflicts of program caches caused by invalid instruction fetch.As the experimental results show,this scheme can improve the throughput of the instruction unit and the performance of the whole processor by 12%~23% relatively,while the program cache’s miss rate increases appreciably, even decreases sometimes. It also reduces the program cache’s accesses by 10%~25%, so the power consumption of the whole processor is decreased.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In a simultaneous multithreaded processor, improving the throughput of the instruction fetch unit usually means that there is more drastic cache competition between threads, but this competition limits the throughput reversely. Based on the characteristics of the current VLIW architectures,this paper presents an instruction fetch scheme that improves the throughput of the fetch unit and the whole processor.By canceling the invalid addresses in the instruction fetching pipeline, it decreases those conflicts of program caches caused by invalid instruction fetch.As the experimental results show,this scheme can improve the throughput of the instruction unit and the performance of the whole processor by 12%~23% relatively,while the program cache’s miss rate increases appreciably, even decreases sometimes. It also reduces the program cache’s accesses by 10%~25%, so the power consumption of the whole processor is decreased.
Key concepts: Fetch, Very long instruction word, Computer science, Throughput, Cache, Parallel computing, Instructions per cycle, Pipeline burst cache