Microarchitecture support for reducing branch penalty in a superscaler processor
Maiko Sakamoto, Yasuhiro Nunomura, Takeshi Yoshida, Yukihiko Shimazu
Abstract
Maiko Sakamoto, Yasuhiro Nunomura, Takeshi Yoshida, Yukihiko Shimazu
Abstract
This paper describes the microarchitecture of the 32-bit superscalar microprocessor GMICRO/400 with simple prejump mechanisms and its performance evaluation. GMICR0/400 has six stages of instruction execution pipeline and implements a dynamic branch prediction scheme, executing jump instructions in early stages. For dynamic branch predictions, GMICR0/400 contains a 1-Kbit table which holds a single history bit for each conditional branch instruction. For dynamic return-address predictions, it contains a 16-entry stack, which holds copies of return addresses of return-from-subroutine instructions. Prediction accuracies of the branch history table and the return-address stack are 81.7% and 99.6% for the SPECCINT92 respectively, and achieve a speed-up of 1.27. This performance is 95% of that of an ideal model, with a much more complex prejump mechanism and perfect accuracies.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This paper describes the microarchitecture of the 32-bit superscalar microprocessor GMICRO/400 with simple prejump mechanisms and its performance evaluation. GMICR0/400 has six stages of instruction execution pipeline and implements a dynamic branch prediction scheme, executing jump instructions in early stages. For dynamic branch predictions, GMICR0/400 contains a 1-Kbit table which holds a single history bit for each conditional branch instruction. For dynamic return-address predictions, it contains a 16-entry stack, which holds copies of return addresses of return-from-subroutine instructions. Prediction accuracies of the branch history table and the return-address stack are 81.7% and 99.6% for the SPECCINT92 respectively, and achieve a speed-up of 1.27. This performance is 95% of that of an ideal model, with a much more complex prejump mechanism and perfect accuracies.
Key concepts: Branch predictor, Microarchitecture, Superscalar, Computer science, Pipeline (software), Parallel computing, Speculative execution, Microprocessor