Automatic Vectorization by Runtime Binary Translation
Takashi Nakamura, Satoshi Miki, Shuichi Oikawa
Abstract
Takashi Nakamura, Satoshi Miki, Shuichi Oikawa
Abstract
Today, most general desktop PCs uses processors that incorporate SIMD (Single Instruction Multiple Data) units. These SIMD units, however, are for the most part underutilized with the exception of a few multimedia applications. As a result, the processing power of modern processors is not fully unleashed. In this research, we propose a system that applies automatic vectorization techniques to binary code at runtime to increase the utilization ratio of SIMD units, which will speed up the execution of programs. We will refer to this system as Selftrans. Selftrans extracts parallelism from binary x86 machine code without requiring its source code, and translates it into a binary code that utilizes the SIMD units. This paper describes the design and implementation of Selftrans. We will show that Selftrans is (1) capable of automatic vectorization of binary machine code, (2) that SIMDization can significantly improve performance, and (3) that it is possible to perform architecture-specific optimization for various processors.
OpenAlex reports 8 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Today, most general desktop PCs uses processors that incorporate SIMD (Single Instruction Multiple Data) units. These SIMD units, however, are for the most part underutilized with the exception of a few multimedia applications. As a result, the processing power of modern processors is not fully unleashed. In this research, we propose a system that applies automatic vectorization techniques to binary code at runtime to increase the utilization ratio of SIMD units, which will speed up the execution of programs. We will refer to this system as Selftrans. Selftrans extracts parallelism from binary x86 machine code without requiring its source code, and translates it into a binary code that utilizes the SIMD units. This paper describes the design and implementation of Selftrans. We will show that Selftrans is (1) capable of automatic vectorization of binary machine code, (2) that SIMDization can significantly improve performance, and (3) that it is possible to perform architecture-specific optimization for various processors.
Key concepts: SIMD, Binary translation, Computer science, Vectorization (mathematics), Parallel computing, x86, Code (set theory), Binary number