2015TU/e Research PortalOpen access

Scalable and bandwidth-efficient memory subsystem design for real-time systems

Manil Dev Gomony

Open full text 0 citations

Abstract

Scalable and Bandwidth-Efficient Memory Subsystem Design for Real-Time Systems In heterogeneous multi-processor platforms for real-time systems, Dynamic Random Access Memory (DRAM) is typically used as a shared resource to reduce cost and enable communication between memory clients, i.e. the processing elements. Since multiple applications with firm real-time requirements run concurrently in such platforms, the memory clients impose strict worst-case requirements on main memory performance in terms of bandwidth and/or latency. These requirements must be guaranteed at design time to reduce the verification effort. This is made possible using a real-time memory subsystem consisting of a real-time memory controller and a memory interconnect in front of it that multiplexes requests arriving from different clients. Existing real-time memory controllers bound the execution time of a memory request by fixing the memory access parameters, such as burst size and page policy, at design time. To bound the response time, predictable arbitration policies, such as Time Division Multiplexing (TDM) and Round-Robin (RR), are employed in the memory interconnect. The performance of real-time memory subsystems can be analyzed using formal performance analysis based on e.g., such as network calculus and data-flow analysis. To meet the ever increasing demand for memory bandwidth with more applications being integrated into multi-core platforms, the maximum clock speeds of memory devices were increased by over a factor of two for every memory generation with the help of technology node scaling. Moreover, memory devices with multiple memory channels (multi-channel memories) and wider interfaces, such as Wide IO, were introduced, targeting battery-operated mobile devices. To support the upcoming memory generations in multi-processor platforms with increasing number of clients, scalable memory subsystems are essential. However, existing bus-based memory interconnects with centralized implementation of predictable arbitration policies are not scalable in terms of clock frequency, and current distributed interconnects either suffer from poor performance in terms of area, power

About this research paper

What this paper is about

Scalable and Bandwidth-Efficient Memory Subsystem Design for Real-Time Systems In heterogeneous multi-processor platforms for real-time systems, Dynamic Random Access Memory (DRAM) is typically used as a shared resource to reduce cost and enable communication between memory clients, i.e. the processing elements. Since multiple applications with firm real-time requirements run concurrently in such platforms, the memory clients impose strict worst-case requirements on main memory performance in terms of bandwidth and/or latency. These requirements must be guaranteed at design time to reduce the verification effort. This is made possible using a real-time memory subsystem consisting of a real-time memory controller and a memory interconnect in front of it that multiplexes requests arriving from different clients. Existing real-time memory controllers bound the execution time of a memory request by fixing the memory access parameters, such as burst size and page policy, at design time. To bound the response time, predictable arbitration policies, such as Time Division Multiplexing (TDM) and Round-Robin (RR), are employed in the memory interconnect. The performance of real-time memory subsystems can be analyzed using formal performance analysis based on e.g., such as network calculus and data-flow analysis. To meet the ever increasing demand for memory bandwidth with more applications being integrated into multi-core platforms, the maximum clock speeds of memory devices were increased by over a factor of two for every memory generation with the help of technology node scaling. Moreover, memory devices with multiple memory channels (multi-channel memories) and wider interfaces, such as Wide IO, were introduced, targeting battery-operated mobile devices. To support the upcoming memory generations in multi-processor platforms with increasing number of clients, scalable memory subsystems are essential. However, existing bus-based memory interconnects with centralized implementation of predictable arbitration policies are not scalable in terms of clock frequency, and current distributed interconnects either suffer from poor performance in terms of area, power

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Scalable and Bandwidth-Efficient Memory Subsystem Design for Real-Time Systems In heterogeneous multi-processor platforms for real-time systems, Dynamic Random Access Memory (DRAM) is typically used as a shared resource to reduce cost and enable communication between memory clients, i.e. the processing elements. Since multiple applications with firm real-time requirements run concurrently in such platforms, the memory clients impose strict worst-case requirements on main memory performance in terms of bandwidth and/or latency. These requirements must be guaranteed at design time to reduce the verification effort. This is made possible using a real-time memory subsystem consisting of a real-time memory controller and a memory interconnect in front of it that multiplexes requests arriving from different clients. Existing real-time memory controllers bound the execution time of a memory request by fixing the memory access parameters, such as burst size and page policy, at design time. To bound the response time, predictable arbitration policies, such as Time Division Multiplexing (TDM) and Round-Robin (RR), are employed in the memory interconnect. The performance of real-time memory subsystems can be analyzed using formal performance analysis based on e.g., such as network calculus and data-flow analysis. To meet the ever increasing demand for memory bandwidth with more applications being integrated into multi-core platforms, the maximum clock speeds of memory devices were increased by over a factor of two for every memory generation with the help of technology node scaling. Moreover, memory devices with multiple memory channels (multi-channel memories) and wider interfaces, such as Wide IO, were introduced, targeting battery-operated mobile devices. To support the upcoming memory generations in multi-processor platforms with increasing number of clients, scalable memory subsystems are essential. However, existing bus-based memory interconnects with centralized implementation of predictable arbitration policies are not scalable in terms of clock frequency, and current distributed interconnects either suffer from poor performance in terms of area, power

Key concepts: Computer science, Interleaved memory, Flat memory model, Registered memory, Extended memory, Memory controller, Uniform memory access, Memory refresh

Related papers

Back to paper searchBrowse research topicsOriginal source
Scalable and bandwidth-efficient memory subsystem design for real-time systems — Research Paper | ScholarLens