Progressive Retrieval and Hierarchical Visualization of Large Remote Data
Hans‐Christian Hege, Andrei Hutanu, Ralf Kähler, André Merzky, Thomas Radke, Edward Seidel, Brygg Ullmer
Abstract
Hans‐Christian Hege, Andrei Hutanu, Ralf Kähler, André Merzky, Thomas Radke, Edward Seidel, Brygg Ullmer
Abstract
The size of data sets produ\ned on remote super\nomputer fa\nilities frequently ex\needs the pro\nessing \napabilities of lo\nal visualization workstations. This phenomenon in\nreasingly limits s\nientists when analyzing results of large-s\nale s\nienti \n simulations. That problem gets even more prominent in s\nienti \n \n ollaborations, spanning large virtual organizations, working on \nommon shared sets of data distributed in Grid environments. In the visualization \nommunity, this problem is addressed by distributing the visualization pipeline. In parti\nular, early stages of the pipeline are exe\nuted on resour\nes \nloser to the initial (remote) lo\nations of the data sets. This paper presents an e \n ient te\nhnique for pla\ning the rst two stages of the visualization pipeline (data a\n\ness and data lter) onto remote resour\nes. This is realized by exploiting the extended retrieve feature of GridFTP for exible, high performan\ne a ess to very large HDF5 les. We redu\ne the number of network transa\ntions for ltering operations by utilizing a server side data pro\nessing plugin, and hen\ne redu\ne laten\ny overhead \nompared to GridFTP partial le a\n\ness. The paper further des\nribes the appli\nation of hierar\nhi\nal rendering te\nhniques on remote uniform data sets, whi\nh make use of the remote data ltering stage. 1. Introdu\ntion. The amount of data produ\ned by numeri\nal simulations on super\nomputing fa\nilities ontinues to in\nrease rapidly in parallel with the in\nreasing \nompute power, main memory, storage spa\ne, and I/O transfer rates available to resear\nhers. These developments in super\nomputing have been observed to ex\need the growth of \nommodity network bandwith and visualization workstation memory/performan\ne by a fa\ntor of 4 [11℄. Hen\ne, it is in\nreasingly \nriti\nal to use remote data a\n\ness te\nhniques for analyzing this data. Among
OpenAlex reports 17 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The size of data sets produ\ned on remote super\nomputer fa\nilities frequently ex\needs the pro\nessing \napabilities of lo\nal visualization workstations. This phenomenon in\nreasingly limits s\nientists when analyzing results of large-s\nale s\nienti \n simulations. That problem gets even more prominent in s\nienti \n \n ollaborations, spanning large virtual organizations, working on \nommon shared sets of data distributed in Grid environments. In the visualization \nommunity, this problem is addressed by distributing the visualization pipeline. In parti\nular, early stages of the pipeline are exe\nuted on resour\nes \nloser to the initial (remote) lo\nations of the data sets. This paper presents an e \n ient te\nhnique for pla\ning the rst two stages of the visualization pipeline (data a\n\ness and data lter) onto remote resour\nes. This is realized by exploiting the extended retrieve feature of GridFTP for exible, high performan\ne a ess to very large HDF5 les. We redu\ne the number of network transa\ntions for ltering operations by utilizing a server side data pro\nessing plugin, and hen\ne redu\ne laten\ny overhead \nompared to GridFTP partial le a\n\ness. The paper further des\nribes the appli\nation of hierar\nhi\nal rendering te\nhniques on remote uniform data sets, whi\nh make use of the remote data ltering stage. 1. Introdu\ntion. The amount of data produ\ned by numeri\nal simulations on super\nomputing fa\nilities ontinues to in\nrease rapidly in parallel with the in\nreasing \nompute power, main memory, storage spa\ne, and I/O transfer rates available to resear\nhers. These developments in super\nomputing have been observed to ex\need the growth of \nommodity network bandwith and visualization workstation memory/performan\ne by a fa\ntor of 4 [11℄. Hen\ne, it is in\nreasingly \nriti\nal to use remote data a\n\ness te\nhniques for analyzing this data. Among
Key concepts: Computer science, Visualization, Pipeline (software), Grid, Rendering (computer graphics), Data access, Workstation, Data visualization