2002•Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIERequires access

Fault-tolerant extensible approach to improving spectral analysis throughput via parallel/distributed processing

Peter G. Raeth, George R. Mikulski, Lorie I. Moffitt, Wayne C. Wood

Open publisher page 0 citations

Abstract

This paper describes a means of achieving fault-tolerance and architecture extensibility for parallel/distributed systems that support spectral analysis. These attributes are essential to critical 24/7/366 operations and they improve upon systems that only enhance throughput. They also address the single-point-of-failure issues attendant upon architectures that commit critical operations to single machines. Graceful throughput degradation is achieved to mitigate all-or-nothing approaches. Parallel/distributed processing has three important goals. The first is for the subject application to provide faster throughput than it would while running on a single CPU or computer. The second goal is to make best use of existing capital equipment. For critical systems, the third goal is fault tolerance via redundancy. This project addresses the third goal. It seeks to demonstrate a means to make parallel/distributed processing systems fault tolerant so that crashes of individual machines ina cluster do not bring the entire system down. In spite of individual machine failures, it also seeks to ensure the completion of all tasks so that system throughput degrades gracefully. These goals can be met by a system composed of a generic TCP/IP LAN connecting some number of ordinary office computers and laboratory workstations that are heterogeneous and of unknown reliability. Described here is concept formulation and design. Other projects in this arena are referenced. These provide essential technology to this present effort. Particular application is made to detecting unspecified anomalies in unspecified data streams drawn from staring continuous-dwell sensors. This application enables the reliable non-stop detection of unexpected events, with the results immediately made available to human analysts or additional automated processing.

About this research paper

What this paper is about

This paper describes a means of achieving fault-tolerance and architecture extensibility for parallel/distributed systems that support spectral analysis. These attributes are essential to critical 24/7/366 operations and they improve upon systems that only enhance throughput. They also address the single-point-of-failure issues attendant upon architectures that commit critical operations to single machines. Graceful throughput degradation is achieved to mitigate all-or-nothing approaches. Parallel/distributed processing has three important goals. The first is for the subject application to provide faster throughput than it would while running on a single CPU or computer. The second goal is to make best use of existing capital equipment. For critical systems, the third goal is fault tolerance via redundancy. This project addresses the third goal. It seeks to demonstrate a means to make parallel/distributed processing systems fault tolerant so that crashes of individual machines ina cluster do not bring the entire system down. In spite of individual machine failures, it also seeks to ensure the completion of all tasks so that system throughput degrades gracefully. These goals can be met by a system composed of a generic TCP/IP LAN connecting some number of ordinary office computers and laboratory workstations that are heterogeneous and of unknown reliability. Described here is concept formulation and design. Other projects in this arena are referenced. These provide essential technology to this present effort. Particular application is made to detecting unspecified anomalies in unspecified data streams drawn from staring continuous-dwell sensors. This application enables the reliable non-stop detection of unexpected events, with the results immediately made available to human analysts or additional automated processing.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

This paper describes a means of achieving fault-tolerance and architecture extensibility for parallel/distributed systems that support spectral analysis. These attributes are essential to critical 24/7/366 operations and they improve upon systems that only enhance throughput. They also address the single-point-of-failure issues attendant upon architectures that commit critical operations to single machines. Graceful throughput degradation is achieved to mitigate all-or-nothing approaches. Parallel/distributed processing has three important goals. The first is for the subject application to provide faster throughput than it would while running on a single CPU or computer. The second goal is to make best use of existing capital equipment. For critical systems, the third goal is fault tolerance via redundancy. This project addresses the third goal. It seeks to demonstrate a means to make parallel/distributed processing systems fault tolerant so that crashes of individual machines ina cluster do not bring the entire system down. In spite of individual machine failures, it also seeks to ensure the completion of all tasks so that system throughput degrades gracefully. These goals can be met by a system composed of a generic TCP/IP LAN connecting some number of ordinary office computers and laboratory workstations that are heterogeneous and of unknown reliability. Described here is concept formulation and design. Other projects in this arena are referenced. These provide essential technology to this present effort. Particular application is made to detecting unspecified anomalies in unspecified data streams drawn from staring continuous-dwell sensors. This application enables the reliable non-stop detection of unexpected events, with the results immediately made available to human analysts or additional automated processing.

Key concepts: Computer science, Fault tolerance, Throughput, Redundancy (engineering), Distributed computing, Troubleshooting, Extensibility, Workstation

Related papers

Back to paper searchBrowse research topicsOriginal source
Fault-tolerant extensible approach to improving spectral analysis throughput via parallel/distributed processing — Research Paper | ScholarLens