Publication
Efficient Irregular Wavefront Propagation Algorithms on Hybrid CPU-GPU Machines
Downloadable Content
- Persistent URL
- Last modified
- 02/20/2025
- Type of Material
- Authors
- Language
- English
- Date
- 2013-04-01
- Publisher
- Elsevier
- Publication Version
- Copyright Statement
- © 2013 Elsevier B.V. Published by Elsevier B.V. All rights reserved.
- License
- Final Published Version (URL)
- Title of Journal or Parent Work
- ISSN
- 0167-8191
- Volume
- 39
- Issue
- 4-5
- Start Page
- 189
- End Page
- 211
- Grant/Funding Information
- This research was funded, in part, by grants from the National Institutes of Health through contract HHSN261200800001E by the National Cancer Institute; and contracts 5R01LM009239-04 and 1R01LM011119-01 from the National Library of Medicine, R24HL085343 from the National Heart Lung and Blood Institute, NIH NIBIB BISTI P20EB000591, RC4MD005964 from National Institutes of Health, and PHS Grant UL1TR000454 from the Clinical and Translational Science Award Program, National Institutes of Health, National Center for Advancing Translational Sciences.
- This research used resources of the Keeneland Computing Facility at the Georgia Institute of Technology, which is supported by the National Science Foundation under Contract OCI-0910735.
- Abstract
- We address the problem of efficient execution of a computation pattern, referred to here as the irregular wavefront propagation pattern (IWPP), on hybrid systems with multiple CPUs and GPUs. The IWPP is common in several image processing operations. In the IWPP, data elements in the wavefront propagate waves to their neighboring elements on a grid if a propagation condition is satisfied. Elements receiving the propagated waves become part of the wavefront. This pattern results in irregular data accesses and computations. We develop and evaluate strategies for efficient computation and propagation of wavefronts using a multi-level queue structure. This queue structure improves the utilization of fast memories in a GPU and reduces synchronization overheads. We also develop a tile-based parallelization strategy to support execution on multiple CPUs and GPUs. We evaluate our approaches on a state-of-the-art GPU accelerated machine (equipped with 3 GPUs and 2 multicore CPUs) using the IWPP implementations of two widely used image processing operations: morphological reconstruction and euclidean distance transform. Our results show significant performance improvements on GPUs. The use of multiple CPUs and GPUs cooperatively attains speedups of 50× and 85× with respect to single core CPU executions for morphological reconstruction and euclidean distance transform, respectively.
- Author Notes
- Keywords
- Research Categories
- Biology, Bioinformatics
Tools
- Download Item
- Contact Us
-
Citation Management Tools
Relations
- In Collection:
Items
| Thumbnail | Title | File Description | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|---|
|
|
Publication File - v1f77.pdf | Primary Content | 2025-02-06 | Public | Download |