Publication

From genome-scale data to models of infectious disease: A Bayesian network-based strategy to drive model development

Downloadable Content

Persistent URL
Last modified
  • 02/20/2025
Type of Material
Authors
    Weiwei Yin, Georgia Institute of TechnologyJessica C. Kissinger, University of GeorgiaC. Moreno, Emory UniversityMary Galinski, Emory UniversityMark P. Styczynski, Georgia Institute of Technology
Language
  • English
Date
  • 2015-12-01
Publisher
  • Elsevier
Publication Version
Copyright Statement
  • © 2015 Published by Elsevier B.V.
License
Final Published Version (URL)
Title of Journal or Parent Work
ISSN
  • 0025-5564
Volume
  • 270
Issue
  • Pt B
Start Page
  • 156
End Page
  • 168
Grant/Funding Information
  • This project has been funded in whole or in part with federal funds from the National Institute of Allergy and Infectious Diseases; National Institutes of Health, Department of Health and Human Services [contract no. HHSN272201200031C].
Supplemental Material (URL)
Abstract
  • High-throughput, genome-scale data present a unique opportunity to link host to pathogen on a molecular level. Forging such connections will help drive the development of mathematical models to better understand and predict both pathogen behavior and the epidemiology of infectious diseases, including malaria. However, the datasets that can aid in identifying these links and models are vast and not amenable to simple, reductionist, and univariate analyses. These datasets require data mining in order to identify the truly important measurements that best describe clinical and molecular observations. Moreover, these datasets typically have relatively few samples due to experimental limitations (particularly for human studies or in vivo animal experiments), making data mining extremely difficult. Here, after first providing a brief overview of common strategies for data reduction and identification of relationships between variables for inclusion in mathematical models, we present a new generalized strategy for performing these data reduction and relationship inference tasks. Our approach emphasizes the importance of robustness when using data to drive model development, particularly when using genome-scale, small-sample in vivo data. We identify the use of appropriate feature reduction combined with data permutations and subsampling strategies as being critical to enable increasingly robust results from network inference using high-dimensional, low-observation data.
Author Notes
Keywords
Research Categories
  • Engineering, Biomedical
  • Biology, Bioinformatics

Tools

Relations

In Collection:

Items