Comprehensive assessment of machine learning methods for diagnosing gastrointestinal diseases through whole metagenome sequencing data.

Jul 08, 2024

Experts: Sungho Lee,Insuk Lee

The gut microbiome, linked significantly to host diseases, offers potential for disease diagnosis through machine learning (ML) pipelines. These pipelines, crucial in modeling diseases using high-dimensional microbiome data, involve selecting profile modalities, data preprocessing techniques, and classification algorithms, each impacting the model accuracy and generalizability. Despite whole metagenome shotgun sequencing (WMS) gaining popularity for human gut microbiome profiling, a consensus on the optimal methods for ML pipelines in disease diagnosis using WMS data remains elusive. Addressing this gap, we comprehensively evaluated ML methods for diagnosing Crohn’s disease and colorectal cancer, using 2,553 fecal WMS samples from 21 case-control studies. Our study uncovered crucial insights: gut-specific, species-level taxonomic features proved to be the most effective for profiling; batch correction was not consistently beneficial for model performance; compositional data transformations markedly improved the models; and while nonlinear ensemble classification algorithms typically offered superior performance, linear models with proper regularization were found to be more effective for diseases that are linearly separable based on microbiome data. An optimal ML pipeline, integrating the most effective methods, was validated for generalizability using holdout data. This research offers practical guidelines for constructing reliable disease diagnostic ML models with fecal WMS data.

Author

admin

View all posts

ABOUT THE EXPERTS

Sungho Lee,Insuk Lee

Sungho Lee

Department of Biotechnology, College of Life Science and Biotechnology, Yonsei University, Seoul, Republic of Korea.

Insuk Lee

Department of Biotechnology, College of Life Science and Biotechnology, Yonsei University, Seoul, Republic of Korea.

POSTECH Biotech Center, Pohang University of Science and Technology (POSTECH), Pohang, Republic of Korea.