This project investigates how gut microbial metabolic enzymes contribute to the production of disease-associated metabolites and the development of cardiometabolic diseases (CMD). By integrating large-scale metagenomic, metabolomic, and clinical datasets from population-based cohorts (e.g., SUPERB, SCAPIS, and IGT), we aim to identify microbial enzymes responsible for key metabolic transformations and provide mechanistic links between microbial genes, metabolites, and disease phenotypes. Rather than relying solely on sequence homology, the project adopts a structure-guided strategy to improve functional annotation of microbial proteins and identify enzymes that may share catalytic functions despite limited sequence similarity.
The analysis begins with cohort-scale processing of shotgun metagenomes, including metagenomic assembly, genome reconstruction, gene catalog generation, and functional annotation to prioritize candidate metabolic enzymes associated with disease-relevant metabolites. These candidates will subsequently undergo large-scale protein structure prediction, protein representation learning, structural similarity searches, active-site and substrate-binding pocket characterization, and protein–metabolite interaction modelling. Structural features will be integrated with metagenomic abundance, metabolomic profiles, and clinical phenotypes to prioritize candidate enzymes and predict their metabolic functions using multimodal machine learning models.
The computational bottleneck of the project lies in large-scale structural modelling of tens of thousands of microbial proteins. Protein structure prediction, generation of protein embeddings using foundation models, binding-pocket analysis, structural comparison, and repeated model training require extensive GPU acceleration and cannot be performed efficiently on conventional CPU-based infrastructure. These analyses involve parallel inference with deep neural networks, high-dimensional structural representations, and repeated optimization of multimodal learning models, resulting in substantial GPU time, memory, and temporary storage requirements. Access to national GPU resources is therefore essential for enabling high-throughput structural annotation and structure-guided discovery of gut microbial metabolic enzymes associated with cardiometabolic disease.