Viruses, known and unknown have a significant impact on human health, through disease, microbiome modulation and shaping the immune system. We study viruses through metagenomics, where large datasets are generated from human samples, followed by identification of viruses through computational methods. The sample sets include type 1 diabetes, pediatric cancers and skin diseases, as well as infections of unknown origin. We have published several novel viruses, and descriptions of viral microbiomes with putative disease associations. We have several ongoing projects, including following up data showing an association between virus infections and the development of psoriasis lesions, indications that circoviruses may be involved in autoimmune diseases and screening of clinical samples where ni clear diagnosis has been possible. Metagenomics is also used to study ubiquitous chronic human virus infections, primarily anelloviruses. These are highly variable DNA viruses that nearly all humans carry in the blood stream and that interact with the immune system. The association of anellovirus variation with human disease is unclear due to the complexity of the virus population.
The bioinformatics analyses have been carried out using an evolving in-house pipeline that combines different nucleotide and protein sequence-based homology and motif-based methods into lists of identities of known viruses and candidates for novel viruses. This pipeline differs from others in that it incorporates methods to find novel virus variants and species to a great extent, and does not only focus on known viruses for diagnostics purposes. A separate pipeline has been developed for the analysis of anelloviruses, based on identification of anellovirus sequences, stringent assemblies, and classification based on phylogenetic analysis at the protein level. This has made it possible to begin to perform statistical analyses of anellovirus types and species and their possible association with human disease.
In both the overall virus microbiome analysis and the anellovirus studies, the improvements in machine learning capacity and methods can be used to further improve and deepen the analyses. Our main research questions in this proposal are: 1. To incorporate machine learning and protein structure prediction methods applied to unclassified sequences to identify potential unknown viruses. 2. To improve the analysis of anelloviruses by incorporating information on protein motifs and protein structure, including immune epitopes using machine learning methods.
In the former case, we see the potential for improved identification of candidate novel viruses that can be verified in the laboratory. Each new virus species or family can, as we have seen previously, turn out to be of importance for human health and open new avenues of research. In case of anelloviruses, these methods will open ip possibilities to identify species and protein variants that interact with the host immune system and may be involved in causing disease.