Tandem mass spectrometry (MS/MS) is a key technology for the identification and characterization of small molecules in metabolomics, environmental chemistry, and the life sciences. This project aims to fine-tune the DreaMS foundation model (https://www.nature.com/articles/s41587-025-02663-3) on large-scale MS/MS datasets to improve structural annotation and molecular characterization from tandem mass spectra.
The work will focus on adapting pre-trained models to domain-specific datasets, evaluating performance on annotation tasks, and developing reproducible machine learning workflows for training and inference. The project will make use of GPU resources for model fine-tuning, hyperparameter optimization, and large-scale prediction experiments.
The expected outcomes include improved annotation performance and computational workflows that can support a range of mass spectrometry applications. The project is conducted within the Chalmers Mass Spectrometry Infrastructure (CMSI) and contributes to the development and use of machine learning methods for computational mass spectrometry.