NAISS
SUPR
NAISS Projects
SUPR
AI for Molecular Engineering
Dnr:

NAISS 2026/3-679

Type:

NAISS Medium

Principal Investigator:

Rocio Mercado

Affiliation:

Chalmers tekniska högskola

Start Date:

2026-09-01

End Date:

2027-09-01

Primary Classification:

10201: Computer Sciences

Secondary Classification:

10203: Bioinformatics (Computational Biology) (Applications at 10610)

Tertiary Classification:

10403: Materials Chemistry

Allocation

Abstract

Note: this is an application to extend the file quota on our existing NAISS Medium Allocation. The compute and storage capacity requests are unchanged; only the number of files is insufficient. The AI Laboratory for Molecular Engineering (AIME), led by Assist. Prof. Rocío Mercado Oropeza in the Section for Data Science and AI, Department of Computer Science and Engineering at Chalmers University of Technology, develops AI-driven methods for molecular engineering at the intersection of machine learning, chemistry, and the life sciences. The group comprises one faculty member, four postdoctoral researchers, eleven PhD students, and approximately eight MSc/BSc thesis students, together with eleven co-advised students and postdocs (of whom 25 are active users on this allocation). Our migration onto Arrhenius was largely completed in June and July 2026, and our Large Storage resources have been consolidated into this allocation as planned. Our research objectives are to: (1) train deep generative and language models for molecular design and optimization, including synthesizability-constrained generation, retrosynthesis, and the design of multi-target therapeutic modalities such as PROTACs and molecular glues; (2) develop large-scale multi-modal and representation-learning models for single-cell data and cell-image (phenotypic) analysis; (3) apply atomistic and coarse-grained molecular dynamics and ab initio methods both to understand biomolecular interactions (e.g., ternary-complex formation, membrane mechanics) and to generate training data for surrogate property models; and (4) discover sustainable materials, including PFAS alternatives and battery electrolytes, by coupling simulation-derived datasets with generative and predictive models. These efforts have produced a substantial body of peer-reviewed and open-access work (https://ailab.bio/publications) and tools, with further publications in preparation. Consistent with the group's open-science commitment, all code, models, and datasets are released open-source (GitHub, Hugging Face, Zenodo) and acknowledge NAISS resources. We ask only for an increase in the file quota on Arrhenius Disk, from 3,000,000 to 10,000,000 files, at the Medium ceiling. The capacity we were otherwise granted is correct and we do not request a change to it. We are currently using 9,036 of 35,000 GiB, about a quarter of the allocation, but 3.9M files against a 3M quota, so we are already over the file limit at a quarter of capacity. This reflects the nature of our data rather than undisciplined use: our mean file size across live working data is about 1 MiB, whereas a 3M quota on 35,000 GiB assumes about 12 MiB, and file counts in our imaging and literature-extraction work scale with the number of images and documents rather than with data volume. The Resource Usage section gives the per-user, per-project audit behind the 10M figure and details the compute request, which is unchanged. We have a data management plan in place.