NAISS
SUPR
NAISS Projects
SUPR
Evaluating and Enhancing AI Models and Agentic Systems for Performance, Usability, and Scale
Dnr:

NAISS 2026/3-657

Type:

NAISS Medium

Principal Investigator:

Thor Wikfeldt

Affiliation:

RISE Research Institutes of Sweden

Start Date:

2026-08-31

End Date:

2027-03-01

Primary Classification:

10210: Artificial Intelligence

Secondary Classification:

10214: Networked, Parallel and Distributed Computing

Allocation

Abstract

Agentic AI systems are increasingly central to scientific computing yet running them efficiently on high-performance computing (HPC) infrastructure remains difficult: software stacks are fragmented, performance behaviour is poorly characterised, and few reusable deployment practices exist, especially when pipelines need to be ported to cloud platforms. This project aims to improve the effectiveness, efficiency, and reusability of agentic AI pipelines and frameworks on HPC systems by systematically benchmarking, optimising, and packaging complete software stacks for large-scale inference and open-weights model fine-tuning. The core activity of this project is a systematic evaluation campaign across the full solution space, covering AI inference software (such as, but not limited to, vLLM, SGLang, llama.cpp), open-weights models and model families, and common frameworks for fine-tuning, quantisation, agentic orchestration, retrieval, and tool integration. Combinations of these components will be evaluated for effectiveness and efficiency on representative datasets and complex tasks, using profiling and sampling where necessary to attribute performance behaviour to specific bottlenecks. We will develop and adapt solutions to build comprehensive workflows integrating the best-performing tools and methods, from low-level kernel and runtime optimisations to high-level agentic harnesses, with particular attention to the NVIDIA Grace Hopper architecture. All outputs will be open source from the start and packaged for reuse as Singularity/Apptainer container images, code templates, customisable recipes, and/or self-study guides. Documented procedures will support the migration of optimised workloads between HPC systems and cloud platforms.