NAISS
SUPR
NAISS Projects
SUPR
A natural-language interface for the Allen Mouse Brain Connectivity Atlas with embedded uncertainty quantification: toward a minimum transparency standard for LLM-assisted neuroscience
Dnr:

NAISS 2026/4-1295

Type:

NAISS Small

Principal Investigator:

Roksana Khalid

Affiliation:

Stockholms universitet

Start Date:

2026-08-13

End Date:

2027-09-01

Primary Classification:

30105: Neurosciences

Webpage:

Allocation

Abstract

The brain's connectivity is high-dimensional and relational, and the questions researchers ask of it are inherently multi-part. Natural-language querying with large language models could give researchers direct access to curated neuroscience data without the technical barrier of programmatic querying, and locally deployable models could extend this to researchers' own sensitive data without exposing it to hosted systems. But these models hallucinate, and those best placed to benefit are least equipped to detect when an output is wrong, so the central problem is not access but knowing when to trust the output. We address this by pairing natural-language querying with validated measures of output reliability. We will validate our workflow against the Allen Mouse Brain Connectivity Atlas, a resource of 2,331 brain-wide anterograde tracer experiments passing Allen's quality control, whose reference data let us measure whether the certainty signal actually tracks correctness. Each experiment maps axonal projections from a defined injection site, registered to a common anatomical framework, and together they compose a region-level map of projection strength between brain areas, with cell-type specificity where Cre driver lines are used. This provides systematically validated reference data against which model outputs can be checked, isolating model reliability as the object of study. Our aim is to provide a validated natural-language workflow which lets researchers ask the database questions directly while preserving the reliability that unvalidated AI access would sacrifice. As part of this project, we aim to assess the accuracy and confidence of open-source LLMs for answering domain-specific neuroscience questions. We use both semantic embedding and LLM-as-a-judge frameworks to estimate model confidence and agreement with ground truth answers. However, evaluating model performance and uncertainty over a large range of prompts is both time-consuming and computationally-expensive on a single GPU. As such, it is impossible to get a timely estimate of model performance without using HPC resources. Therefore, we request access to Arrhenius to assess multiple LLMs in parallel across separate GPUs.