Large Language Models (LLMs) are increasingly considered for knowledge-intensive and decision-support tasks in organisations. However, their suitability for such use remains constrained by reliability and safety concerns, particularly hallucinated or unsupported claims and incorrect application of rules, instructions, and contextual information.
This project investigates the reliability and safety of open-weight LLMs through systematic empirical evaluation. The research focuses primarily on correctness, factual grounding, hallucination behaviour, and the ability of models to correctly interpret and apply provided rules and instructions. Secondary dimensions include robustness to variations in prompts and context, information disclosure, prompt injection, inappropriate autonomous actions, bias and inconsistent treatment, uncertainty handling, and computational efficiency.
The empirical setting includes a research collaboration with Statens servicecenter (SSC), complemented by more general tasks to enable broader generalisation. Evaluation scenarios will cover document understanding, administrative and case-support tasks, retrieval-augmented generation (RAG), text generation, and agent/tool-use workflows. Experiments will use synthetic material and de-identified or sanitised research material that is suitable for processing on the ordinary Arrhenius environment. Swedish will be the primary evaluation language, with English experiments used for comparison and established benchmarks.
A central part of the project is systematic comparison across open-weight model families and model sizes, including very large models in the 100B+ parameter class. Experiments will investigate inference configurations, quantisation, prompt and system-prompt variation, retrieval-augmented generation, parameter-efficient adaptation such as LoRA/PEFT, and agent configurations.
Reliability will primarily be assessed through expert evaluation by researchers and domain experts, complemented where appropriate by reference answers, automated verification, and validated LLM-based evaluation methods. Repeated experimental runs will be used to investigate variability and robustness rather than relying on single model responses.
Expected outputs include peer-reviewed publications, empirical comparisons of open-weight LLMs, reusable evaluation software and benchmark suites, synthetic and sanitised evaluation datasets, and evidence-based guidance concerning the reliability and safety of LLM-based systems.