NAISS
SUPR
NAISS Projects
SUPR
Efficient Multimodal AI for Autonomous Robotics
Dnr:

NAISS 2026/3-703

Type:

NAISS Medium

Principal Investigator:

Martin Magnusson

Affiliation:

Örebro universitet

Start Date:

2026-08-28

End Date:

2027-09-01

Primary Classification:

10201: Computer Sciences

Secondary Classification:

20208: Computer Vision and learning System (Computer Sciences aspects in 10207)

Tertiary Classification:

20201: Robotics and automation

Allocation

Abstract

The goal of this project is to develop efficient and robust AI approaches for autonomous robotics by investigating multimodal perception, reasoning, and action generation. The project focuses on several complementary research directions, including the development of Vision-Language-Action models for real-world robotic manipulation, multimodal radar-lidar-camera perception for safe navigation in challenging environments, and robotic planning and decision making using Small Language Models. The first direction investigates how pre-trained Vision-Language Models can be extended with action-generation capabilities using robot demonstration data. The goal is to understand how multimodal knowledge from Vision-Language Models can be transferred to robotic action generation and how different model architectures and training strategies affect robotic manipulation performance and generalization to unseen tasks and environments. The second direction focuses on multimodal radar perception for autonomous navigation in harsh environments, such as mining and construction, and under challenging weather and visibility conditions, where dust, smoke, fog, and rain can significantly degrade the performance of RGB and thermal cameras and lidar sensors. A particular focus will be placed on radar-based object detection and motion prediction, with the aim of improving the detection and tracking of dynamic objects under adversarial conditions. The work will explore enhanced radar representations, fusion with complementary sensors, and generative approaches for improving sparse radar point clouds. The third direction examines whether Small Language Models can provide effective reasoning for robotics while requiring substantially fewer computational resources than Large Language Models. The central hypothesis is that, for robotics tasks such as task planning, navigation, task-and-motion planning, and decision making, much of the reasoning problem can be simplified by providing the model with structured information describing the environment, available actions, robot capabilities, constraints, and task objectives. The project will evaluate small open-source language models across representative robotics tasks and compare their performance against larger language models used as baselines. The aim is to determine under which conditions a relatively small model can provide comparable practical performance to a larger model, while requiring substantially fewer computational resources.