Vision-language models (VLMs) have achieved strong performance in multimodal benchmarks, but they often rely on learned language priors rather than the actual image. The ViLP (Probing Visual Language Priors) benchmark specifically exposes this failure case by including strong distractor facts in the questions, in order to trick the model to answer based on its linguistic knowledge, without consulting the image. So, despite near-perfect human performance, modern VLMs show low performance. This project aims to develop novel methods for visually grounded reasoning in VLMs.