The Illusion of Visual Understanding: Mirages in Vision-Language Models

Based on: Asadi et al., "Mirage: The Illusion of Visual Understanding in Vision-Language Models," arXiv:2603.21687v3, April 2026

Duration: ~15 minutes Prerequisites: CLIP, LLaVA, image tokenization, visual reasoning over video (Chain-of-Steps)


Learning Objectives

By the end of this lecture, students will be able to:

  1. Define the mirage effect and distinguish it from hallucination in VLMs

  2. Explain how standard VLM benchmarks can be solved without visual input

  3. Critically evaluate VLM benchmark results using the B-Clean framework

  4. Connect mirage findings to the broader question of whether models genuinely reason over visual input


Lecture Outline

Section 1: Motivation — A Provocative Question (~2 min)

Slide 1: Title slide Slide 2: The core question

Section 2: Defining the Mirage Effect (~2 min)

Slide 3: Mirage vs. hallucination

Section 3: The Phantom-0 Benchmark (~3 min)

Slide 4: Phantom-0 design and Figure 1 heatmap Slide 5: Key results

Section 4: The Super-Guesser Experiment (~3 min)

Slide 6: Mirage-scores on standard benchmarks Slide 7: The super-guesser

Section 5: Why This Happens — Mirage Enablers (~2 min)

Slide 8: Mirage-mode vs. guess-mode

Section 6: B-Clean Mitigation Framework (~1.5 min)

Slide 9: B-Clean pipeline and impact

Section 7: Connecting the Threads (~1.5 min)

Slide 10: Synthesis and discussion


Slide Summary

SlideTitleContent
1Title SlidePaper title, authors, course context
2The Core QuestionCan VLMs answer visual questions without seeing?
3Mirage vs. HallucinationDefinitions, epistemic frame distinction
4Phantom-0 BenchmarkDesign + Figure 1 heatmap
5Phantom-0 ResultsKey mirage rates, pathology bias
6Mirage-ScoresStandard benchmarks without images
7The Super-Guesser3B text-only model beats radiologists
8Why It HappensMirage-mode vs. guess-mode
9B-Clean FrameworkMitigation, impact on scores
10SynthesisConnection to video reasoning, discussion