EMNLP 2026 Findings

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

Erfan Nourbakhsh1, Ke Yang1, Anthony Rios1

¹University of Texas at San Antonio (UTSA)

The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
Findings Paper

Introduction

Motivation

Prior methods generate free-text answers; MedProb probes frozen layer representations with a linear classifier.

Figure 1. Prior methods (left) give a medical image and question to a VLM or multi-agent system, which then generates a free-text answer. MedProb (right) keeps the VLM frozen, extracts hidden representations from each layer, and predicts the multiple-choice answer with a linear classifier.

Paper

Abstract

Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with MedProb, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text generation. Across PATH-VQA, SLAKE, and VQA-RAD, MedProb recovers substantially more answer-relevant signal than prompting and performs stronger than medical VLMs and agentic systems. Probing also reduces the apparent gap between small and large models compared to prompting, suggesting that smaller VLMs contain more recoverable Med-VQA signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve this linear decodability. Finally, free-text generation exhibits an answer-position bias of up to 10 percentage points, whereas MedProb also has positional bias, however, it is impacted differently than prompting. Our main results target the multiple-choice/multiclass Med-VQA setting; we additionally show the probe can be extended to open-ended generation via a rejection-sampling scoring procedure.

Experiments

Method Overview

We study Med-VQA as multimodal multiple-choice classification and compare what a VLM produces through prompting with what can be read from its internal representations using simple linear probes.

Overview of MedProb: evaluated VLMs and prompting, probing, and fine-tuned prompting settings.

Figure 2. Overview of the MedProb evaluation framework. Left: general-purpose and medical VLMs evaluated, grouped by backbone. Right: prompts with system prompt, medical image, clinical question, and options (A–D), under three settings: direct prompting (top), MedProb linear probing of frozen hidden states (middle), and fine-tuned then prompted VLMs (bottom).

Results: Main Results

Prompting vs. MedProb

Across PATH-VQA, SLAKE, and VQA-RAD, MedProb recovers substantially more answer-relevant signal than prompting and outperforms agentic systems and many medical VLMs. The gap is especially large on PATH-VQA, where prompting often collapses while probing stays high.

Prompting vs MedProb accuracy for general and medical VLMs on VQA-RAD and PATH-VQA.

Figure 3. Prompting vs. MedProb accuracy for general (Gemma-4B/27B) and medical (MedGemma-4B/27B) VLMs on VQA-RAD and PATH-VQA.

Model Method PATH-VQA SLAKE VQA-RAD Average
F1AccF1AccF1AccF1Acc
MAMAgentic51.8551.8039.9259.5139.3159.3139.6256.87
MMedAgentAgentic36.1137.0028.3043.1359.8058.1741.4046.10
UCAgentsAgentic60.5561.1978.0977.1064.1771.1062.3669.80
HuatuoGPT-Vision-34BMedical VLM60.6161.8077.8976.8683.5775.6674.0271.44
Aloe-Vision-72B-ARMedical VLM67.7069.2086.8080.2475.7379.4676.7476.30
UniMedVL-14BMedical VLM52.8448.0066.4658.0766.2458.9461.8555.00
MedVLM-R1-2BMedical VLM†46.7358.0049.2961.2058.7151.7151.5856.97
InfiMed-RL-3BMedical VLM†79.7979.6088.1681.9283.4374.9083.7978.81
MedMo-8BMedical VLM†65.9967.8080.6472.2973.3958.9473.3466.34
Medix-R1-30BMedical VLM†73.0073.0082.4683.6170.3272.2475.2676.28
InternVL3-1BPrompting62.0162.8083.6875.6676.1265.0273.9467.83
InternVL3-1BMedProb82.0982.2088.7083.1381.5172.6284.1079.32
Qwen3-VL-2BPrompting55.7856.8068.7377.8364.7264.6463.0866.42
Qwen3-VL-2BMedProb83.6783.8089.0284.1080.7973.0084.4980.30
Gemma-4BPrompting53.9457.0068.9361.6957.9652.4760.2857.05
Gemma-4BMedProb86.3386.4086.4379.7680.3372.2484.3679.47
Qwen3-VL-4BPrompting59.8854.6086.0377.3546.2467.3064.0566.42
Qwen3-VL-4BMedProb85.6985.8090.6386.0283.1474.9086.4982.24
Qwen3-VL-8BPrompting58.9858.6070.3676.6346.2569.5858.5368.27
Qwen3-VL-8BMedProb86.1286.2089.3584.0982.3475.2885.9381.86
Llama-3.2-11B-VisionPrompting48.9459.8068.4369.8867.6068.0661.6665.91
Llama-3.2-11B-VisionMedProb86.1286.2088.2182.4083.8976.0486.0781.55
Gemma-27BPrompting61.1862.0073.1275.4269.5669.5867.9569.00
Gemma-27BMedProb86.8687.0089.1782.8983.1476.0486.3981.98
Qwen3-VL-30BPrompting63.2361.2088.0480.9674.5377.9575.2773.37
Qwen3-VL-30BMedProb85.7085.8091.1286.7586.4681.3787.7684.64

Table 1. Performance comparison on medical VQA benchmarks across PATH-VQA, SLAKE, and VQA-RAD. Bold: best per column; MedProb rows highlighted; † trained on PATH-VQA, VQA-RAD, and SLAKE.

Results: Effect of Fine-tuning

Does Fine-tuning Close the Gap?

Fine-tuning the base VLM does not close the probing gap. After in-domain or multi-domain fine-tuning, the probe still outperforms prompting on SLAKE and VQA-RAD. Elicitation, not encoded knowledge, remains the bottleneck.

Probing vs prompting accuracy after single-domain and multi-domain fine-tuning.

Figure 4. Probing vs. prompting accuracy after single-domain and multi-domain fine-tuning.

Results: Layer-wise Probing

Where Is the Signal?

Probing accuracy rises with layer depth. Both models start near chance at layer 0; clinically informative signal concentrates in deeper layers.

Layer-wise probing accuracy for Qwen2.5-VL-7B and MedVLThinker-7B.

Figure 5. Layer-wise probing accuracy for Qwen2.5-VL-7B and MedVLThinker-7B.

Results: Answer Option Ordering

Robustness to Option Order

Free-text prompting shows large positional bias (up to ~10 points). MedProb also has positional effects, but is far less sensitive because it reads activations without generating text.

Probing vs prompting under random, correct-first, and correct-last answer orderings.

Figure 6. Probing vs. prompting accuracy under random, correct-first, and correct-last answer orderings.

Results: General vs. Medical

General vs. Medical Models

Across 14 matched pairs, medical adaptation does not consistently improve linear decodability. Six pairs favor the general model; where medical helps, gains are often small (except RL-adapted MediX-R1).

General ModelMedical ModelΔ
Gemma-4BMedGemma-4B−4.92
Gemma-27BMedGemma-27B−0.95
Open-Flamingo-9BMed-Flamingo-9B−1.93
InternVL3-1BBioMed-InternVL3-1B+2.36
LLaVA-7BLLaVA-Med-7B+0.17
LLaMA3.2-11B-VisionBioMed-LLaMA3.2-11B−2.30
Qwen2-VL-2BBioMed-Qwen2-VL-2B+0.05
Qwen2.5-VL-3BMedVLThinker-RL-3B+2.16
Qwen2.5-VL-7BMedVLThinker-7B+1.15
Qwen2.5-VL-32BMedVLThinker-RL-32B+1.65
Qwen3-VL-2BMediX-R1-2B+11.31
Qwen3-VL-4BMedMo-4B−1.65
Qwen3-VL-8BMedMo-8B−0.19
Qwen3-VL-30BMediX-R1-30B+3.72

Table 2. Probing accuracy delta (Δ) between matched general/medical pairs (averaged across benchmarks). Positive: medical better; negative: general better.

Results: Ablation Controls

Does the Probe Use the Image?

Blanking or shuffling images drops accuracy by 6–18 points vs. the original probe. The probe needs both image and question, not text alone.

ModelMethodSLAKEVQA-RAD
Llama-3.2-11B-VisionMedProb82.476.0
Text-only (image-ablated)67.268.4
Image-shuffled66.565.8
Qwen3-VL-8BMedProb84.175.3
Text-only (image-ablated)67.768.4
Image-shuffled65.869.1

Table 3. Best-layer probe accuracy (%) with the original image, an image-ablated (blank) control, and an image-shuffled (mismatched) control.

Results: Sample Efficiency

How Much Labeled Data?

Probe accuracy improves from 50 to 1,000 examples (~12 points). Even at 50 examples, models exceed 71% on PATH-VQA. Gains diminish above 500.

PATH-VQA probe accuracy versus training-set size.

Figure 7. PATH-VQA probe accuracy (%) vs. training-set size (mean ± std over random seeds).

BibTeX

@misc{nourbakhsh2026medprobprobinginternalrepresentations,
  title={MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering},
  author={Erfan Nourbakhsh and Ke Yang and Anthony Rios},
  year={2026},
  eprint={2609.04336},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2609.04336},
}