BioNLP 2026 · ACL 2026

When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

Erfan Nourbakhsh1, Rocky Slavin1, Ke Yang1, Anthony Rios1

¹University of Texas at San Antonio (UTSA)

BioNLP 2026
colocated with ACL 2026 · San Diego, California

Introduction

Motivation

RAG is widely assumed to help medical QA. Across models from 7B to 72B, we find only small gains, pointing to evidence use, not retrieval quality alone, as the bottleneck.

Overview of motivation: retrieval yields only small gains across model scales.

Figure 1. Overview of our motivation and main finding: across models from 7B to 72B, retrieval yields only small gains, suggesting that the main bottleneck is evidence use rather than retrieval quality.

Paper

Abstract

Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely viewed as a promising solution, and prior work has reported substantial gains for large medical QA models. We revisit this assumption across a broad range of open-weight instruction-tuned models spanning 7B to 72B parameters. Across five models, ten biomedical QA datasets, four retrieval methods, and four retrieval corpora, we find that retrieval yields only small and inconsistent improvements over a no-retrieval baseline, typically within 1-2 points. In contrast, the choice of backbone model has a much larger effect than the choice of retriever or corpus, and expert and layman retrieval sources perform similarly in most settings. These results suggest that the main bottleneck is not retrieval quality alone, but the model's limited ability to use retrieved evidence effectively. Code is available at github.com/erfan-nourbakhsh/BioMedicalRAG.

Experiments

Experimental Pipeline

Five open-weight models (7B-72B), ten expert/layman QA datasets, four retrievers (BM25, TF-IDF, MedCPT, Hybrid RRF), and four corpora (BioASQ/PubMed, Medical Textbooks, Yahoo Answers, HealthCareMagic).

Experimental pipeline overview for biomedical RAG evaluation.

Figure 2. Experimental pipeline overview.

Results

Retrieval Helps Little

On close-ended medical QA, retrieval often hurts smaller models and gives at most ~1-2 points for larger ones. Backbone choice dominates retriever and corpus choice.

Model Data MCQA MQA MMLU Avg
LLaMA3.1-8Bw/o RAG80.883.883.782.8
BioASQ75.984.682.380.9
HCM73.784.274.077.3
Textbook74.184.183.380.5
Yahoo74.683.781.579.9
LLaMA3.1-70Bw/o RAG81.089.189.286.4
BioASQ79.890.890.286.9
HCM78.980.587.582.3
Textbook80.281.189.583.6
Yahoo79.191.788.786.5
Mistral-7Bw/o RAG72.677.976.775.7
BioASQ61.372.472.268.6
HCM63.373.272.069.5
Textbook66.073.777.372.3
Yahoo65.574.074.271.2
Qwen2.5-7Bw/o RAG80.083.786.383.3
BioASQ75.881.182.279.7
HCM74.981.482.879.7
Textbook76.681.185.781.1
Yahoo77.581.485.281.4
Qwen2.5-72Bw/o RAG82.581.992.585.6
BioASQ80.181.691.284.3
HCM77.385.890.784.6
Textbook80.183.591.284.9
Yahoo79.881.890.584.0

Table 1. Accuracy by model and retrieval corpus. MCQA denotes MedMCQA, MQA denotes MedQA, and HCM denotes HealthCareMagic. Bold: best per model column block.

Ablations

Shots and Top-k

Few-shot count and retrieval depth (top-k) do not overturn the main finding: gains stay small across close- and open-ended settings.

Close-ended accuracy across shot counts 1, 3, 5, 10.

Figure 3. Close-ended accuracy across shot counts (1, 3, 5, 10).

Open-ended ROUGE-L across shot counts 1, 3, 5, 10.

Figure 4. Open-ended ROUGE-L across shot counts (1, 3, 5, 10).

Close-ended accuracy across top-k values.

Figure 5. Close-ended accuracy across top-k (1, 3, 5, 10, 25, 50).

Open-ended ROUGE-L across top-k values.

Figure 6. Open-ended ROUGE-L across top-k values.

Conclusion

Main Takeaways

BibTeX

@inproceedings{nourbakhsh-etal-2026-retrieval,
    title = "When Retrieval Doesn{'}t Help: A Large-Scale Study of Biomedical {RAG}",
    author = "Nourbakhsh, Erfan  and
      Slavin, Rocky  and
      Yang, Ke  and
      Rios, Anthony",
    booktitle = "{B}io{NLP} 2026",
    month = jul,
    year = "2026",
    address = "San Diego, California",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.bionlp-1.72/",
    doi = "10.18653/v1/2026.bionlp-1.72",
    pages = "890--910",
}