Introduction
RAG is widely assumed to help medical QA. Across models from 7B to 72B, we find only small gains, pointing to evidence use, not retrieval quality alone, as the bottleneck.
Figure 1. Overview of our motivation and main finding: across models from 7B to 72B, retrieval yields only small gains, suggesting that the main bottleneck is evidence use rather than retrieval quality.
Paper
Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely viewed as a promising solution, and prior work has reported substantial gains for large medical QA models. We revisit this assumption across a broad range of open-weight instruction-tuned models spanning 7B to 72B parameters. Across five models, ten biomedical QA datasets, four retrieval methods, and four retrieval corpora, we find that retrieval yields only small and inconsistent improvements over a no-retrieval baseline, typically within 1-2 points. In contrast, the choice of backbone model has a much larger effect than the choice of retriever or corpus, and expert and layman retrieval sources perform similarly in most settings. These results suggest that the main bottleneck is not retrieval quality alone, but the model's limited ability to use retrieved evidence effectively. Code is available at github.com/erfan-nourbakhsh/BioMedicalRAG.
Experiments
Five open-weight models (7B-72B), ten expert/layman QA datasets, four retrievers (BM25, TF-IDF, MedCPT, Hybrid RRF), and four corpora (BioASQ/PubMed, Medical Textbooks, Yahoo Answers, HealthCareMagic).
Figure 2. Experimental pipeline overview.
Results
On close-ended medical QA, retrieval often hurts smaller models and gives at most ~1-2 points for larger ones. Backbone choice dominates retriever and corpus choice.
| Model | Data | MCQA | MQA | MMLU | Avg |
|---|---|---|---|---|---|
| LLaMA3.1-8B | w/o RAG | 80.8 | 83.8 | 83.7 | 82.8 |
| BioASQ | 75.9 | 84.6 | 82.3 | 80.9 | |
| HCM | 73.7 | 84.2 | 74.0 | 77.3 | |
| Textbook | 74.1 | 84.1 | 83.3 | 80.5 | |
| Yahoo | 74.6 | 83.7 | 81.5 | 79.9 | |
| LLaMA3.1-70B | w/o RAG | 81.0 | 89.1 | 89.2 | 86.4 |
| BioASQ | 79.8 | 90.8 | 90.2 | 86.9 | |
| HCM | 78.9 | 80.5 | 87.5 | 82.3 | |
| Textbook | 80.2 | 81.1 | 89.5 | 83.6 | |
| Yahoo | 79.1 | 91.7 | 88.7 | 86.5 | |
| Mistral-7B | w/o RAG | 72.6 | 77.9 | 76.7 | 75.7 |
| BioASQ | 61.3 | 72.4 | 72.2 | 68.6 | |
| HCM | 63.3 | 73.2 | 72.0 | 69.5 | |
| Textbook | 66.0 | 73.7 | 77.3 | 72.3 | |
| Yahoo | 65.5 | 74.0 | 74.2 | 71.2 | |
| Qwen2.5-7B | w/o RAG | 80.0 | 83.7 | 86.3 | 83.3 |
| BioASQ | 75.8 | 81.1 | 82.2 | 79.7 | |
| HCM | 74.9 | 81.4 | 82.8 | 79.7 | |
| Textbook | 76.6 | 81.1 | 85.7 | 81.1 | |
| Yahoo | 77.5 | 81.4 | 85.2 | 81.4 | |
| Qwen2.5-72B | w/o RAG | 82.5 | 81.9 | 92.5 | 85.6 |
| BioASQ | 80.1 | 81.6 | 91.2 | 84.3 | |
| HCM | 77.3 | 85.8 | 90.7 | 84.6 | |
| Textbook | 80.1 | 83.5 | 91.2 | 84.9 | |
| Yahoo | 79.8 | 81.8 | 90.5 | 84.0 |
Table 1. Accuracy by model and retrieval corpus. MCQA denotes MedMCQA, MQA denotes MedQA, and HCM denotes HealthCareMagic. Bold: best per model column block.
Ablations
Few-shot count and retrieval depth (top-k) do not overturn the main finding: gains stay small across close- and open-ended settings.
Figure 3. Close-ended accuracy across shot counts (1, 3, 5, 10).
Figure 4. Open-ended ROUGE-L across shot counts (1, 3, 5, 10).
Figure 5. Close-ended accuracy across top-k (1, 3, 5, 10, 25, 50).
Figure 6. Open-ended ROUGE-L across top-k values.
Conclusion
@inproceedings{nourbakhsh-etal-2026-retrieval,
title = "When Retrieval Doesn{'}t Help: A Large-Scale Study of Biomedical {RAG}",
author = "Nourbakhsh, Erfan and
Slavin, Rocky and
Yang, Ke and
Rios, Anthony",
booktitle = "{B}io{NLP} 2026",
month = jul,
year = "2026",
address = "San Diego, California",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.bionlp-1.72/",
doi = "10.18653/v1/2026.bionlp-1.72",
pages = "890--910",
}