Answering real-world clinical questions using large language model, retrieval-augmented generation, and agentic systems
Special-purpose AI systems far outperformed general-purpose chatbots in answering real-world clinical questions
YS Low, ML Jackson, RJ Hyde, RE Brown, NM Sanghavi, JD Baldwin, CW Pike, J Muralidharan, G Hui, N Alexander, H Hassan, RV Nene, M Pike, CJ Pokrzywa, S Vedak, AP Yan, DH Yao, AR Zipursky, C Dinh, P Ballentine, DC Derieg, V Polony, RN Chawdry, J Davies, BB Hyde, NH Shah, S Gombar
Digital Health, 11: 20552076251348850 (2025)
Plain-language summary
We often struggle to find evidence-based answers for specific patient questions, especially when published research is scarce. In our study, we compared general AI chatbots to specialized systems: one that searches existing medical literature and another that analyzes real-world patient data to generate new insights. We found that the general chatbots rarely provided relevant, evidence-based answers, while our literature-searching tool was effective when studies already existed, and our data-analyzing tool excelled when published evidence was lacking. Our findings suggest that combining these specialized approaches could greatly improve the availability of reliable, actionable information for clinical decision-making.
Abstract
OBJECTIVE: The practice of evidence-based medicine can be challenging when relevant data are lacking or difficult to contextualize for a specific patient. Large language models (LLMs) could potentially address both challenges by summarizing published literature or generating new studies using real-world data. MATERIALS AND METHODS: We submitted 50 clinical questions to five LLM-based systems: OpenEvidence, which uses an LLM for retrieval-augmented generation (RAG); ChatRWD, which uses an LLM as an interface to a data extraction and analysis pipeline; and three general-purpose LLMs (ChatGPT-4, Claude 3 Opus, Gemini 1.5 Pro). Nine independent physicians evaluated the answers for relevance, quality of supporting evidence, and actionability (i.e., sufficient to justify or change clinical practice). RESULTS: General-purpose LLMs rarely produced relevant, evidence-based answers (2-10% of questions). In contrast, RAG-based and agentic LLM systems, respectively, produced relevant, evidence-based answers for 24% (OpenEvidence) to 58% (ChatRWD) of questions. OpenEvidence produced actionable results for 48% of questions with existing evidence, compared to 37% for ChatRWD and <5% for the general-purpose LLMs. ChatRWD provided actionable results for 52% of questions that lacked existing literature compared to <10% for other LLMs. DISCUSSION: Special-purpose LLM systems greatly outperformed general-purpose LLMs in producing answers to clinical questions. Retrieval-augmented generation-based LLM (OpenEvidence) performed well when existing data were available, while only the agentic ChatRWD was able to provide actionable answers when preexisting studies were lacking. CONCLUSION: Synergistic systems combining RAG-based evidence summarization and agentic generation of novel evidence could improve the availability of pertinent evidence for patient care.
Details
Formatted
YS Low, ML Jackson, RJ Hyde, RE Brown, NM Sanghavi, JD Baldwin, CW Pike, J Muralidharan, G Hui, N Alexander, H Hassan, RV Nene, M Pike, CJ Pokrzywa, S Vedak, AP Yan, DH Yao, AR Zipursky, C Dinh, P Ballentine, DC Derieg, V Polony, RN Chawdry, J Davies, BB Hyde, NH Shah, S Gombar. Answering real-world clinical questions using large language model, retrieval-augmented generation, and agentic systems. Digital Health. 2025;11:20552076251348850. doi:10.1177/20552076251348850BibTeX
@article{pike2025answering,
title = {Answering real-world clinical questions using large language model, retrieval-augmented generation, and agentic systems},
author = {Low, YS and Jackson, ML and Hyde, RJ and Brown, RE and Sanghavi, NM and Baldwin, JD and Pike, CW and Muralidharan, J and Hui, G and Alexander, N and Hassan, H and Nene, RV and Pike, M and Pokrzywa, CJ and Vedak, S and Yan, AP and Yao, DH and Zipursky, AR and Dinh, C and Ballentine, P and Derieg, DC and Polony, V and Chawdry, RN and Davies, J and Hyde, BB and Shah, NH and Gombar, S},
journal = {Digital Health},
year = {2025},
volume = {11},
pages = {20552076251348850},
doi = {10.1177/20552076251348850},
}
RIS
TY - JOUR
AU - Low, YS
AU - Jackson, ML
AU - Hyde, RJ
AU - Brown, RE
AU - Sanghavi, NM
AU - Baldwin, JD
AU - Pike, CW
AU - Muralidharan, J
AU - Hui, G
AU - Alexander, N
AU - Hassan, H
AU - Nene, RV
AU - Pike, M
AU - Pokrzywa, CJ
AU - Vedak, S
AU - Yan, AP
AU - Yao, DH
AU - Zipursky, AR
AU - Dinh, C
AU - Ballentine, P
AU - Derieg, DC
AU - Polony, V
AU - Chawdry, RN
AU - Davies, J
AU - Hyde, BB
AU - Shah, NH
AU - Gombar, S
TI - Answering real-world clinical questions using large language model, retrieval-augmented generation, and agentic systems
JO - Digital Health
PY - 2025
VL - 11
SP - 20552076251348850
DO - 10.1177/20552076251348850
ER -
- Publication Type
- Journal Article
- Journal
- Digital Health
- Date Published
- June 9, 2025
- Volume
- 11
- Pages
- 20552076251348850
- Digital Object Identifier
- 10.1177/20552076251348850
- PubMed Identifier
- 40510193
- PubMed Central Identifier
- PMC12159471
