OpenEvidence as a Tool for Physicians
In the field of medicine, the volume of trials, systematic reviews, and society guidelines far exceeds what any clinician can hold in working memory, and much of it changes from year to year. This abundance of information is especially acute in perioperative care, where recommendations may rest on a shifting and sometimes thin evidence base.1 The explosion of artificial intelligence (AI) tools in recent years additionally complicates navigating the wealth of available information, as these tools provide quick and easy answers to questions while drawing on different sources and offering varying levels of accuracy. OpenEvidence is a large language model that seeks to act as an accurate knowledge base of medical information for physicians and a useful tool to support clinical decision-making. It is designed for physicians to ask clinical questions in plain language and be presented with a direct answer with inline citations to the primary sources.2 The model is trained on peer-reviewed clinical research papers.
The practical appeal of OpenEvidence is in its potential to condense the search for and synthesis of clinical evidence. Point-of-care decisions require clinicians to analyze the specific case at hand under time pressure, drawing on their experience, practice guidelines, and a large and changing body of literature. A tool that compresses that process could make evidence more accessible. For example, OpenEvidence lets a physician ask a specific, patient-framed question and get a referenced response in seconds. Instead of searching "aspiration risk GLP-1," a clinician can ask, "How long should semaglutide be held before elective surgery to reduce aspiration risk?" and receive an answer that seeks to incorporate the most current statements and studies. That specificity reduces the cognitive load of translating a broad query into a usable recommendation.
The research is promising but still in early stages. Existing studies assess answer quality, physician ratings, and performance on clinical cases; they do not yet show that OpenEvidence improves patient outcomes. The evidence currently supports using it to augment, rather than replace, physician judgment.3,4,5
A 2025 study examined five retrospective primary-care cases involving hypertension, hyperlipidemia, type 2 diabetes, depression, and obesity. Four physicians rated the answers for clarity, relevance, evidentiary support, effect on decision-making, and satisfaction. OpenEvidence scored highly for clarity, relevance, support, and satisfaction, and its recommendations aligned with the physicians’ existing plans. The tool had little effect on the decisions themselves, scoring lower for impact on clinical decision-making; it mainly reinforced plans rather than changing them.3
A larger blinded evaluation used 620 real-world questions submitted by physicians across 30 specialties. One hundred forty-nine practicing physicians compared OpenEvidence with three general-purpose AI tools. They rated OpenEvidence highest for accuracy, clinical utility, source quality, verifiability, and completeness.5 These findings suggest that clinical specialization, targeted retrieval, and reference presentation can improve usefulness compared with general-purpose chatbots. They do not prove that OpenEvidence is correct or safest in every specialty or workflow.
The most cautionary evidence comes from a pilot study of 100 complex medical subspecialty scenarios. OpenEvidence’s rapid mode achieved 34% accuracy, while Deep Consult achieved 41%; evaluator concordance was 77% and 72%, respectively.4 These results indicate that performance may decline when questions require complex specialty reasoning or integration of clinical nuance.
A separate comparison of four clinical chatbots, including OpenEvidence, assessed 128 question-and-answer pairs across orthopedics, pediatrics, gynecology, and psychiatry. Specialists evaluated correctness, standards-of-care agreement, bias, timeliness, patient safety, citation authenticity, and contextual appropriateness. OpenEvidence performed strongly on citation authenticity in that evaluation, but ratings for bias and patient safety varied across systems and specialties.6
Because performance varies by question complexity and evaluation dimension, physicians should always treat an AI answer as a draft. They should verify the cited evidence, check for omitted harms or alternatives, and confirm that it applies to the individual patient. OpenEvidence shows promise as a physician-facing tool for quickly locating and synthesizing clinical evidence. Its value is clearest when it helps physicians verify or challenge an initial plan. Yet weaker performance on complex subspecialty scenarios and limitations in the broader question-answering literature argue for careful oversight. Physicians can use OpenEvidence to accelerate evidence-oriented work, but they should verify important references and retain responsibility for the final clinical decision. Prospective trials are needed to evaluate its utility in complex cases and multidisciplinary settings.
References
1. Laserna A, Rubinger DA, Barahona-Correa JE, Wright N, Williams MR, Wyrobek JA, Hasman L, Lustik SJ, Eaton MP, Glance LG. Levels of Evidence Supporting the North American and European Perioperative Care Guidelines for Anesthesiologists between 2010 and 2020: A Systematic Review. Anesthesiology. 2021 Jul 1;135(1):31-56. doi: 10.1097/ALN.0000000000003808. PMID: 34046679.
2. Jagarapu J, Babata K, Chamarthi S, Hoyt R. The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios: a pilot study. medRxiv [Preprint]. 2025 Dec 4. doi:10.64898/2025.11.29.25341091.
3. Hurt RT, Stephenson CR, Gilman EA, Aakre CA, Croghan IT, Mundi MS, Ghosh K, Edakkanambeth Varayil J. The use of an artificial intelligence platform OpenEvidence to augment clinical decision-making for primary care physicians. J Prim Care Community Health. 2025;16:21501319251332215. doi:10.1177/21501319251332215.
4. Jagarapu J, Babata K, Chamarthi S, Hoyt R. The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios: a pilot study. medRxiv [Preprint]. 2025 Dec 4. doi:10.64898/2025.11.29.25341091.
5. Feng J, Patel VR, Heagerty P, Mai Y, Sivaraman V, Vossler P, Ouyang J, Jena A. Expert evaluation of clinical AI tools on real point-of-care clinical queries. arXiv [Preprint]. 2026. arXiv:2606.28960. doi:10.48550/arXiv.2606.28960.
6. Castaño-Villegas N, Villa MC, Monsalve Barrientos K, Llano I, Zea J. Arkangel AI, OpenEvidence, ChatGPT, Medisearch: Are they objectively up to medical standards? A real-life assessment of LLM chatbots in health care. Mayo Clin Proc Digit Health. 2026;4.