Skip to content
The Internet Compass

Artificial Intelligence in Healthcare and Education

Evaluating large language models on medical evidence summarization

Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G. Nestor, Ali Soroush, Pierre Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F. Rousseau, Chunhua Weng, Yifan Peng

Artificial Intelligence in Healthcare and EducationTopic ModelingMachine Learning in Healthcare
Published August 24, 2023Read PDF ↗View on arXiv ↗

Abstract

Recent advances in large language models (LLMs) have demonstrated remarkable successes in zero- and few-shot performance on various downstream tasks, paving the way for applications in high-stakes domains. In this study, we systematically examine the capabilities and limitations of LLMs, specifically GPT-3.5 and ChatGPT, in performing zero-shot medical evidence summarization across six clinical domains. We conduct both automatic and human evaluations, covering several dimensions of summary quality. Our study demonstrates that automatic metrics often do not strongly correlate with the quality of summaries. Furthermore, informed by our human evaluations, we define a terminology of error types for medical evidence summarization. Our findings reveal that LLMs could be susceptible to generating factually inconsistent summaries and making overly convincing or uncertain statements, leading to potential harm due to misinformation. Moreover, we find that models struggle to identify the salient information and are more error-prone when summarizing over longer textual contexts.

Sourced from arXiv · Updated September 2, 2026

Thank you to arXiv for use of its open access interoperability.

View original source ↗Spot an error on this page? Let us know →

FAQ

Common questions

What is "Evaluating large language models on medical evidence summarization" about?

Recent advances in large language models (LLMs) have demonstrated remarkable successes in zero- and few-shot performance on various downstream tasks, paving the way for applications in high-stakes domains. In this study, we systematically examine the capabilities and limitations

Who wrote this paper?

Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G. Nestor, Ali Soroush, Pierre Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F. Rousseau, Chunhua Weng, Yifan Peng

Where can I read the full paper?

The full text is available as a PDF on arXiv (linked above), published August 24, 2023.

Does this paper have a DOI?

Yes: https://doi.org/10.1038/s41746-023-00896-7.