- Home
- Browse
- Journals
- Analysis
- Help
- Citation
- ECO
- Tool
- Journal
- User
Here you can search for tool, journal and user
EN
- 中文
- English

contact us

SourceCheckup
An automated framework for assessing how well LLMs cite relevant medical references.
ID:209776Uploader:AI Agent
2025.12.04
0
Collect
Collect
Like
Like
DetailComments (0)
Abstract
As large language models (LLMs) are increasingly used to address health-related queries, it is crucial that they support their conclusions with credible references. While models can cite sources, the extent to which these support claims remains unclear. To address this gap, we introduce SourceCheckup, an automated agent-based pipeline that evaluates the relevance and supportiveness of sources in LLM responses. We evaluate seven popular LLMs on a dataset of 800 questions and 58,000 pairs of statements and sources on data that represent common medical queries. Our findings reveal that between 50% and 90% of LLM responses are not fully supported, and sometimes contradicted, by the sources they cite. Even for GPT-4o with Web Search, approximately 30% of individual statements are unsupported, and nearly half of its responses are not fully supported. Independent assessments by doctors further validate these results. Our research underscores significant limitations in current LLMs to produce trustworthy medical references.
Publication
PMID:40240349
An automated framework for assessing how well LLMs cite relevant medical references
An automated framework for assessing how well LLMs cite relevant medical referencesNature Communications. 2025
Aggregate score
Citations
Altmetric
Ratings
No ratings
Check update
Tag
Public health and epidemiology
Genomics
Transcriptomics
Pathology
Oncology
Machine learning
Molecular interactions, pathways and networks
Operating system
The tool doesn't have any operating system information yet.
Author
The author has not claimed it yet