AI-enabled clinical decision support in breast cancer care: a blinded multicenter benchmarking study comparing medically specialized with a general-purpose system
A general-purpose AI chatbot outperformed specialized cancer AI systems at recommending breast cancer treatments, suggesting simpler tools may be surprisingly effective in clinical practice.
In a blinded multicenter study across 7 university breast cancer centers, a general-purpose LLM (ChatGPT-5 Thinking) significantly outperformed two regulatory-cleared, specialized AI-CDS systems across safety, guideline adherence, completeness, and overall quality for breast cancer treatment planning. These findings challenge the assumption that medical specialization and regulatory clearance translate to superior clinical AI performance, with broad implications for AI medical device adoption strategies.
What the study was
- Study design
- Blinded multicenter benchmarking study (7 university breast cancer centers, 20 standardized cases, 3 AI-CDS systems)
- Population
- Breast cancer patients; 7 academic center raters evaluating AI-CDS outputs
- Sample size
- 20
- Category
- Diagnostics
- Maturity
- Validated
- Journal
- J Med Syst
Why it surfaced
Rigorous blinded 7-center benchmarking showing GPT-5 outperforms FDA-cleared specialized AI-CDS in breast cancer; challenges AI regulatory certification assumptions; highly relevant to clinical AI watchlist.
A plain-language summary of published research — not medical advice. Talk to a clinician about your own care.