The Role of Advanced Language Models in MedicalDiagnostics: A Case Study on Breast Cancer Prediction

dc.contributor.affiliationMinistry of National Education; Mi̇lli̇ Eği̇ti̇m Bakanliği
dc.contributor.affiliationBaşkent University
dc.contributor.authorHabibe Karayiğit; Ministry of National Education; Mi̇lli̇ Eği̇ti̇m Bakanliği
dc.contributor.authorFiliz Kalelioğlu; Başkent University
dc.contributor.orcidhttps://orcid.org/0000-0002-7518-8699
dc.contributor.orcid
dc.contributor.rorhttps://ror.org/00jga9g46
dc.contributor.rorhttps://ror.org/02v9bqx10
dc.date.accessioned2026-09-07T14:07:17Z
dc.date.issued2026-08-28
dc.date.updated2026-09-07T14:07:17Z
dc.description.abstractBreast cancer remains one of the most common and life threatening cancers worldwide, and early detection is strongly associated with improved survival and reduced treatment burden. This study  investigates the ability of Large Language Models to perform diagnostic prediction from structured breast cancer related data. We systematically evaluated 12 LLMs across three public datasets with different clinical characteristics: the Wisconsin Breast Cancer Dataset (WBCD) based on cytological features, the Breast Cancer Coimbra Dataset (BCCD) based on metabolic biomarkers, and the Mammographic Mass Dataset (MMD) based on mammographic attributes. The evaluation covered 13 prompting strategies, including three zero-shot variants, few-shot prompting, and multiple Chain-of-Thought (CoT) and knowledge-enhanced reasoning settings. Performance was assessed using confusion-matrix-based metrics, including accuracy, precision, recall, F1-score, specificity, and Matthews Correlation Coefficient. The results showed that performance was strongly dependent on both dataset type and prompting design, and no single model dominated all tasks. The best model–strategy pair varied by dataset: Cogito-v1-preview-qwen-32B achieved the highest F1-score on WBCD with an F1-score of 92.00%, GPT 4.1 and GPT 4o on BCCD with an F1-score of 85.39%, and Gemini 2.5 Flash Lite on MMD with an F1-score of 82.91%. Prompt engineering had a substantial effect on outcomes, but its benefit varied across models, with some systems improving under knowledge-enhanced few shot prompting and others performing best under simpler strategies. Comparison with state-of-the-art traditional ML baselines showed that, while LLMs do not yet surpass supervised methods, the performance gap has narrowed substantially, particularly on MMD, where the best single-run gap was 2.04 percentage points in F1 and the mean gap under robustness analysis was approximately 5.3 points. Robustness analysis across multiple few-shot example sets confirmed stable performance on WBCD (F1 = 91.37 ± 0.71%) and BCCD (F1 = 87.46 ± 1.91%), while revealing moderate sensitivity on MMD (F1 = 79.62 ± 2.87%). Although the evaluated LLMs did not outperform traditional supervised models, the study provides a clear performance baseline for future research on structured clinical prediction with language models. The results show that LLMs may offer value as complementary exploratory tools, but their outputs should be interpreted only with expert oversight because clinically significant errors remain.
dc.description.endingpagee2221
dc.description.startingpagee2221
dc.identifier.urihttps://doi.org/10.9781/ijimai.2026.2221
dc.identifier.urihttps://reunir.unir.net/handle/123456789/20567
dc.publisherUniversidad Internacional de La Rioja
dc.relation.ispartof10
dc.relation.ispartofvolume1
dc.rightsopenAccess
dc.rights.uriopenAccess
dc.subjectAI-based Medical Diagnosis
dc.subjectAI-based Prediction
dc.subjectBreast Cancer Diagnosis,
dc.subjectComparative Analysis
dc.subjectLarge Language Models
dc.titleThe Role of Advanced Language Models in MedicalDiagnostics: A Case Study on Breast Cancer Prediction

Archivos

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Nombre:
ijimai10_1_2026_2221.pdf
Tamaño:
723.56 KB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Nombre:
license.txt
Tamaño:
1.65 KB
Formato:
Item-specific license agreed upon to submission
Descripción: