Vollständiger Abstract
Worum geht es in dieser Arbeit?
Breast cancer as such, is one of the two fallaciously debilitating diseases known to woman (the other being cardiac disease) so in that context and with some qualifications it really shouldn't come as that much of a surprise. The concern, however, is: Yes, you could do it right — early detection is our nearest weapon to the Silent Killer (cancer). However, standard diagnostics are far from 100% accurate, can be mislead by virtually any kind of fake and often miss occasions that cry for assistance but which have not yet proceeded to the malignancy that will eventually confirm. More recently you would see the triggers of real momentum with machine learning(ML) one of the last great fortresses where the majority of frontier artificial intelligence (AI) resides. ML can also be leveraged to guide the accuracy and precision of a breast cancer diagnosis when considering a dataset carefully that includes patient data. The aim in this study was to compare three popular supervised ML algorithms on Logistic Regression (LR), Support Vector Machines (SVM) and K-Nearest Neighbors (KNN). In this domain, the main objective was to effectively categorize malignant breast depending on the benign breast tissue of widely used Wisconsin Breast Cancer Dataset (WBCD) features. An accurate extraction and integration of records belonging to 699 patients was performed for data formatting suitable for analysis. All the required cleaning of data was performed first. This is key to building reliable and credible outcomes, we standardized our functions in a thoroughgoing manner before feeding the model. For feature selection, Gain Ratio is used. Most data points can be covered in a fair manner and performance has to be squeezed up fully. It is an old-school style but it fits our case perfectly: we keep about 30% of records as test samples and the rest (70%) for training (so, we split by 70/30). With each algorithm you performed a full 10-fold cross validation. No stone unturned! We took a fairly wide range of metrics. You covered sensitivity and specificity but what got the most play was with benchmarked net accuracy rates, much less so about false positives/negatives. When we went to so much trouble with other parts of it, it seems an odd time to write off Precision or AUC-ROC as adjuncts to fill in this sector of the picture. They made use of ANOVA to provide statistical comparison among those sets of algorithms that were tested against the data — with a McNemar's test functioning as an objective measuring stick. Now the exciting part began, SVM model absolutely jumped to end up at amazing 97.1% accuracy and 0.994 on AUC score! Logistic Regression followed, with 96.4% accuracy(AUC=0.992), KNN takes the third place with a combined score of 95.6%(AUC=0.978).
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Nada A. Muneam, Asma S. Atteah, Ekhlas Ali Hussein, Suha A. Muneam, Reem Ali Haddad, Nawar Sahib Khalil, Mahmood S. Jameel
- Quelle
- Al-Khwarizmi Engineering Journal
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2312-0789, 1818-1171
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Nada A. Muneam, Asma S. Atteah, Ekhlas Ali Hussein, Suha A. Muneam, Reem Ali Haddad, Nawar Sahib Khalil, Mahmood S. Jameel (2026). Artificial Intelligence in Breast Tumor Classification: A Comparative Study of Logistic Regression, SVM, and KNN Algorithms Using the WBCD Dataset. Al-Khwarizmi Engineering Journal. https://doi.org/10.22153/kej.2026.02.003
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1