A Narrative Review on Machine Learning Methods for Pre-Malignant Blood Cells Identification in Hematological Malignancies
Samuel Bamidele Afolabi, TOBECHI BRENDAN NNANNA, Glory Ojoma, Simon, Damian Ndubuisi NWAJEI, Victor Damilare Oladele, Adibia, Umoroye Nathan, Tobiloba Philip Olatokun
International Research Journal of Oncology · pp. 134–150 · Published 30 Apr 2026
10.9734/irjo/2026/v9i1203Abstract
Hematological malignancies, including leukemias, lymphomas, myelomas, and myelodysplastic syndromes, impose a substantial global health burden, accounting for approximately 10% of all new cancer diagnoses worldwide. Early identification of pre-malignant blood cells a critical window for preventive intervention remains clinically challenging due to the limitations of conventional diagnostic tools such as light microscopy, flow cytometry, and next-generation sequencing, none of which is individually optimized for risk stratification prior to overt disease manifestation. This review examines machine learning (ML) approaches for classifying pre-malignant blood cells, synthesizing evidence from 25 studies encompassing 38,417 participants across diverse clinical settings. Ensemble methods and Random Forest algorithms demonstrated consistently strong discriminative performance, achieving AUC-ROC values ranging from 0.856 to 0.932. Multi-omics integration combining morphological, immunophenotypic, genetic, and epigenetic data systematically outperformed single-domain approaches, underscoring the biological complexity of pre-malignant transformation. Key predictive biomarkers identified across studies included CD34 expression levels, telomere length attrition, and TP53 mutation status, consistent with established pathways of clonal hematopoietic evolution. Despite these promising findings, significant methodological limitations were identified: external validation was reported in only 44% of studies, and open-source code availability was documented in just 40%, raising concerns about reproducibility and generalizability. Additionally, most training cohorts lacked demographic diversity, limiting applicability across varied populations. Successful translation of ML-based pre-malignant cell classification into routine clinical practice will require prospective validation trials, standardized reporting frameworks aligned with existing diagnostic criteria, and the development of ethnically and geographically diverse training datasets.
Cited by 0
No indexed citations yet.
Related research
- Detecting Dental Caries through Captured Images Using the Machine Learning Technology Teachable Machine — shares topic coverage
- Prediction of Radiotherapy Dose Distribution for Glioblastoma Using Convolutional Neural Network Model — shares topic coverage
- A Systematic Literature Review of Machine Learning Methods in Healthcare — shares topic coverage
- Diagnostic Accuracy of Artificial Intelligence for Breast Cancer Detection: A Systematic Review — shares topic coverage
- Artificial Intelligence in the Analysis of the Fetal Genome in Utero: A Critical Review of Current Paradigms, Clinical Utility and Future Horizons — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
0
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.