Skip to content
Research Article Open access CC BY 4.0

A Narrative Review on Machine Learning Methods for Pre-Malignant Blood Cells Identification in Hematological Malignancies

Samuel Bamidele Afolabi, TOBECHI BRENDAN NNANNA, Glory Ojoma, Simon, Damian Ndubuisi NWAJEI, Victor Damilare Oladele, Adibia, Umoroye Nathan, Tobiloba Philip Olatokun

International Research Journal of Oncology · pp. 134–150 · Published 30 Apr 2026

10.9734/irjo/2026/v9i1203

Abstract

Hematological malignancies, including leukemias, lymphomas, myelomas, and myelodysplastic syndromes, impose a substantial global health burden, accounting for approximately 10% of all new cancer diagnoses worldwide. Early identification of pre-malignant blood cells a critical window for preventive intervention remains clinically challenging due to the limitations of conventional diagnostic tools such as light microscopy, flow cytometry, and next-generation sequencing, none of which is individually optimized for risk stratification prior to overt disease manifestation. This review examines machine learning (ML) approaches for classifying pre-malignant blood cells, synthesizing evidence from 25 studies encompassing 38,417 participants across diverse clinical settings. Ensemble methods and Random Forest algorithms demonstrated consistently strong discriminative performance, achieving AUC-ROC values ranging from 0.856 to 0.932. Multi-omics integration combining morphological, immunophenotypic, genetic, and epigenetic data systematically outperformed single-domain approaches, underscoring the biological complexity of pre-malignant transformation. Key predictive biomarkers identified across studies included CD34 expression levels, telomere length attrition, and TP53 mutation status, consistent with established pathways of clonal hematopoietic evolution. Despite these promising findings, significant methodological limitations were identified: external validation was reported in only 44% of studies, and open-source code availability was documented in just 40%, raising concerns about reproducibility and generalizability. Additionally, most training cohorts lacked demographic diversity, limiting applicability across varied populations. Successful translation of ML-based pre-malignant cell classification into routine clinical practice will require prospective validation trials, standardized reporting frameworks aligned with existing diagnostic criteria, and the development of ethnically and geographically diverse training datasets.

Machine learning hematological malignancies pre-malignant blood cells cancer risk stratification multi-omics integration

Cited by 0

No indexed citations yet.

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

0

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.