Background: Early miscarriage is a common complication following assisted reproductive technology (ART), affecting 20-30% of conceptions and creating challenges for patient counseling and treatment planning. Accurate prediction remains difficult because of the multifactorial nature of miscarriage and variability in ART outcomes. Machine learning (ML) offers promising opportunities for individualized risk assessment.
Objective: This study aimed to develop an ML-based framework for early prediction of chemical miscarriage in infertile women undergoing ART using only pre-treatment baseline clinical data.
Materials and Methods: In this developmental study, clinical data from 3000 infertile couples who attended the Yazd Reproductive Sciences Institute, Yazd, Iran, between March 2020 and September 2022 were extracted from medical records. After preprocessing, including missing data handling, outlier correction, normalization, and class balancing, a benchmark dataset of 1234 samples with 32 features were constructed and made publicly available. Multiple ML algorithms and feature selection methods were evaluated.
Results: Across both imbalanced and balanced datasets using synthetic minority oversampling technique and adaptive synthetic sampling, the random forest model achieved the best performance, with area under the receiver operating characteristic curve values of 70.9%, 70.3%, and 70.8%, and corresponding accuracies of 82.2%, 81.1%, and 80.0%, respectively. 12 clinically relevant predictors were identified, including maternal age, body mass index, previous abortions, thyroid-stimulating hormone levels, sperm parameters, ART-related protocols, history of intrauterine insemination, and menstrual regularity.
Conclusion: ML models, particularly random forest, can provide meaningful early prediction of miscarriage risk using baseline clinical data. The publicly available benchmark dataset supports reproducible research and future advances in predictive reproductive medicine.
Send email to the article author