Machine learning classification of mutagen treatment types in crop breeding: A comparative analysis

Abstract
Mutation breeding is a vital tool for crop improvement, yet mutagen selection still relies largely on trial-and-error. This study developed a machine learning (ML) framework to predict mutagen treatment types using 2815 curated entries from the FAO/IAEA Mutant Variety Database. The dataset included crop descriptors, mutagen details, and trait outcomes. Preprocessing involved removal of incomplete entries, standardization of categories, and class balancing with SMOTE and weighted models. Three classifiers—Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression (LR)—were trained using stratified 80/20 splits and 10-fold cross-validation. Model performance was assessed with accuracy, precision, recall, F1-score, and ROC-AUC, with statistical comparisons via Friedman and Wilcoxon tests. RF and SVM achieved the highest accuracies (96.3 %), while LR performed slightly lower (95.7 %). SVM demonstrated superior recall (0.695) and F1-score (0.624), improving detection of minority classes such as EMS and somaclonal variation (p < .05). Gamma rays dominated overall, but EMS produced broader phenotypic and agronomic improvements. These findings demonstrate ML—particularly SVM—as a decision-support tool to guide mutagen selection, reduce experimental inefficiency, and accelerate breeding for food and climate resilience.

Author
Mehdi Rahimi

DOI
https://doi.org/10.1016/j.egg.2025.100427

ISSN
2405-9854

Publish Date: 17-Nov-2025

اتصل بنا

التسجيل: +964 750 3000 600
التسجيل: +964 750 3000 700
الرئاسة: +964 750 3000 800

ارسل لنا عبر البريد الإلكتروني

[email protected]