International Journal For Multidisciplinary Research
E-ISSN: 2582-2160
•
Impact Factor: 9.24
A Widely Indexed Open Access Peer Reviewed Multidisciplinary Bi-monthly Scholarly International Journal
Home
Research Paper
Submit Research Paper
Publication Guidelines
Publication Charges
Upload Documents
Track Status / Pay Fees / Download Publication Certi.
Editors & Reviewers
View All
Join as a Reviewer
Get Membership Certificate
Current Issue
Publication Archive
Conference
Publishing Conf. with IJFMR
Upcoming Conference(s) ↓
Conferences Published ↓
DePaul-2026
IC-AIRCM-T3-2026
NSSFIGTMA-2025
SPHERE-2025
AIMAR-2025
SVGASCA-2025
ICRTET-4
ICCE-2025
Chinai-2023
PIPRDA-2023
ICMRS'23
Contact Us
Plagiarism is checked by the leading plagiarism checker
Call for Paper
Volume 8 Issue 5
September-October 2026
Indexing Partners
Machine Learning-Based Crop Yield Prediction Using FAO Global Agricultural Data: A Comparative Study of Linear Regression, Random Forest, and XGBoost
| Author(s) | Abdulhafedh Mohammed Alsarfi, Archana Harsing Sable, Aymen Mosleh Al-Hejri |
|---|---|
| Country | India |
| Abstract | Accurate crop yield prediction is essential for food security planning, agricultural policy-making, and sustainable farming management. This study develops and compares three machine learning models Linear Regression, Random Forest, and XGBoost for predicting crop yield using the FAO (Food and Agriculture Organization) global agricultural dataset. The dataset spans 212 countries and regions, covers 10 major crop types, and encompasses the period 1961–2016, comprising 56,717 records. Following a systematic preprocessing pipeline including missing value verification, duplicate removal, and label encoding, all models were trained and evaluated under identical experimental conditions using an 80:20 train-test split (45,373 training / 11,344 testing records). Performance was assessed using Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and the coefficient of determination (R²). Random Forest achieved the best predictive performance with R² = 0.9586, MAE = 6,724.19 hg/ha, and RMSE = 13,829.60 hg/ha. XGBoost demonstrated competitive performance with R² = 0.8115, while Linear Regression showed limited effectiveness with R² = 0.0608, confirming its inability to model the nonlinear dynamics of global agricultural data. Feature importance analysis identified Item Code (crop type) as the most influential predictor, followed by geographic region and temporal factors. These findings confirm that ensemble learning methods, particularly Random Forest, are well-suited for modelling complex agricultural yield patterns at global scale. |
| Keywords | Crop yield prediction1; Machine learning2; Random Forest3; XGBoost4; Linear Regression5; FAO dataset6; Precision agriculture7; Feature importance8; Food security9. |
| Field | Computer |
| Published In | Volume 8, Issue 4, July-August 2026 |
| Published On | 2026-08-28 |
| DOI | https://doi.org/10.36948/ijfmr.2026.v08i04.86667 |
Share this

E-ISSN 2582-2160
CrossRef DOI prefix of IJFMR is 10.36948/ijfmr
All research papers published on this website are licensed under Creative Commons Attribution-ShareAlike 4.0 International License, and all rights belong to their respective authors/researchers.
Powered by Sky Research Publication and Journals