Research on the Application of Agricultural Big Data Processing with Hadoop and Spark
Yan Cheng, Qiang Zhang, Ziming Ye
Abstract
Yan Cheng, Qiang Zhang, Ziming Ye
Abstract
Numerous terminal equipment in the agricultural park collect environmental data that affects crop growth every day. Proper analysis of these massive amounts of data can acquire useful information on the status of crop growth. In this paper, two cloud computing frameworks, Apache Hadoop and Apache Spark are used to study agricultural big data analysis. This paper developed applications for real agricultural park big data analysis in both frameworks and implemented a yield prediction model based on multiple linear regression using Spark MLlib. The performance of the two frameworks in agricultural big data processing was studied and compared through various experiments. The experiments show that the comprehensive performance of Spark is higher than Hadoop, and the model can obtain better prediction results.
OpenAlex reports 9 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Numerous terminal equipment in the agricultural park collect environmental data that affects crop growth every day. Proper analysis of these massive amounts of data can acquire useful information on the status of crop growth. In this paper, two cloud computing frameworks, Apache Hadoop and Apache Spark are used to study agricultural big data analysis. This paper developed applications for real agricultural park big data analysis in both frameworks and implemented a yield prediction model based on multiple linear regression using Spark MLlib. The performance of the two frameworks in agricultural big data processing was studied and compared through various experiments. The experiments show that the comprehensive performance of Spark is higher than Hadoop, and the model can obtain better prediction results.
Key concepts: SPARK (programming language), Big data, Computer science, Agriculture, Cloud computing, Database, Data mining, Operating system