KENDALL RANK CORRELATION ANALYSIS BASED DISCRIMINANT RANDOM FOREST MAPREDUCE CLASSIFICATION FOR BIG DATA ANALYTICS
Abstract
Classification is a key problem to be resolved in data mining for efficient big data analytics. Recently, many research works have been done for big data classification. However, the error rate and time complexity involved during the classification process while taking big data as input using conventional algorithms was very higher. To overcome such limitations, The Kendall Rank Correlation Analysis Based Discriminant Random Forest Mapreduce (KRCA-DRFM) Classifier is proposed in this work. The KRCA-DRFM Classifier is designed for reducing the dimensionality of big data with minimal execution time and error rate. Initially,KRCA-DRFM Classifier obtains a large size of the dataset as input. After taking input, KRCA-DRFM Classifier arbitrarilybuilds bootstrap samples usingseveral data in a given dataset. Next, KRCA-DRFM Classifier generatesm’ number of regression tree results for all the data in the bootstrap sample using Kendall Rank Correlation Analysis. To precisely determine the homogeneity of each input data, Kendall Rank Correlation Analysis is utilized in KRCA-DRFM Classifier on the contrary to traditionalworks. Subsequently, KRCA-DRFM Classifier uses a voting scheme aiming at applying vote for each regression tree output. At last, KRCA-DRFM Classifier accurately classifiesall the input data into anassociated class with a lower amount of timeaccording to the majority voting result. By an effective classification of big data,the proposed KRCA-DRFM Classifier increases big data analytics performance when compared to state-of-the-art works. The KRCA-DRFM Classifier carried out experimental processes using factors such as classification accuracy, classification time, error rate and space complexity concerninga varied number of data from the El-Nino dataset and Beijing PM2.5 dataset.Downloads
Published
2020-12-30
Issue
Section
Articles

