A Novel Approach For Clustering Big Data With Enhanced Genetic Algorithm For Data Selection
Abstract
Due to the modernization of data handling and volume, data management over
servers has become an issue. Unstructured and unlabeled data adds more complexity to the
storage and the retrieval part. This unlabelled data calls for more effective and efficient
analysis tools. Clustering has always been a favorite way to handle the data but, bigger volume
may produce unexpected low accuracies for traditional clustering approaches. In this work, a
novel framework EGAKM is proposed for the clustering of big data Preprocessing is one of
the earnest steps for data filtration and hence a global stop-word list is utilized to filter the
contents before passing them to cluster distribution. To remove unwanted rows from the
datasets, a meta-heuristic oriented Genetic Algorithm (GA) is used. An attribute-based novel
fitness function (f) is also designed to accompany GA. To evaluate the performance of
EGAKM, it is compared with other clustering methods. Standard Error (SE), root mean
squared error (RMSE), and adjusted R squared error is computed between the distributions of
clusters for evaluation.

