An Approach To Active Learning From Imbalanced Data Using Machine Learning Techniques

Authors

  • P. Aravind, Dr.B Sankara Babu

Abstract

In general, unbalanced data has a severe impact on the predictive model. Most basic sampling techniques are time-consuming and lose samples containing important information while processing the unbalanced data. Active mastery increases labeling efficiency, while only subsets of unlabeled data sets can be manually processed. As far as we know, the current algorithms are designed to assume that data sets are balanced. Several algorithms have been developed online, for example, the Perceptron rule set, exponential weighted average algorithm, and net gradient ratios. We examined the core packages of those algorithms in the device study and data analysis in the remaining ten years, such as online classification. The traditional inline class targets to reduce the range of errors by minimizing the cumulative convex loss function. However, many real-existence datasets are surely imbalanced and we support Principal component analysis (PCA) for dimensionality reduction technique without any loss of given facts, where it allows observing the correlations and methods within the given data set. The proposed method is compared with previous classifier of support vector machine (SVM) in the parameter of dimensionality reduction on imbalanced dataset.

Published

2020-12-30

Issue

Section

Articles