Certain investigations on optimizing feature selection for sentiment analysis using hybrid adaboost framework in big data

Abstract

An essential element for modern decision support systems associated newlinewith social networks and other data sources is Big Data Mining. The newlineprocedure in which text analytics is employed for mining numerous data newlinesources for opinions is referred to as Sentiment Analysis (SA). SA is usually newlineconducted on data gathered from the Internet and numerous social media newlineplatforms. Product feedback analysis has feasibility for SA on big data. newlineRecently, it has been useful for individuals to share product information on newlinesocial networking networks. The objective is either the automatic or semi newlineautomatic extraction of user opinions from a large volume of data texts. The newlinedevelopment of efficient information extraction systems capable of processing newlinea huge data volume and extracting customer opinions from the Websites newlineavailable is quite challenging. As a critical data mining research domain, newlinefeature selection will pick a subset of relevant features for model construction. newlineThe extraction of features can be a crucial step in the task of SA. Term newlineFrequency-Inverse Document Frequency (TF-IDF) based technique can newlineeliminate the common terms and extract the relevant terms from a corpus. newlineFeature selection can be a crucial area of research found in data mining and newlinethis selects a new subset of relevant features. These are used for model newlinebuilding. The Genetic Algorithm (GA) is a very efficient method of heuristics newlinethat is employed widely in getting global solutions inside the solution space. newlineThus, to solve the problem of feature selection, there is a GA that is proposed. newlineThe model of Naïve Bayes (NB) classification will compute the class and its newlineposterior probability which is dependent on the sharing of words. newlineClassification and Regression Tree (CART) is formed on a new statistical newlineapproach that had been devised for classification as well as with specific newlinecategorical outcomes or regression with continuous outcomes newline

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced