Identifying neutral reviews from unlabeled data: An exploratory study on user ratings and word-level polarity scores
The presence of the reviews containing mixed or contrasting opinions, also known as neutral reviews, is prevalent in user feedback data. By leveraging annotated data, supervised machine learning (ML) classifiers can learn implicit patterns to identify these neutral reviews. However, labeled data are barely available in most circumstances. When annotated data are unavailable, unsupervised approaches such as lexicon-based methods are employed that utilize word-level polarity scores with a set of rules. As a preliminary study for developing a sophisticated unsupervised framework for recognizing neutral reviews, here, we scrutinize the performances of the existing lexicon-based methods. When app


ACM source attribution. Complete text was converted from the supplied ACM publisher HTML and verified against the authorized ACM archival PDF for 10.1145/3511095.3536367. Complete source-page facsimiles preserve figures, tables, captions, equations, and layout. © 2022 Association for Computing Machinery.

Identifying neutral reviews from unlabeled data: An exploratory study on user ratings and word-level polarity scores

Authors: Salim Sazzed

Complete visual source pages

Complete source page 1: figures, tables, equations, and captions
Complete source page 2: figures, tables, equations, and captions
Complete source page 3: figures, tables, equations, and captions
Complete source page 4: figures, tables, equations, and captions
Complete source page 5: figures, tables, equations, and captions
Complete source page 6: figures, tables, equations, and captions

Complete text

Abstract

The presence of the reviews containing mixed or contrasting opinions, also known as neutral reviews, is prevalent in user feedback data. By leveraging annotated data, supervised machine learning (ML) classifiers can learn implicit patterns to identify these neutral reviews. However, labeled data are barely available in most circumstances. When annotated data are unavailable, unsupervised approaches such as lexicon-based methods are employed that utilize word-level polarity scores with a set of rules. As a preliminary study for developing a sophisticated unsupervised framework for recognizing neutral reviews, here, we scrutinize the performances of the existing lexicon-based methods. When applied to four multi-domain review datasets, we observe that all of them perform poorly for identifying neutral reviews. We manually inspect the semantic attributes of a subset of neutral reviews classified wrong by these lexicon-based methods. The experimental results and manual analysis reveal that determining neutrality utilizing the lexical rule-based methods is often ineffective due to numerous reasons, such as user preferences on certain aspects, coverage of the sentiment lexicon, irregularly in the efficacy of aggregation rules, and the context-sensitive polarity of words. As a preliminary study, this analysis reveals traits of neutral reviews and limitations of existing approaches and provides insights to develop methods for neutral review identification from the unlabeled data.

CCS Concepts

CCS Concepts: • Social and professional topics → User characteristics ; • Information systems → World Wide Web ;

Keywords

Keywords: neutral review , unlabeled data , sentiment analysis , sentiment lexicon , context sensitivity

ACM Reference Format

ACM Reference Format: Salim Sazzed. 2022. Identifying neutral reviews from unlabeled data: An exploratory study on user ratings and word-level polarity scores. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT '22), June 28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA 5 Pages. https://doi.org/10.1145/3511095.3536367

Sentiment analysis systematically analyzes people's sentiments or opinions towards various entities, such as products, organizations, events, and activities. Based on the opinions and views present in a piece of textual content, it attempts to identify its semantic orientation (e.g., positive, neutral, or negative). The researchers utilized various approaches for sentiment classification which can be broadly categorized into the following three categories: i) lexicon-based [ 19 ] ii) machine learning(ML)-based [ 1 , 17 ], and iii) hybrid approach [ 2 , 16 , 18 , 22 ] combining both i) and ii). The lexicon-based approach [ 15 ] relies on a sentiment dictionary that contains words with weighted or unweighted polarity scores. Lexicon-based approaches play a key role for sentiment classification [ 19 ]. Although supervised ML classifiers usually perform better than the lexicon-based methods, they require annotated data, which are not always available due to the cost associated with the data labeling process. Thus, approaches which do not require labeled data, such as lexicon-based approaches, are still important for sentiment analysis.

Neutrality may be considered somewhere between positivity and negativity (e.g., a user rating of 3 on a scale of 1-5 [ 9 , 10 , 10 , 12 ]). When users undergo contrasting attitudes towards different aspects of an entity, they often rate the entire experience as neutral or mixed (e.g., a rating of 3 out of 5). The definitions of neutrality are often not consistent in the various literature as the term ’neutrality has been employed to refer to the absence of any kind of sentiment, opinions expressing mixed or conflicting sentiments or both [ 21 ]. In this study, by neutrality, we refer to the reviews that contain both positivity and negativity to some extent which makes the reviewer select a rating close to the middle of the possible range.

For analyzing sentiment at the document level, existing unsupervised approaches mainly considered the binary classification task of positive or negative class prediction. Even when the 3-class prediction (by adding the neutral class) was reckoned, the datasets represented primarily short texts (mostly single sentences), and the aggregated results were reported. Since the neutral class usually constitutes a comparatively low proportion of the dataset, they have a low impact on the overall performance of the classifiers. Thus, a comprehensive analysis regarding neutrality detection in the reviews was mostly ignored. However, it is still important to analyze neutral examples for a better understanding of the sentiments expressed in a text [ 7 ]. Besides, even though they exist in comparatively lower numbers in a dataset, the poor performance in the neutral class prediction will negatively affect the performance of the classifiers, specifically, when the neutral class is treated equally with the positive and negative classes (e.g., macro F1 score).

Although few studies investigated the presence of neutral reviews in various datasets [ 3 , 5 , 21 ], they focused on dissimilar facets of neutrality than this work. For example, [ 5 ] and [ 3 ] used supervised ML classifiers with annotated data. With the availability of labeled data, ML classifiers can predict neutrality with an acceptable level of accuracy. The work of [ 21 ] showed that excluding neutral reviews helps improve the performance of ML classifiers for binary-level sentiment classification.

Unlike previous studies, here, we emphasize neutrality determination from unlabeled data. We examine the performances of a set of popular lexicon-based methods, VADER [ 6 ], Textblob 1 , Sentistrength [ 20 ], AFINN [ 11 ], and LRsentiA [ 18 ], which can classify reviews without using any labeled data. Four review datasets comprising user opinions towards various entities from multiple domains are utilized. Besides, we study the consensus of the predictions among these sentiment analysis methods to see whether the predictions are consistent among the various lexicon-based methods. We observe that all the lexicon-based methods perform poorly in identifying the neutral reviews. Moreover, we notice very low agreements between each pair of these methods; Cohen κ scores between 0.15 to 0.50 are observed between various pairs. We manually analyze a set of wrongly classified neutral reviews to pinpoint why the lexicon methods fail (most if not all). We notice that several factors, including the simplistic aggregation rules, coverage of lexicon, word-level polarity strengths, and user's choice of particular aspects, make the neutral review identification task very challenging for the lexicon-based methods.

To find how the predictions considering aggregated word-level polarity scores perform for sentiment classification, we consider five popular lexicon-based methods, AFINN [ 11 ], VADER, SentiStrength [ 20 ], Textblob, and LRSentiA [ 18 ]. All of these methods (except the recently proposed LRSentiA) showed comparatively better performance than other lexicon-based methods on a number of benchmark datasets [ 14 ].

AFINN: The AFINN lexicon consists of 2477 English words and phrases annotated for valence with an integer score between -5 and +5, where -5 refers to (strongly negative ) and +5 refers to (strongly positive ). We utilize the PyPI implementation of AFINN 2 to compute the polarity scores of the reviews. For 3-class predictions, polarity score above 0 refers positive prediction, polarity score below 0 means negative class,and a 0 polarity score represents neutral class.

SentiStrength: SentiStrength is a sentiment analysis tool that leverages a dictionary of sentiment words with associated polarity strength to determine positive and negative sentiment strength from the short informal text. SentiStrength employs a dual 5-point scoring system for positive and negative sentiment. Based on the authors, a positive sentiment score refers to positive positive prediction, a negative score refers to a negative class, and a 0 polarity score indicates neutral prediction.

VADER: VADER (Valence Aware Dictionary and sEntiment Reasoner) [ 6 ] is a lexicon and rule-based sentiment analysis tool which was specially designed for determining sentiments expressed in social media.

According to the authors of VADER, for 3-class prediction, a compound score greater than +0.05 indicates positive prediction, a score less than -0.05 indicates as negative score, and a score between -0.05 and +0.05 as neutral prediction. We utilize the PyPI VADER package implementation 3 .

TextBlob:

TextBlob calculates both the polarity and subjectivity scores from the text utilizing weighted word-level polarity scores. In binary-level, a non- negative polarity score refers to positive prediction. When 3-class is considered, a polarity score above 0.05 indicates the positive class, below -0.05 refers to the negative class, and values in between these two ranges attribute to the neutral class.

LRSentiA: LRSentiA is a lexicon-based classifier that can predict the sentiment of a review without using any labeled data. LRSentiA utilizes a binary-level sentiment lexicon and a set of rules to predict the polarity of a review. In addition to predicting the class, LRSentiA provides the confidence score of the prediction. For the LRSentiA, a predicted polarity score of 0 refers to the neutral class.

Dataset

#Neutral Reviews

MP3

2529

Food

2150

Hotel

4759

Restaurant

1453

Dataset

Method

#Wrong Prediction

# Correct Prediction

Acc(%)



Predicted as Neg / Pos




VADER

544 / 1913

72

2.85%

MP3

Textblob

294 / 2191

44

1.74%


SentiStrength

1291 / 1173

65

2.57%


AFINN

436 / 1936

157

6.20


LRSentiA

588 / 1585

356

14.00%


VADER

371 / 1693

85

3.95%

Food

Textblob

305 / 1814

31

1.44%


SentiStrength

734 / 1321

95

4.42%


AFINN

226 / 1729

135

6.27%


LRSentiA

393 / 1368

389

18.01%


VADER

730 / 3972

56

1.18%

Hotel

Textblob

516 / 4238

4

0.8%


SentiStrength

2211 / 2515

33

0.69%


AFINN

580 / 4005

174

3.65%


LRSentiA

786 / 3570

403

8.46%


VADER

127 / 1297

28

1.93%

Restaurant

Textblob

94 / 1344

15

1.32%


SentiStrength

463 / 964

26

1.79%


AFINN

88 / 1299

66

4.54%


LRSentiA

153 / 1154

146

10.03%

We utilize four datasets 4 : MP3 , Hotel , Food and Restaurant in this study. The reviews represent user feedback towards entities from varied domains. The datasets used in this study are selected from diverse domains for a robust analysis. The original MP3 dataset consists of over 27000 reviews, with ratings between 1 and 5 (integer). The TripAdvisor hotel dataset consists of 23600 reviews with 5 different ratings (-4, -2, 0, 2, 4). The Yelp dataset contains 10000 reviews with a rating in 5-point scale (1 to 5). The food review dataset contains 20000 reviews in 5-point scale between 1-5. To select the neutral reviews from the above datasets, we used the guideline provided by the existing literature [ 9 , 10 , 12 , 13 ]. [ 12 ] considered rating above 2 stars and below 3.5 stars as neutral in 5-point scale 5 . [ 9 ] considered < = 4 as negative and > = 7 as positive in their study on a 10-point rating. [ 10 ] considered 3 as neutral in their study. The statistics of neutral reviews in various datasets are shown in Table 1 .

We report the accuracy (recall) of various methods for the neutral review prediction. Note that the precision score for the neutral class is not informative since we find the classifiers hardly predict anything as neutral (shown in the Table 2 ). A low false positive (FP) prediction is observed for the neutral class that yield a high precision score for the neutral class, which is not a real indicator of the performance of the classifiers for neutral class prediction.

The accuracy (recall) of a method for the neutral class prediction is defined as follows -

$accuracy = frac{ Numbers; of; identified ;neutral; reviews}{Total; Numbers; of; neutral; reviews ; in ; corpus } ; times 100$

Besides, we report the breakdown of the false-negative predictions of the neutral class, such as how many neutral reviews are predicted as positive and how many are predicted as negative.

Table 2 shows the performances of various lexicon-based methods for detecting neutral reviews. We can see (Table 2 ) that all the lexicon-based methods perform poorly for predicting neutral reviews. We analyze predictions of various methods to find the most dominant false-negative group of the neutral class. As we can see from Table 2 , the types of false-negative predictions (positive or negative) for the neutral class vary across various lexicon-based methods. We observe most of the methods are inclined toward the positive class (except SentiStrength).

We investigate the agreements in the predictions of various lexicon-based methods. We seek to find if all the lexicon-based methods show similar types of errors for the predictions. To find the agreements in the predictions of various pairs of lexicon-based methods, we compute Cohen's kappa statistics. Cohen's kappa is a statistical measure used to gauge inter-rater reliability, where a score of 1 refers to a perfect agreement.

Based on [ 8 ], κ value can be interpreted as follows: < 0 = No agreement, 0.0 − 0.20 = Slight agreement, 0.21 − 0.40 = Fair agreement, 0.41 − 0.60 = Moderate agreement, 0.61 − 0.80 = Substantial agreement and 0.81 − 1.0 = Perfect agreement.

Dataset

Food

MP3

Hotel

Restaurant

VADER-TextBlob

0.29

0.30

0.37

0.29

VADER-SentiStrength

0.25

0.21

0.19

0.17

VADER-AFINN

0.41

0.48

0.49

0.41

VADER-LRSentiA

0.27

0.34

0.38

0.28

TextBlob-SentiStrength

0.22

0.15

0.16

0.13

TextBlob-AFINN

0.30

0.35

0.39

0.33

TextBlob-LRSentiA

0.23

0.25

0.33

0.24

SentiStrength-AFINN

0.28

0.24

0.19

0.22

SentiStrength-LRSentiA

0.24

0.20

0.19

0.21

AFINN-LRSentiA

0.25

0.31

0.37

0.27

As seen by the Table 3 , there exist minimal agreements between various methods in the predictions. For most of the pairs, the κ values are in the range of 0.20-0.40, which indicates fair agreements. The best agreement is observed between VADER and AFINN. For all the datasets, they show κ values between 0.40-0.50, which indicates moderate agreement. The worst consensus is found between TextBlob and SentiStrength, which ranges from 0.13 to 0.22 across the datasets.

We manually analyze a number of neutral reviews, which are predicted wrong by most (if not all) of the lexicon-based methods.

Example 1: ”These things are just too darn cheesy. If you like a lot of flavor, you'll love these. Otherwise, you will overdose on their cheddary goodness.”

Analysis: The above review is predicted positive with very high confidence by the VADER. Textblob correctly classifies it as neutral; however, among a list of possible opinion words in the reviews, it deems polarity for only two words, ’cheesy’ and ’love’. It does not consider other words such as ’like’, ’darn’, or ’overdose’ as opinion words due to the lexicon coverage issue. VADER has a better lexicon coverage in this review; it identifies the above-mentioned three words as opinion conveying words. However, the aggregation of word-level polarity scores provides a high compound score of 0.865, which indicates a positive prediction. LRSentiA correctly classifies it as a neutral; however, since it uses a binary lexicon, a zero polarity score refers to that it finds the same number of positive and negative words/phrases, which could happen due to lexicon coverage or simple aggregation rules. AFINN indicates this review as positive, it also suffers from the lexicon-coverage issue. The prediction suggests that in addition to the lexicon coverage, aggregation rules play an important part in the final decision-making process.

Example 2: ”I love, love, love the idea of this product and my little one loves, loves, loves to use it. She gets fresh fruit and other tasty things and I don't have to worry about her choking. The only drawback is that bits of food get stuck in the mesh and the seams and are impossible to get out. Granted, some foods are more difficult than others and some don't cause that much of a problem. Bananas however, have pretty much destroyed a couple of them. While I love this product, I think that next time I will get the one that has the removable and replaceable mesh bags.”

Analysis: The above review is predicted as positive by the VADER with high a compound score of 0.9758 (where > 0.05 means a positive prediction). A high frequency of the word ’love’ in the review makes most of the lexicon-based methods assume it highly positive. LRSentiA correctly identifies it as a neutral review that indicates it notices an equal presence of positive and negative words or phrases in the review based on its lexicon.

Example 3: ”The taste was great, but the berries had melted. May order again in winter. If you order in cold weather you should enjoy flavor.”

Analysis: This is an example of neutral review, where context-dependency of the word plays a crucial role in determining the overall polarity of a text. Although ’melt’ is not an opinion word, in the text, ”The taste was great, but the berries had melted”, considering the context, ’melted’ indicates a negative opinion. As anticipated, none of the lexicon-based methods can capture it.

Example 4: ”To my mind, a fine caramel should be creamy and melting. These are grainy. The flavor on these was fine, although a little sweet to my taste.”

Analysis: This is another neutral review that does not hold any negative opinion words. The implicit negative sentiment carrying words here are ’should be’ and ’grainy’. Unsurprisingly, all the lexicon-based methods except LRSentiA fail to capture these.

Example 5: ”The taste is harsh as it is a lesser grade of Yunnan. It is ok to drink, but not to really enjoy. I only make it for when I am in the hurry and will not be able to fully enjoy my tea anyway.”

Analysis: VADER predicts this review as negative with a high confidence. Surprisingly, both Textblob and AFINN think it is positive. The presence of the word ’enjoy’ makes Textblob and AFINN think it of as positive. VADER, on the other hand, can comprehend the context better. It captures that the entire phrase ”not to really enjoy” expresses a negative opinion.

Example 6: ”The texture is like rice crackers, which I personally don't like. Taste isn't too bad, but every single bags are too salty, even the original flavor. I'd give four stars if they make low sodium chips. The salt.. is one of your worst enemies. I bet it does more harms to your body than MSG's.”

Analysis: VADER, TextBlob, LRSentiA, and AFINN classify the above review as negative. Although the true label of this review is neutral (based on the user rating), from the review text, a negative true label seems more appropriate. This review indicates that the user rating may not always be fully consistent with the text.

Based on the experimental results and analysis, we find that numerous aspects make the identification of neutral review very challenging such as-

4.5.1 Precedence of Opinion: We observe that among the multiple contrasting opinions present in the review towards various entities, it is often difficult to identify which ones play a more influential role in the conclusive polarity label. Aggregating the word-level polarity to infer the orientation often fail to correctly identify that, which is reflected in the performances of various lexicon-based methods. Besides, the final rating (i.e., true class) of a review often depends more on the reviewer's discretion than the text content, which makes predicting neutrality even more challenging.

4.5.2 Context Dependency: One of the issues of the lexicon-based method is that they use the prior polarity scores of words [ 4 ]. A list of statically defined word-level polarity scores out of context may misguide classifiers in many circumstances. To give an example, the polarity score of ”cold” varies between ”cold beer” and ”cold pizza” due to context differences [ 4 ]. For a highly polarized review (i.e., extremely positive or negative), the incorrect prior polarity score(s) of one or multiple words may have a limited impact on the final semantic orientation. However, for the highly sensitive neutral review, ignoring the context may change the orientation of the prediction.

4.5.3 Polarity Scores of Opinion Words: We observe that using the weighted (i.e., any value within a particular range) or binary-level (discrete value of -1 and +1 for the negative and positive words, respectively) polarity scores of the words do not substantially affect the prediction accuracy of the neutral class. Although using the binary polarity lexicon, LRSentiA performs a bit better than others, yet its performance is poor. Besides, in addition to the polarity lexicon, aggregation rules influence the final prediction.

4.5.4 The Ranges of Polarity Scores for the Neutral Class: Lexicon-based methods assign the neutral label to a review when the combined polarity values within some particular ranges. For example, VADER assumes a predicted compound score between < − 0.05, +0.05 > as neutral. In this study, we follow the original author(s) suggested ranges for deciding neutrality. Since we obtain poor results for the neutral class prediction, we explore whether we can achieve better results by utilizing different ranges of values. For example, instead of using a 0 polarity score for determining the neutral class by Textblob/Sentistrength, we employ some other ranges such as < − 0.05, +0.05 >, < − 0.10, +0.10 > to assigning the neutral class. Similar ranges are employed for other methods too. However, we do not observe any noticeable difference by varying the range of the neutral class. A wide range for the neutral class decreases the false-negative prediction, but at the same time, it raises the false-positive count; thus, the overall classification performance does not improve much compared to the original results.

Our manual and automatic analysis reveal that expecting a neutral review to have close to 0 polarity score is an over-simplified assumption. A number of factors such as context-dependent polarity scores of opinion words, aggregation rules, or the complexity of natural language may cause a non-zero polarity score of the neutral reviews. However, what polarity value should be used to denote the neutral review is another unresolved research question.

4.5.5 Discrepancy and Ensemble of the Predictions of Various Methods. We find that the intensity of the word-level polarity, coverage of the sentiment lexicon, and sentiment aggregation procedures influence the discrepancy in the predictions of various classification methods (low κ scores). In addition, the performances of the lexicon-based methods indicate that combining predictions of multiple approaches is not an effective prospect for identifying neutral reviews as all of them perform miserably.

As a preliminary study of determining neutral reviews from unlabeled data, here, we explore diverse aspects of neutral reviews from various perspectives which is required to develop sophisticated methods. We demonstrate that the existing unlabeled approaches (i.e., lexicon-based) perform very poorly for the neutral review predictions. We scrutinize the effectiveness of a set of popular lexicon-based methods for neutral review prediction and dissect the rationales behind the results. Our analysis reveals that detecting neutrality using rule-based methods is very challenging due to many facets, such as the impact and bias of contrasting opinions on the overall polarity, context-dependency of the word-level polarity scores, and inconsistent relationships among aggregation rules and ground-truth labels, and complexity of natural language. Our future work will focus on generating pseudo-label utilizing the aspect-level sentiment. We will consider the influences of various aspects on the overall semantic orientations and incorporate that information to train ML classifiers for better predictions.

    Apoorv Agarwal, Boyi Xie, Ilia Vovsha, Owen Rambow, and Rebecca J Passonneau. 2011. Sentiment analysis of twitter data. In Proceedings of the workshop on language in social media (LSM 2011) . 30–38.

    Orestes Appel, Francisco Chiclana, Jenny Carter, and Hamido Fujita. 2018. Successes and challenges in developing a hybrid approach to sentiment analysis. Applied Intelligence 48, 5 (2018), 1176–1188.

    Luis Chiruzzo, Mathias Etcheverry, and Aiala Rosá. 2020. Sentiment analysis in Spanish tweets: Some experiments with focus on neutral tweets. (2020).

    Marco Guerini, Lorenzo Gatti, and Marco Turchi. 2013. Sentiment analysis: How to derive prior polarities from SentiWordNet. arXiv preprint arXiv: 1309.5843 (2013).

    AR Hamed, Renxi Qiu, and Dayou Li. 2016. The importance of neutral class in sentiment analysis of Arabic tweets. Int. J. Comput. Sci. Inform. Technol 8 (2016), 17–31.

    Clayton J Hutto and Eric Gilbert. 2014. Vader: A parsimonious rule-based model for sentiment analysis of social media text. In Eighth international AAAI conference on weblogs and social media .

    Moshe Koppel and Jonathan Schler. 2006. The importance of neutral examples for learning sentiment. Computational Intelligence 22, 2 (2006), 100–109.

    J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977), 159–174.

    Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies . 142–150.

    Andrius Mudinas, Dell Zhang, and Mark Levene. 2012. Combining lexicon and learning based approaches for concept-level sentiment analysis. In Proceedings of the first international workshop on issues of sentiment discovery and opinion mining . 1–8.

    Finn Årup Nielsen. 2011. A new ANEW: Evaluation of a word list for sentiment analysis in microblogs. arXiv preprint arXiv: 1103.2903 (2011).

    Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. arXiv preprint cs/0409058 (2004).

    Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. arXiv preprint cs/0506075 (2005).

    Filipe N Ribeiro, Matheus Araújo, Pollyanna Gonçalves, Marcos André Gonçalves, and Fabrício Benevenuto. 2016. Sentibench-a benchmark comparison of state-of-the-practice sentiment analysis methods. EPJ Data Science 5, 1 (2016), 1–29.

    Salim Sazzed. 2020. Development of sentiment lexicon in bengali utilizing corpus and cross-lingual resources. In 2020 IEEE 21st International conference on information reuse and integration for data science (IRI) . IEEE, 237–244.

    Salim Sazzed. 2021. Improving sentiment classification in low-resource bengali language utilizing cross-lingual self-supervised learning. In International Conference on Applications of Natural Language to Information Systems . Springer, 218–230.

    Salim Sazzed and Sampath Jayarathna. 2019. A sentiment classification in bengali and machine translated english corpus. In 2019 IEEE 20th international conference on information reuse and integration for data science (IRI) . IEEE, 107–114.

    Salim Sazzed and Sampath Jayarathna. 2021. Ssentia: a self-supervised sentiment analyzer for classification from unlabeled data. Machine Learning with Applications 4 (2021), 100026.

    Maite Taboada, Julian Brooke, Milan Tofiloski, Kimberly Voll, and Manfred Stede. 2011. Lexicon-based methods for sentiment analysis. Computational linguistics 37, 2 (2011), 267–307.

    Mike Thelwall, Kevan Buckley, Georgios Paltoglou, Di Cai, and Arvid Kappas. 2010. Sentiment strength detection in short informal text. Journal of the American society for information science and technology 61, 12 (2010), 2544–2558.

References

    Ana Valdivia, M Victoria Luzón, Erik Cambria, and Francisco Herrera. 2018. Consensus vote models for detecting and filtering neutrality in sentiment analysis. Information Fusion 44(2018), 126–135.

    Lei Zhang, Riddhiman Ghosh, Mohamed Dekhil, Meichun Hsu, and Bing Liu. 2011. Combining lexicon-based and learning-based methods for Twitter sentiment analysis. HP Laboratories, Technical Report HPL-2011 89 (2011).

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime