Learning to Adapt Domain Shifts of Moral Values via Instance Weighting
Xiaolei Huang, Alexandra Wormley, Adam Cohen
Published in: HT ’22: Proceedings of the 33rd ACM Conference on Hypertext and Social Media · DOI: 10.1145/3511095.3531269
License: © 2022 Copyright held by the owner/author(s). Publication rights licensed to ACM.
ACM source attribution: Complete text and visual/table renderings were converted directly from the authorized ACM HyperText 2022 proceedings PDF.
---
<!-- PDF page 1 -->
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting
Xiaolei Huang∗
University of Memphis Memphis, Tennessee, United States xiaolei.huang@memphis.edu
ABSTRACT
Classifying moral values in user-generated text from social media is critical in understanding community cultures and interpreting user behaviors of social movements. Moral values and language usage can change across the social movements; however, text classifiers are usually trained in source domains of existing social movements and tested in target domains of new social issues without considering the variations. In this study, we examine domain shifts of moral values and language usage, quantify the effects of domain shifts on the morality classification task, and propose a neural adaptation framework via instance weighting to improve cross-domain classification tasks. The quantification analysis suggests a strong correlation between morality shifts, language usage, and classification performance. We evaluate the neural adaptation framework on a public Twitter data across 7 social movements and gain classification improvements up to 12.1%. Finally, we release a new data of the COVID-19 vaccine labeled with moral values and evaluate our approach on the new target domain. For the case study of the COVID-19 vaccine, our adaptation framework achieves up to 5.26% improvements over neural baselines. This is the first study to quantify impacts of moral shifts, propose adaptive framework to model the shifts, and conduct a case study to model COVID-19 vaccine-related behaviors from moral values.
CCS CONCEPTS
• Computing methodologies →Natural language processing; • Social and professional topics →Cultural characteristics.
KEYWORDS
morality, moral values, classification, domain variation, adaptation, instant weighting
ACM Reference Format: Xiaolei Huang, Alexandra Wormley, and Adam Cohen. 2022. Learning to Adapt Domain Shifts of Moral Values via Instance Weighting. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT ’22), June
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. HT ’22, June 28-July 1, 2022, Barcelona, Spain © 2022 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-9233-4/22/06...$15.00 https://doi.org/10.1145/3511095.3531269
121
Alexandra Wormley Adam Cohen awormley@asu.edu adam.cohen@asu.edu Arizona State University Tempe, Arizona, United States
28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA, 11 pages. https: //doi.org/10.1145/3511095.3531269
1 INTRODUCTION
Moral values defined by the Moral Foundations Theory (MFT) [16, 17] provide a unique perspective to interpret human behaviors in various social movements (e.g., gun control [6, 40]) by connecting cultures and reflecting human ideology. The theory summarizes moral ideologies with five opposing pairs (virtues and vices) of foundations, including authority/subversion, care/harm, fairness/cheating, loyalty/betrayal, and purity/degradation (Table 1). Language usage in user-generated texts can reflect emotions, moral values, and cultures that explain behavior motivations and shape people’s decision-making. Emotional intuitions and moral values, known as “foundation” in the MFT, can influence people’s judgments and decision-making. Studies are leveraging online usergenerated texts (e.g., tweets) at scale by building morality classifiers to examine online users’ reactions and behaviors across different social movements, such as political stances [22, 36, 37], social protest sentiments [15, 30, 39, 50], and natural disaster attitudes [10, 12, 49]. However, limited classifiers have explicitly considered the shifts of moral values across social movements. Shifts in language use and morality across social movements can impact the applications of well-trained morality classifiers. Studies show that moral values vary across domains of social issues, such as variations between #ALLLivesMatter and #BlackLivesMatter [38], political polarization in immigration and gun control [40], and moral differences between liberals and conservatives [36, 45]. The domain is implicitly embedded in the classification process that may impact classifiers to assess user behaviors and ideologies: classifiers are often built to be applied to a new target domain that doesn’t yet exist, and performance on held-out data is measured to estimate performance on the new domain whose distributions of moral values and language use may vary from existing domains of training data. It is also hard to annotate moral values of a new domain corpus in the field of computational social sciences, which requires extensive human labor and professional training procedures [18, 22]. For example, a study [22] shows that demographics and political biases of annotators yield skewed annotations towards care/harm or fairness/cheating using the Amazon Mechanical Turk, which is a widely-used crowdsourcing platform to obtain data annotations. The domain shift is a challenge to build robust classifiers when the domain of training data is different from the domain of testing data and obtaining high-quality annotations is hard.
<!-- PDF page 2 -->
HT ’22, June 28-July 1, 2022, Barcelona, Spain Xiaolei Huang, Alexandra Wormley, and Adam Cohen
Moral Foundation Description Virtue Vice
authority subversion The values underlies virtues of leadership and followership, deference to legitimate authority and respect for traditions
care harm This foundation describes ability to feel (and dislikes) pain of others. It underlies kindness, gentleness, and nurturance.
fairness cheating The values generate ideas of justice, rights, and autonomy.
loyalty betrayal The foundation underlies virtues of patriotism and self-sacrifice for the group.
purity degradation The foundation underlies 1) an idea that the body is a temple that can be desecrated by immoral activities and contaminants, and 2) religious notions of striving to live in an elevated, less carnal, more noble way.
Table 1: Principles and descriptions of moral foundations.
In this study, we focus on the morality classification task to provide insights into how performance varies across social movement domains and propose a neural adaptation framework that applies instance weighting [21] to model the domain shift. The instance weighting is a domain adaptation approach that re-weights the source-domain training instances to approximate the targetdomain distribution. We examine the changes on a publicly available data [18], which contains tweets with moral value labels for 7 social issues (All Lives Matter, Baltimore Protest, Black Lives Matter, Hate Speech, US Presidential Election, MeToo movement, and Natural Disaster). Our analysis considers domain shifts of language usage and morality annotations. The study addresses the following research questions: 1) To what extent do the language use and morality vary across the social issues, 2) In what ways does the shift influence text classification across domains, and 3) Can morality classifiers be adapted to perform better in a new target domain. To address question 1, we compare cross-domain distributions of language use and moral values and statistically examine relations between language and morality shifts. To address question 2, we conduct cross-domain classification experiments that train a classifier on existing source domains and evaluate the classifier on a new target domain. Our analysis statistically quantifies the effects of domain shifts on classification performance by a two-tailed ttest. To address question 3, we apply a standard instance weighting approach treating social issues as domains and adapting classifiers trained on existing source domains on a new target domain. We show that this approach can significantly improve classification performance on the public dataset, even on a new social issue, the COVID-19 vaccine.
2 MORALITY DATA
In this study, we retrieved the publicly anonymized data [18] that have morality annotations for 35,108 tweets and 11 moral annotation types. The tweets come from seven discourse domains that focus on major US social movements in the past few years: 1) ALM, tweets related to All Lives Matter movement that typically contains with conservative views and oppose to the ideas of the Black Lives Matter movement; 2) Baltimore, tweets related to the Baltimore protests against police brutality leading to the death of Freddie Gray; 3) BLM, tweets related to Black Lives Matter that fights against discrimination and inequality experienced by African
122
American people; 4) Davidson, tweets of hate speech and offensive language collected by [9]; 5) Election, tweets posted during the 2016 U.S. presidential election; 6) MeToo, tweets related to the #MeToo movement against sexual abuse and harassment; 7) Sandy, tweets related to the Hurricane Sandy, which was a natural disaster happened in 2012. We replace any user mentions and URLs in the tweet documents with “USER” and “URL” to preserve user privacy. We then lowercased all tweet documents and tokenized each document by NLTK [4]. We removed tweet documents if they had fewer than 3 tokens. The annotations have one no-moral and the other 10 moral annotations defined by the Moral Foundations Theory [16, 17]. The label set proposes five bipolar factor pairs (virtues / vices): care/harm, fairness/cheating, loyalty/betrayal, authority/subversion, purity/degradation. We present descriptions of the moral foundations in Table 1. The no-moral label indicates a tweet is nonmoral if the tweet does not fall under any foundations. At least three annotators annotated each tweet document. To ensure annotation quality, we followed the previous study [18] and used the majority vote to keep any labels with votes by at least two annotators. We dropped documents with empty annotations after applying the majority vote. If most annotators label a tweet with the no-moral label, we mark the tweet as no-moral and drop other labels. We summarize the data statistics in Table 2. The documents are short, with 17 tokens per document on average. However, we can find that skewed distributions exist in labels that the 11 document labels are imbalanced. For example, the percentage of no-moral ranges from 13.38% in the Sandy domain to 92.16% in the Davidson domain. The distributions of moral values are also different across the domains. For example, the top 3 moral values of Election are cheating, harm, and fairness, but subversion, degradation, and cheating are the 3 top moral values for MeToo. To better understand moral variations across domains, we extract percentages of moral values for each domain and group the values using the Moral Foundations Theory into virtue and vice moral values. We group the social events according to their virtue-vice ratio and summarize the distributions of moral values in Table 3. Significant distribution variations of moral labels exist across 7 domains in Table 3. For example, the three social protests (BLM, MeToo, and Baltimore) and the hate speech (Davidson) have shown very different moral value distributions than the other domains, such as presidential election (Election) and natural disaster (Sandy). The ratio comparison between virtue-related and vice-related moral
<!-- PDF page 3 -->
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting HT ’22, June 28-July 1, 2022, Barcelona, Spain
Domain Doc Token authority betrayal care cheating degradation fairness harm loyalty purity subversion no-moral
ALM 3,389 16.76 6.37 1.04 12.07 13.22 3.16 13.51 19.03 6.29 2.14 2.35 20.82 Baltimore 4,815 15.53 0.42 11.66 3.39 9.81 0.57 2.70 5.09 7.13 0.73 5.57 52.95 BLM 4,980 15.41 5.31 2.81 6.09 14.06 4.32 8.49 19.74 7.73 2.84 5.91 22.7 Davidson 4,641 15.64 0.41 0.82 0.19 1.26 1.36 0.08 2.74 0.78 0.08 0.12 92.16 Election 4,994 17.87 2.65 2.00 6.27 9.53 2.09 8.79 9.19 3.28 6.40 2.56 47.23 MeToo 4,313 23.50 7.17 5.29 3.53 10.93 14.83 6.55 7.07 5.42 3.08 14.89 21.23 Sandy 3,847 14.49 9.60 3.11 21.41 9.9 1.93 3.74 16.83 8.93 1.45 9.73 13.38
Overall 30,979 17.02 4.45 4.02 6.96 10.01 4.34 6.23 11.25 5.68 2.56 6.16 38.34
Table 2: Statistics summary of the multi-class morality dataset [18]. We list number of total document (Doc), average number of tokens per tweet document (Token), and label distributions of 11 annotations in percentage (%).
Domain Virtue (%) Vice (%) Virtue-Vice Ratio authority care fairness loyalty purity betrayal cheating degradation harm subversion
Sandy 11.08 24.72 4.31 10.30 1.68 3.58 11.43 2.23 19.43 11.23 1.09 Election 5.03 11.88 16.67 6.21 12.12 3.79 18.06 3.97 17.42 4.85 1.08 ALM 8.04 15.24 17.07 7.94 2.70 1.32 16.69 3.99 24.03 2.97 1.04
BLM 6.87 7.88 10.99 10.01 3.67 3.63 18.19 5.58 25.54 7.64 0.65 MeToo 9.10 4.48 8.31 6.88 3.91 6.72 13.88 18.83 8.98 18.91 0.49 Baltimore 0.88 7.21 5.74 15.15 1.54 24.78 20.85 1.21 10.81 11.84 0.44 Davidson 5.25 2.36 1.05 9.97 1.05 10.50 16.01 17.32 34.91 1.57 0.25
Overall 7.22 11.28 10.10 9.21 4.15 6.52 16.23 7.04 18.25 9.99 0.72
Table 3: Distributions of moral values across the 7 social domains. We group moral values into two groups, virtue and vice. The virtue-vice ratio is a divide operation between virtue-related and vice-related moral value counts. Overall, the data shows higher ratio of the vice-related moral values.
values can reflect that the social movements have higher vicerelated moralities in their language. Studies have shown that group identities relating to their cultural and social backgrounds can lead to moral variations [47]. For example, ALM and BLM show various moral value distributions in virtue-related and vice-related moralities; degradation and subversion are the top 2 moral values in the MeToo domain, while cheating and harm are the top 2 moral values in the Election domain. Our finding generalizes and aligns with a recent study [38] that moral values in ALM- and BLM-related social events are different. The moral variations motivate us to characterize effects and examine the impacts of moral variations. We conducted the following analyses, moral value shifts and moral shift impacts, to understand two important questions to quantify the effects and impacts:
• To what extent do the moral values shift across domains? • To what extent can the moral shift impact morality modeling?
2.1 Analysis 1: Moral Value Shifts
To quantify moral shifts, we measure cosine similarities of moral values and language usage. Language usage is an important indicator to reflect and identify moral values [15, 22, 37, 41]. We extract linguistic patterns of language usage by a Latent Dirichlet Allocation (LDA) [5], which abstracts language usage into thematic vectors. We train a LdaModel from GenSim [35] with 200 topics and default parameters over all available corpus. The model learns a multinomial topic distribution P(Z |D) from a Dirichlet prior, where
123
Z refers to each topic, and D refers to each document. After training the topic model, we can encode each document d with a probability distribution over the 200 topics. For each domain (e.g., BLM or MeToo), we calculate the average topic distribution across the documents from that domain. Finally, we compute cosine similarities of the topic distributions between every two domains. To measure morality similarities across domains, we reuse the moral label distributions (Table 3) and calculate cosine similarities between every two domains of the moral labels. Figure 1a and 1b illustrate the cosine similarities across domains for morality (11 values) and language usage (topic) respectively. We can find that similarities are low and varying across domains. The low similarities indicate both language and moral values shift significantly across domains. For example, similarities between ALM and BLM in moral values and language usage are 0.021 and 0.019, respectively. The Davidson has comparatively high similarities from other domains. We infer this as the domain includes diverse topics related to hate speech and offensive language. We conduct a significance test to verify if the language and moral values variations correlate. The significance test requires the normality test to decide whether to use parametric or nonparametric methods. We conduct a normality test via normaltest from statsmodel [43]. The test yields p-values as 0.00873 and 3.318e-05 for language and moral value variations respectively, which demonstrates that the significance test will use non-parametric methods. We use the spearmanr from statsmodel to test our null hypothesis that the moral variations have no significant correlation
<!-- PDF page 4 -->
HT ’22, June 28-July 1, 2022, Barcelona, Spain Xiaolei Huang, Alexandra Wormley, and Adam Cohen
1 0.2 0.021 0.41 0.18 0.024 0.014
ALM
0.2 1 0.16 0.047 0.004 0.17 0.18
Baltimore
0.021 0.16 1 0.36 0.15 0.012 0.016
BLM
0.41 0.047 0.36 1 0.06 0.37 0.38
Davidson
0.18 0.004 0.15 0.06 1 0.15 0.16
Election
0.024 0.17 0.012 0.37 0.15 1 0.015
MeToo
0.014 0.18 0.016 0.38 0.16 0.015 1
Sandy
ALM Baltimore BLM Davidson Election MeToo Sandy
(a) Similarities of moral values across domains.
1 0.3 0.019 0.47 0.053 0.038 0.25
ALM
0.3 1 0.29 0.22 0.12 0.42 0.26
Baltimore
0.019 0.29 1 0.48 0.065 0.059 0.31
BLM
0.47 0.22 0.48 1 0.28 0.51 0.38
Davidson
0.053 0.12 0.065 0.28 1 0.11 0.19
Election
0.038 0.42 0.059 0.51 0.11 1 0.36
MeToo
0.25 0.26 0.31 0.38 0.19 0.36 1
Sandy
ALM Baltimore BLM Davidson Election MeToo Sandy
(b) Similarities of topic distributions across domains.
Figure 1: Analyses to quantify effects of moral shifts (Section 2.1). Higher scores indicate domains are more similar.
with the topic variations. Finally, the results show p-value=0.488 and coefficient=0.165 for the test between domain variations of the label and linguistic expression. Therefore, we can not reject the null hypothesis of the significant test. The observations suggest that: while we can observe variations exist in both moral labels and language use, the two variations across domains do have a strong correlation.
2.2 Analysis 2: Moral Shift Impacts
While language use and moral values vary across social event domains, it is a separate question whether those variations will affect morality classification. Prior work [18] has shown that classification performance varies across the social event domains when training and testing sets are from the same domain. However, it is unknown how the domain variations impact the classification of moral values when the testing set is from a new domain. To achieve this, we conduct a morality classification evaluation under the cross-domain setting that trains a classifier on one domain and tests the classifier on the other domains. We split 80% of documents as the training set and hold out 20% of documents as the testing set for each domain corpus. We extract TF-IDF weighted uni-, bi-, and tri-gram features on each domain corpus with the most frequent 15K features during the training. We then build a logistic regression classifier using
124
.676
.591
0.6
0.4
.347
.299
.233
0.2
Baltimore Election BLM ALM Sandy
(a) In-domain (Baltimore in light blue) vs. Cross-domain (others in grey) classification performance.
.66
.621
0.6
.412
.383
0.4
.246
0.2
Election Baltimore BLM ALM Sandy
(b) In-domain (Election in light blue) vs. Cross-domain (others in grey) classification performance.
Figure 2: Analyses to quantify impacts of moral shifts (Section 2.2) . Higher scores indicate better performance.
LogisticRegression from scikit-learn [33] with default parameters. Finally, we evaluate the classifier across each domain’s test set using the F1 score. We visualize cross-domain classification performance in Figure 2 with more details in the appendix’s Figure 5. In general, we can observe that in-domain classification evaluations outperform out-domain evaluations. For example, the in-domain evaluation on the Baltimore achieves 67.6% versus the out-domain evaluation on the Election, 59.1%. The finding applies to the other cross-domain evaluations. The observation indicates domain variations in moral values and language usage can impact morality classifier performance when training and testing sets of classifiers are from different social event domains. To qualitatively examine the domain
<!-- PDF page 5 -->
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting HT ’22, June 28-July 1, 2022, Barcelona, Spain
shifts of language use, we extract top word features with regard to morality labels for each domain and present a qualitative study in Appendix A.2. To statistically quantify the impacts of domain variations on cross-domain classification, we conduct a two-tailed t-test on the domain shifts. The statistical test examines the domain and classification performance variations to test our null hypothesis that the domain shifts have no significant relation with the classification performance variations. We first obtain performance variations by subtracting the in-domain score from the out-domain score. For example, to get the performance variation of Baltimore (in-domain) and Election (out-domain), we subtract the in-domain score (.676) by the out-domain score (.591). To decide if the significant test uses parametric or non-parametric methods, we conduct a normality test via scipy.stats.normaltest from statsmodel [43]. The normality test yields p-value=0.533 for the domain variations and p-value=0.409 for the performance variations, which indicates that we will use parametric methods. We then fit the two variables to a linear model (Yperf ormance = f (Xdomain) + e) that predicts performance variations by the domain variations (topic + label variations). Finally, the test results show p-value=1.21e −10 and coefficient=1.010 for the correlation between the domain and classification performance variations.1 Therefore, we reject the null hypothesis of the significant test. The observations suggest that domain variations may cause significant variations for cross-domain classification performance.
3 LEARNING TO ADAPT FRAMEWORK (L2AF)
We propose a neural adaptation framework (L2AF in Figure 3) that applies instance weighting [21] to encounter the domain variations for the morality classification task when training and testing sets are from different domains. The framework adapts domain variations of language usage and moral values through four main modules: neural feature extractor, prediction network, weighting network, and joint optimization.
3.1 Neural Feature Extractor
Our framework treats neural models as feature extractors. The feature extractor encodes input documents into feature representations, X. In this study, we experiment with two types of neural models, recurrent neural network (RNN) [7] and BERT [11], which have achieved promising performance in morality classification [18, 22]. The RNN is a bidirectional network with GRU [7] units. We then feed the document representations X for two joint prediction tasks, prediction and weighting networks.
3.2 Prediction Network
The prediction network takes the document representations X to predict moral values (fθ (c|X)) via a full-connected network, where θ is the network parameters and c is the moral categories. The network applies a dropout [44] on the feature vectors X and employs a softmax function to predict the 11 moral labels. We apply the cross-entropy to calculate morality prediction loss and optimize
1p-value=3.9e −5 for the topic variation, and p-value=3.43e −3 for the moral label variation; coefficient=1.012 for the topic and coefficient=1.008 for the moral variation.
125
the prediction network, L(yc, fθ (c|X)), where L is the loss and yc is the ground truth of moral values.
3.3 Weighting Network
The L2AF deploys a weighting network that dynamically adapts the domain shifts of language and moral values. Based on our observations (Section 2.1), we assume that two domains sharing similar language usage will also share similar moral values. And therefore, it is reasonable to assign higher weights to out-domain training instances sharing similar language usage with in-domain data and lower weights to out-domain training instances with different language usage from in-domain data. The weighting network yields training weights (w) to leverage shifts between in-domain and out-domain data for the prediction network as w · L(yc, fθ (c|X)). We obtain the weights by predicting in-domain probabilities of tweet documents, fϕ(d|X), where d is domain label and ϕ refers to weight network parameters. The fϕ(d|X) is a binary domain classification task that predicts the domains (in- vs. out-domain) of the tweet documents, where the f is a sigmoid function to calculate in-domain probabilities, instance weights (w). For example, if our target (test) domain is ALM (in-domain), we will treat the other six domains (e.g., BLM and Election) as out-domain. We deploy a binary cross-entropy to optimize the weighting network.
3.4 Joint Optimization
L = argmin γ,ϕ,θ α · L(yd, fϕ(d|X)) + w · L(yc, fθ (c|X)) (1)
We train the L2AF via two optimization tasks (Equation 1), moral value and domain predictions, where γ refers to model parameters of the neural extractor, and α is a coefficient factor to leverage the importance of the domain prediction task. We first converge the weighting network during the training phase to obtain instance weights. We then train the prediction network with the instance weights. As the joint optimization continuously updates the neural extractor and weighting network parameters, the weighting network will dynamically adjust its predicted weights. The dynamic procedure will ensure the framework adjusts and balance multi-domain data accordingly. The framework trains the weighting network with Adam optimizer [24] and uses a separate optimizer for the prediction network. If the neural extractor uses RNN, we optimize the prediction network with the RMSprop [46]. Otherwise, we optimize the prediction network with the AdamW [28] for the BERT-based extractor. Separate optimizers for the joint prediction tasks can give us more controls on the model training process. We only use the prediction network during the test phase to obtain moral value predictions.
4 EXPERIMENTS
We focus on a multi-domain adaptation scenario in which morality classifiers trained on existing domains often adapt to a new domain quickly. The goal is to evaluate the ability of a classifier to adapt from several observed domains to a new unobserved domain, which can be especially common for examining social movements by classifiers [30, 36, 45]. We use the annotated morality data (Table 2) that contains 30,979 tweets from 7 domains. For each target domain, we randomly split 80% of its domain data as a test set and the rest as a
<!-- PDF page 6 -->
HT ’22, June 28-July 1, 2022, Barcelona, Spain Xiaolei Huang, Alexandra Wormley, and Adam Cohen
Backpropagation
𝑋
Neural Extractor
Tweets
𝑋
Backpropagation
Prediction Network
𝒘· 𝑳(𝒚𝒄, 𝒇𝜽𝒄𝑿)
Backpropagation
⊙
Morality
Prediction
𝑓𝜃(𝑐|𝑋)
Loss
𝑤
Weighting Network
𝑳(𝒚𝒅, 𝒇𝝓𝒅𝑿)
Domain
Prediction
𝑓𝜙(𝑑|𝑋)
Loss Backpropagation
Figure 3: Illustration of learning to adapt framework via instance weighting. We use blue and orange to represent optimization workflows of prediction and weighting networks respectively.
validation set for tuning model parameters to simulate limited annotated data from a new domain. Under the multi-domain scenario, we train models on the 6 existing source domains, validate and tune model parameters on the validation set, and evaluate classification performance on the test set of the held-out target domain. We report model performance by F1 score, which is a standard measurement in existing studies of morality classification [18, 37, 42].
4.1 Model Settings
We compare our neural framework (denote as Adapt) with two baseline types, In-Domain and No-adapt. The In-Domain approach trains, validates, and tests classification models on in-domain (target domain) data. The No-adapt shares the same modules except for using the weighting network. We train the No-adapt approach by the same data splits with our proposed framework that trains classification models on out-domain (source domain) data, tune model parameters on the in-domain validation set, and evaluate models on the in-domain test set. The No-adapt approach follows the existing study [18] that predicts labels for multiple outcomes in a multitasking manner. We experiment with two types of neural classifiers, RNN and BERT. For the base models (RNN and BERT), we utilize the pre-trained embeddings [11, 34] to initialize model parameters. The RNN and BERT baselines follow model architectures of existing studies, [18] and [29], respectively, which have achieved state-of-art performance results in morality classification. We trained the models on an Nvidia 3090 GPU and evaluated the model on CPUs. Our experiments set the training batch as 16, the number of training epochs as 30, the max document as 60, and the dropout rate as 0.2. We tune the learning rate in a range of [1e-6, 1e-4] on the validation set to get the best performance.
4.2 Results
We report classification performance results in Table 4. Our learning to adapt framework leads to performance improvements (from 2.276% to 12.128%) over the comparable baselines across 7 domains. The adaptation approach outperforms the No-adapt approach by a
126
large margin in both RNN and BERT models. The improvements indicate that adapting domain shifts can improve morality classifiers and demonstrate that our approach can effectively adapt shifts in moral values and language usage across social events. We find that our adaptation approach achieves minor improvements on the Davidson domain, and we infer this as the domain is more similar to other domains in moral values (Figure 1a) and language usage (Figure 1b). Training morality classifiers on multiple out-domain data can outperform the models trained on in-domain data. Compared to the single-cross-domain evaluation (Figure 5), classifiers trained multi-domain data can perform better than in-domain-only data by leveraging common patterns across multiple domains.
5 CASE STUDY: MORALITY OF COVID-19 VACCINE
In this section, we present a case study on the COVID-19 vaccine and examine if our L2AF can effectively adapt morality classifiers trained on existing domains to a new target domain. To verify our approach, we annotate Twitter data with morality labels following the definitions and annotation standards of the previous studies [16– 18]. We randomly sampled 500 vaccine-related tweets from a public repository [20] that has retrieved COVID-19 data using the Twitter streaming API since 2020. We only keep English tweets and tweet authors from the US to reduce cultural and multilingual issues. Two professional social psychologists (the second and third authors) participated in the data annotation process. The annotators have demonstrated extensive expertise and published journal articles related to culture and morality [8, 25, 31]. We utilized a double annotation strategy [13]. The first annotator reviewed each tweet and labeled the document with the 11 categories (10 moral values and one no-moral label). The second annotator reviewed the annotations to ensure label qualities. The two annotator setting fits this pilot study and reduces annotation biases beyond a single domain expert. We follow the same steps in Section 2 to preprocess the tweet documents and present the data summary in Table 5.
<!-- PDF page 7 -->
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting HT ’22, June 28-July 1, 2022, Barcelona, Spain
Base Model Type Target Domain
RNN
BERT
∆↑(%) 3.879 4.811 11.001 2.276 12.128 3.274 10.058
ALM Baltimore BLM Davidson Election MeToo Sandy
In-Domain 0.700 0.633 0.705 0.918 0.551 0.549 0.518
No-adapt 0.716 0.674 0.711 0.932 0.653 0.553 0.526
Adapt 0.746 0.706 0.816 0.955 0.712 0.578 0.570
In-Domain 0.748 0.71 0.783 0.936 0.683 0.559 0.564
No-adapt 0.776 0.727 0.828 0.949 0.735 0.630 0.649
Adapt 0.781 0.732 0.864 0.955 0.758 0.655 0.672
Table 4: Performance (F1 score) comparisons between three approaches across the domains of seven social movements. The ∆↑refers to average improvements of our approach over baselines.
Domain Doc Token Virtue Vice Virtue-Vice Ratio authority care fairness loyalty purity betrayal cheating degradation harm subversion no-moral
Vaccine 500 18.59 7.88 3.11 0.37 5.86 0.55 3.30 2.93 6.78 9.16 3.85 56.23 0.683
Table 5: Data and annotation summary of COVID-19 vaccine.
model = RNN
0.85
0.820
0.80
0.779
0.743
0.75
score
0.70
0.65
0.60
Indomain No-adapt Adapt
type
Figure 4: Performance (F1 score) comparison on the COVID-19 vaccine data.
We can find that the morality distribution is different from the existing 7 domains of social movements. The table shows that: 1) authority and loyalty are the top two for virtue-related morality, 2) and degradation and harm are the top two for vice-related morality. Such a finding aligns with a previous study [1] on parental vaccine hesitancy that traditional communications and interventions focusing on harm and fairness moralities may not fit for the new moral domain of vaccine hesitancy. According to the Moral Foundation Theory [16], the corresponding moral values of authority, loyalty, harm, and degradation are subversion, betrayal, care, and purity. The distribution differs from the previous study [1] on morality and vaccine hesitancy showing that moral values can also shift across the vaccine domain. The domain shift of moral values highlights the necessity of developing generalizable and robust morality classifiers for understanding the vaccine challenge.
127
model = BERT
0.805
0.790
0.626
Indomain No-adapt Adapt
type
We follow the same experimental settings as Section 4.1 and evaluate the three approach types on the COVID-19 vaccine data. We present F1 scores for the two base models in Figure 4. Our approach (Adapt) consistently outperforms the other two baselines across the two neural encoders. The performance improvement demonstrates the effectiveness of our approach on adapting classifiers from existing source domains to a new target domain. We also observe that the RNN-based classifiers generally outperform the BERT-based classifiers. We infer this as the data annotation size that the BERT-based classifiers need more annotated data to achieve a similar performance than the RNN-based classifiers. The adaptable and generalizable capabilities of classification models can be essential for studying the morality of new social movements when annotated data is rare or hard to obtain.
<!-- PDF page 8 -->
HT ’22, June 28-July 1, 2022, Barcelona, Spain Xiaolei Huang, Alexandra Wormley, and Adam Cohen
6 ETHIC AND PRIVACY CONCERNS
In this study, we only use the tweet documents and morality labels for evaluation purposes without any other user profile, such as user IDs. All experimental information has been anonymized before training text classifiers. Specifically, we used anonymized tweet IDs and replaced any user mentions and URLs with two generic symbols, “USER” and “URL”, respectively. To preserve user privacy, we will follow Twitter’s privacy policy to release our annotated data of COVID-19 vaccine and provide instructions to access the public data [18] in our experiments. We provided preprocessed datasets with the professional annotators to obtain high-quality annotations. We report all preprocessing steps, hyperparameter settings, and other technical details of analysis procedures and will release our code repository to allow for replications. Our quantitative and qualitative analyses rely on the machine learning toolkits, including topic model [35], scikit-learn [33], and statsmodel [43]. Implementations of our approach and the baselines rely on the deep learning toolkits, including PyTorch [32] and Transformers [48]. Therefore the result report does not represent the authors’ personal views. All our observations derive from the reported results in this paper, which will be reproducible using our code. Under Institutional Review Boards (IRB) guidelines, we believe that there is no code of ethics violation throughout the experiments and annotations.
7 RELATED WORK
Emerging research studies have adopted the Moral Foundation Theory (MFT) to study moral values of users’ behaviors and opinions among social issues, such as abortion [37, 39, 41, 42], immigration [29, 40, 41], climate change [6, 37, 39], and natural disaster [15, 27]. Neural models (e.g., RNN and BERT) have dominated the classification performance of multiple natural language processing tasks [11], text generation [14], and question answering [26]. Recent efforts have expanded the neural-based models to categorize moral values of online user-generated documents [18, 27, 29]. However, studies have found moral values can vary across domains of social issues [36, 38] that may impact classification performance. Current methods focus on augmenting document features for morality classifiers, such as incorporating background knowledge from Wikipedia [27], extending moral lexicon by crowdsourcing [19], and vectorizing moral lexicons [2] to diversify token representations. While the existing approaches focus on document feature augmentation, limited studies focus on the model level to explicitly incorporate the domain shifts into morality classifiers and adapt the classifiers to a new target domain. Our work applies the instance weighting method [3, 21] and proposes a neural adaptation framework to augment the cross-domain performance of morality classifiers. While previous study [23] recruited participants by survey collections and identified moral values on vaccine attitudes from their liked Facebook pages, there is no prior study that examines moral values on COVID-19 from social media texts and develops classification models. This study is the first work examine moral values in COVID-19 related tweets and has a great potential to probe hesitancy of COVID-19 vaccine from the social psychology aspects.
128
8 CONCLUSION
In this paper, we have examined domain shifts of language use and moral values across social issues, quantified the effects of the domain shifts on classification models, and proposed a neural adaptation framework to augment morality classifiers. Differing from existing work of focusing on variations of moral values [36, 38], our study investigates both language use and morality shifts and analyzes connections across language use, morality, and classification performance. Our approach dynamically adjusts weights on training instances from source domains regarding the new target domain and demonstrates its effectiveness on public data [18] and our annotated COVID-19 vaccine data. To our best knowledge, the case study is the first pilot work investigating the moral values of the COVID-19 vaccine. In our future work, we intend to extend the current morality annotations on the COVID-19 vaccine and explore shifts in other settings, such as geolocation and temporality. Our code and data instructions will be available at https://github.com/xiaoleihuang/MoralCausality.
8.1 Limitations
While we have quantified of moral variations across social movements and proposed a neural adaptive framework to model the variations, we must acknowledge several limitations to appropriately interpret our findings. First, we extracted language use patterns in Section 2.1 via topic model, which are not accurate enough for short texts. Using the derived topic distributions may contribute to the correlation test failure between language use and label variations. Our future work will derive more robust feature representations for language use. Second, the moral labels of the COVID-19 data by social psychologists may also bring annotation biases. During the labeling process, one domain expert annotated the data and the second expert checked the annotations. While annotation studies usually recruit at least three non-expert annotators by crowdsourcing methods (e.g., Amazon MTurk), recruiting multiple domain experts can be a challenge. For example, the moral data [18] had multiple undergraduate students to assign moral labels to each tweet, which aims to reduce non-expert noisy and bias. Combining experts and crowdsourcing annotators to build moral values of COVID-19 will be our future work.
ACKNOWLEDGMENTS
The authors want to thank for reviewers’ valuable comments.
REFERENCES
[1] Avnika B. Amin, Robert A. Bednarczyk, Cara E. Ray, Kala J. Melchiori, Jesse Graham, Jeffrey R. Huntsinger, and Saad B. Omer. 2017. Association of moral values with vaccine hesitancy. Nature Human Behaviour 1, 12 (dec 2017), 873–880. https://doi.org/10.1038/s41562-017-0256-5
[2] Oscar Araque, Lorenzo Gatti, and Kyriaki Kalimeri. 2020. MoralStrength: Exploiting a moral lexicon and embedding similarity for moral foundations prediction. Knowledge-Based Systems 191 (mar 2020), 105184. https://doi.org/10.1016/j. knosys.2019.105184
[3] Himanshu Sharad Bhatt, Manjira Sinha, and Shourya Roy. 2016. Cross-domain Text Classification with Multiple Domains and Disparate Label Sets. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Berlin, Germany, 1641–1650. https://doi.org/10.18653/v1/P16-1155
[4] Steven Bird and Edward Loper. 2004. NLTK: The Natural Language Toolkit. In Proceedings of the ACL Interactive Poster and Demonstration Sessions. Association
<!-- PDF page 9 -->
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting HT ’22, June 28-July 1, 2022, Barcelona, Spain
for Computational Linguistics, Barcelona, Spain, 214–217. https://aclanthology. org/P04-3031
[5] David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of Machine Learning Research 3, Jan (2003), 993–1022. http://www.jmlr. org/papers/volume3/blei03a/blei03a.pdf
[6] William J. Brady, Julian A. Wills, John T. Jost, Joshua A. Tucker, and Jay J. Van Bavel. 2017. Emotion shapes the diffusion of moralized content in social networks. Proceedings of the National Academy of Sciences 114, 28 (jul 2017), 7313–7318. https://doi.org/10.1073/pnas.1618923114
[7] Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio.
On the Properties of Neural Machine Translation: Encoder–Decoder Approaches. In Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and
Structure in Statistical Translation. Association for Computational Linguistics, Doha, Qatar, 103–111. https://doi.org/10.3115/v1/W14-4012
[8] Adam B. Cohen and Jordan W. Moon. 2017. Psychology: Atheism and moral intuitions. Nature Human Behaviour 1, 8 (aug 2017), 0157. https://doi.org/10. 1038/s41562-017-0157
[9] Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated Hate Speech Detection and the Problem of Offensive Language. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 11. AAAI, Montreal, Canada, 512–515. https://ojs.aaai.org/index.php/ICWSM/article/ view/14955
[10] Sharon Dawson and Graham Tyson. 2012. Will Morality or Political Ideology Determine Attitudes to Climate Change? Australian Community Psychologist: The official journal of the APS College Of Community Psychologists 24, 2 (nov 2012), 8–25.
[11] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186. https://doi.org/10.18653/v1/N19-1423
[12] Janis L. Dickinson, Poppy McLeod, Robert Bloomfield, and Shorna Allred. 2016. Which Moral Foundations Predict Willingness to Make Lifestyle Changes to Avert Climate Change in the USA? PLOS ONE 11, 10 (10 2016), 1–11. https: //doi.org/10.1371/journal.pone.0163852
[13] Dmitriy Dligach and Martha Palmer. 2011. Reducing the Need for Double Annotation. In Proceedings of the 5th Linguistic Annotation Workshop. Association for Computational Linguistics, Portland, Oregon, USA, 65–73. https: //aclanthology.org/W11-0408
[14] Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi.
Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their
Consequences. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 698–718. https://doi.org/10.18653/v1/2021. emnlp-main.54
[15] Justin Garten, Reihane Boghrati, Joe Hoover, Kate M Johnson, and Morteza Dehghani. 2016. Morality between the lines: Detecting moral sentiment in text. In Proceedings of IJCAI 2016 workshop on Computational Modeling of Attitudes. IJCAI, New York, US, 6 pages. https://www.morteza-dehghani.net/wp-content/ uploads/morality-lines-detecting.pdf
[16] Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P. Wojcik, and Peter H. Ditto. 2013. Chapter Two - Moral Foundations Theory: The Pragmatic Validity of Moral Pluralism. In Advances in Experimental Social Psychology, Patricia Devine and Ashby Plant (Eds.). Vol. 47. Academic Press, Cambridge, MA, 55–130. https://doi.org/10.1016/B978-0-12-407236-7.00002-4
[17] Jonathan Haidt. 2007. The New Synthesis in Moral Psychology. Science 316, 5827 (may 2007), 998–1002. https://doi.org/10.1126/science.1137651
[18] Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, Gabriela Moreno, Christina Park, Tingyee E Chang, Jenna Chin, Christian Leong, Jun Yen Leung, Arineh Mirinjian, and Morteza Dehghani. 2020. Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment. Social Psychological and Personality Science 11, 8 (nov 2020), 1057–1071. https://doi.org/10.1177/1948550619876629
[19] Frederic R. Hopp, Jacob T. Fisher, Devin Cornell, Richard Huskey, and René Weber.
The extended Moral Foundations Dictionary (eMFD): Development and
applications of a crowd-sourced approach to extracting moral intuitions from text. Behavior Research Methods 53, 1 (feb 2021), 232–246. https://doi.org/10. 3758/s13428-020-01433-0
[20] Xiaolei Huang, Amelia Jamison, David Broniatowski, Sandra Quinn, and Mark Dredze. 2020. Coronavirus Twitter Data: A collection of COVID-19 tweets with automated annotations. Johns Hopkins University. https://doi.org/10.5281/ zenodo.5860720 http://twitterdata.covid19dataresources.org/index.
[21] Jing Jiang and ChengXiang Zhai. 2007. Instance Weighting for Domain Adaptation in NLP. In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics. Association for Computational Linguistics, Prague, Czech Republic, 264–271. https://aclanthology.org/P07-1034
129
[22] Kristen Johnson and Dan Goldwasser. 2018. Classification of Moral Foundations in Microblog Political Discourse. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Melbourne, Australia, 720–730. https://doi.org/ 10.18653/v1/P18-1067
[23] Kyriaki Kalimeri, Mariano G. Beiró, Alessandra Urbinati, Andrea Bonanomi, Alessandro Rosina, and Ciro Cattuto. 2019. Human Values and Attitudes towards Vaccination in Social Media. In Companion Proceedings of The 2019 World Wide Web Conference (San Francisco, USA) (WWW ’19). Association for Computing Machinery, New York, NY, USA, 248–254. https://doi.org/10.1145/3308560.3316489
[24] Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. In Proceedings of 3rd International Conference for Learning Representations. ICLR, San Diego, CA, 15 pages.
[25] Jung Yul Kwon, Alexandra S. Wormley, and Michael E.W. Varnum. 2021. Changing cultures, changing brains: A framework for integrating cultural neuroscience and cultural change research. Biological Psychology 162 (2021), 108087. https: //doi.org/10.1016/j.biopsycho.2021.108087
[26] Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 7871–7880. https://doi.org/10.18653/v1/2020.acl-main.703
[27] Ying Lin, Joe Hoover, Gwenyth Portillo-Wightman, Christina Park, Morteza Dehghani, and Heng Ji. 2018. Acquiring Background Knowledge to Improve Moral Value Prediction. In 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). IEEE, Barcelona, Spain, 552–559. https://doi.org/10.1109/ASONAM.2018.8508244
[28] Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In Proceedings of 7th International Conference on Learning Representations. ICLR, New Orleans, LA, 18 pages. https://openreview.net/forum?id=Bkg6RiCqY7
[29] Negar Mokhberian, Andrés Abeliuk, Patrick Cummings, and Kristina Lerman. 2020. Moral Framing and Ideological Bias of News. In Social Informatics, Samin Aref, Kalina Bontcheva, Marco Braghieri, Frank Dignum, Fosca Giannotti, Francesco Grisolia, and Dino Pedreschi (Eds.). Springer International Publishing, Cham, 206–219. https://doi.org/10.1007/978-3-030-60975-716
[30] Marlon Mooijman, Joe Hoover, Ying Lin, Heng Ji, and Morteza Dehghani. 2018. Moralization in social networks and the emergence of violence during protests. Nature Human Behaviour 2, 6 (jun 2018), 389–396. https://doi.org/10.1038/s41562018-0353-0
[31] Stefanie B. Northover, William C. Pedersen, Adam B. Cohen, and Paul W. Andrews.
Effect of artificial surveillance cues on reported moral judgment: Experimental failures to replicate and two meta-analyses. Evolution and Human Behavior
38, 5 (2017), 561–571. https://doi.org/10.1016/j.evolhumbehav.2016.12.003
[32] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Advances in Neural Information Processing Systems (NeuriPs), H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., Vancouver, Canada, 8024–
http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-highperformance-deep-learning-library.pdf
[33] Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011. Scikit-learn: Machine learning in Python. Journal of machine learning research 12, Oct (2011), 2825–2830. http://www.jmlr.org/ papers/volume12/pedregosa11a/pedregosa11a.pdf
[34] Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Doha, Qatar, 1532–1543. https://doi.org/10.3115/v1/ D14-1162
[35] Radim Rehurek and Petr Sojka. 2010. Software framework for topic modelling with large corpora. In In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. European Language Resources Association, Valletta, Malta, 5 pages. https://radimrehurek.com/gensim/lrec2010final.pdf
[36] Markus Reiter-Haas, Simone Kopeinik, and Elisabeth Lex. 2021. Studying Moralbased Differences in the Framing of Political Tweets. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 15. AAAI, Virtual Event, USA, 1085–1089. https://ojs.aaai.org/index.php/ICWSM/article/view/18135
[37] Rezvaneh Rezapour, Ly Dinh, and Jana Diesner. 2021. Incorporating the Measurement of Moral Foundations Theory into Analyzing Stances on Controversial Topics. In Proceedings of the 32nd ACM Conference on Hypertext and Social Media (Virtual Event, USA) (HT ’21). Association for Computing Machinery, New York, NY, USA, 177–188. https://doi.org/10.1145/3465336.3475112 — ACM HyperText copy
<!-- PDF page 10 -->
HT ’22, June 28-July 1, 2022, Barcelona, Spain Xiaolei Huang, Alexandra Wormley, and Adam Cohen
[38] Rezvaneh Rezapour, Priscilla Ferronato, and Jana Diesner. 2019. How Do Moral Values Differ in Tweets on Social Movements?. In Conference Companion Publication of the 2019 on Computer Supported Cooperative Work and Social Computing (Austin, TX, USA) (CSCW ’19). Association for Computing Machinery, New York, NY, USA, 347–351. https://doi.org/10.1145/3311957.3359496
[39] Rezvaneh Rezapour, Saumil H. Shah, and Jana Diesner. 2019. Enhancing the Measurement of Social Effects by Capturing Morality. In Proceedings of the Tenth Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis. Association for Computational Linguistics, Minneapolis, USA, 35–45. https://doi.org/10.18653/v1/W19-1305
[40] Shamik Roy and Dan Goldwasser. 2021. Analysis of Nuanced Stances and Sentiment Towards Entities of US Politicians through the Lens of Moral Foundation Theory. In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media. Association for Computational Linguistics, Online, 1–13. https://doi.org/10.18653/v1/2021.socialnlp-1.1
[41] Shamik Roy, Maria Leonor Pacheco, and Dan Goldwasser. 2021. Identifying Morality Frames in Political Tweets using Relational Learning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 9939–9958. https://doi.org/10.18653/v1/2021.emnlp-main.783
[42] Wesley Santos and Ivandré Paraboni. 2019. Moral Stance Recognition and Polarity Classification from Twitter and Elicited Text. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019). INCOMA Ltd., Varna, Bulgaria, 1069–1075. https://doi.org/10.26615/978-954452-056-4123
[43] Skipper Seabold and Josef Perktold. 2010. Statsmodels: Econometric and statistical modeling with python. In Proceedings of the 9th Python in Science Conference, Vol. 57. Scipy, Austin, Texas, 61. http://conference.scipy.org/proceedings/ scipy2010/pdfs/seabold.pdf
[44] Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. Journal of Machine Learning Research 15, 56 (2014), 1929–1958. http://jmlr.org/papers/v15/srivastava14a.html
[45] Brandon D. Stewart and David S. M. Morris. 2021. Moving Morality Beyond the In-Group: Liberals and Conservatives Show Differences on Group-Framed Moral Foundations and These Differences Mediate the Relationships to Perceived Bias and Threat. Frontiers in Psychology 12 (apr 2021), 28–32. https://doi.org/10.3389/ fpsyg.2021.579908
[46] Tijmen Tieleman and Geoffrey Hinton. 2012. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning 4, 2 (2012), 26–31. https://www.cs.toronto.edu/ ~tijmen/csc321/slides/lectureslideslec6.pdf
[47] Martijn van Zomeren, Tom Postmes, and Russell Spears. 2012. On conviction’s collective consequences: Integrating moral conviction with the social identity model of collective action. British Journal of Social Psychology 51, 1 (mar 2012), 52–71. https://doi.org/10.1111/j.2044-8309.2010.02000.x
[48] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Online, 38–45. https://doi.org/10.18653/v1/2020.emnlp-demos.6
[49] Christopher Wolsko, Hector Ariceaga, and Jesse Seiden. 2016. Red, white, and blue enough to be green: Effects of moral framing on climate change attitudes and conservation behaviors. Journal of Experimental Social Psychology 65 (2016), 7–19. https://doi.org/10.1016/j.jesp.2016.02.005
[50] Jing Yi Xie, Renato Ferreira Pinto Junior, Graeme Hirst, and Yang Xu. 2019. Textbased inference of moral sentiment change. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 4654–4663. https://doi.org/ 10.18653/v1/D19-1472
A APPENDIX ANALYSIS A.1 Cross-domain Performance Analysis
We report cross-domain performance of morality classifiers in Figure 5. The cross-domain evaluations simulate the scenario that trains a classifier on existing source domain and applies the classifier on the new target domain. The x-axis is for training domains, and the y-axis is for testing domains. For example, given a pair of ALM (x-axis) and BLM (y-axis), the 0.38 indicates that we trained
130
0.49 0.3 0.54 0.38 0.16
ALM
0.077 0.68 0.41 0.62 0.023
Baltimore
0.38 0.35 0.77 0.41 0.078
BLM
0.085 0.59 0.38 0.66 0.069
Election
0.054 0.23 0.19 0.25 0.41
Sandy
ALM Baltimore BLM Election Sandy
Figure 5: Cross-domain classification performance by the F1 score.
a classifier on ALM data and tested the classifier on the BLM domain. The diagonal line indicates that we evaluate the classifiers within the same source domains. We can find that evaluating classifiers within the same domain achieves better performance than out-domain evaluations. We can also find that domains sharing similar topic features and moral values have closer performance. For example, ALM and BLM are closer topics, and applying BLMtrained classifier on the ALM achieves 0.54, which is closer to the in-domain performance (0.77) than the other social issues.
A.2 Top Feature Analysis
The quantitative analysis in Section 2 and Appendix A.1 motivates us to explore further the domain shifts based on word features. To achieve this, we first binarize morality labels and build logistic regression classifiers using TF-IDF-weighted n-gram features (uni- , bi-, and tri-grams). The classifier keeps the most frequent 15K features during the training. We then train a classifier for each social issue domain and rank the word features by mutual information classification [33]. Finally, we remove stopwords by the NLTK [4] and present the top unigram features in Table 6. The qualitative results show the most predictable word features towards the 10 moral values. Each domain shows its unique patterns of language use that reflect the morality, ideologies, and sentiments of online users. We notice that the top features may reflect the societal and cultural backgrounds of the social issues. For example, the top features of the Election domain correlate with political topics, such as maga (make America Great Again) and trump (US president); the Sandy happened in 2012 during Obama’s presidency; and “god” connects to religious belief and ranks among top features in the multiple social issues. We can observe that variations may also exist within the same social issue domain. For example, while the hashtag “baltimoreuprising” supports the Baltimore protest, the hashtag “baltimoreriots” criticizes the protest as riots and violence. The language use shifts suggest that such word feature variations may impact extracted feature representations and weaken morality classifiers for new target domains.
<!-- PDF page 11 -->
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting HT ’22, June 28-July 1, 2022, Barcelona, Spain
Table 6: Top predictable unigram features regarding moral values extracted by the mutual information. The Vaccine refers to our annotated COVID-19 vaccine data.
Domain Features
ALM people, god, justice, police, love, black, respect, human, racist, violence Baltimore freddiegray, baltimore, baltimoreriots, baltimoreuprising, police, people, justice, black, deray, man BLM blm, injustice, solidarity, police, black, people, ferguson, iuic, respect, blacktwitter Davidson bitch, hoes, bitches, pussy, fuck, shit, nigga, ass, lol, trash Election realdonaldtrump, trump, potus, president, justice, gop, maga, america, obey, donaldtrump MeToo god, love, people, justice, us, rights, human, hurt, respect, women, equality Sandy sandy, hurricanesandy, liberty, hurricane, holy, bitch, frankenstorm, god, love, obama Vaccine get, got, first, getting, vaccines, coronavirus, people, today, dose, worry
131
<!-- PDF page 12 -->
@{
---
Rendered visual and table regions
Complete source-page renderings preserve figures, tables, equations, and their layouts where PDF text extraction cannot faithfully express geometry.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime