Source and license. “A Simple Language Independent Approach for Distinguishing Individuals on Social Media” by Guangyuan Piao, Proceedings of the 32nd ACM Conference on Hypertext and Social Media (HT ’21) (2021), DOI: 10.1145/3465336.3475092. Original source (version of record): https://dl.acm.org/doi/10.1145/3465336.3475092. Source PDF: /workspace/HT-2021_35-30_3465336/3465336.3475092.pdf. Licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). Changes: this is a format-converted adaptation from the source PDF; page layout, typography, links, equations, tables, and accessible Markdown markup were changed, with no substantive changes intended.
Guangyuan Piao Department of Computer Science, Maynooth University Maynooth, Co Kildare, Ireland guangyuan.piao@mu.ie
Abstract
Nowadays, the large-scale human activity traces on social media platforms such as Twitter provide new opportunities for various research areas such as mining user interests, understanding user behaviors, or conducting social science studies in a large scale. However, social media platforms contain not only individual accounts but also other accounts that are associated with non-individuals such as organizations or brands. Therefore, distinguishing individuals out of all accounts is crucial when we conduct research such as understanding human behavior based on data retrieved from those platforms. In this paper, we propose a language-independent approach for distinguishing individuals from non-individuals with the focus on leveraging their profile images, which has not been explored in previous studies. Extensive experiments on two datasets show that our proposed approach can provide competitive performance with state-of-the-art language-dependent methods, and outperforms alternative language-independent ones.
CCS Concepts: Human-centered computing → Social media; Social networking sites; Computing methodologies → Supervised learning by classification.
Keywords: Account Classification, Deep Learning, Social Media Analysis
ACM Reference Format: Guangyuan Piao. 2021. A Simple Language Independent Approach for Distinguishing Individuals on Social Media. In Proceedings of the 32nd ACM Conference on Hypertext and Social Media (HT ’21), August 30–September 2, 2021, Virtual Event, Ireland. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3465336.3475092
1. Introduction
Social media platforms such as Twitter have been widely used in different research areas to study users from various perspectives in a large scale based on the big data generated by user activities on those platforms. For example, research areas such as predicting substance usage [10], mining user interests [14, 15], and understanding user visiting behaviors [16] or personalities [7]. Although those studies usually assume all accounts retrieved from social media platforms are individuals, McCorriston et al. estimated that 9.4% of accounts on Twitter are non-individual ones (e.g., brands or organizations) [11]. Therefore, distinguishing individual users from retrieved accounts on social media platforms is crucial for studying different user behaviors such as mining user interests or understanding substance usage, and deriving conclusions out of those studies.
Previous studies for distinguishing individuals from non-individuals can be classified into two categories based on whether an approach is language dependent or independent. For example, language-dependent approaches utilize textual information, such as social posts and/or the profile description (biography) of a user in addition to a set of statistical features, e.g., the number of followees and followers on Twitter, to classify whether a given account belongs to an individual. As one might expect, the profile description of an account can provide crucial information to distinguish individuals. For example, we can assume that an account belongs to an individual if its profile description contains words such as “my”, “I”, “Dad”, “Mum”, etc. However, this line of approaches depends on language and the majority of the previous works have been focused on English users. Although English is the most popular language on Twitter, it is used in only 32% of all Twitter messages. In contrast to relying on textual information, recent studies [3, 4] have proposed leveraging statistical features for classifying accounts on social media platforms such as Twitter.
Our focus in this paper falls into the second category, i.e., language-independent approaches for classifying individual accounts. To this end, we leverage the visual content of a user (i.e., profile image), which is critical information but has not been explored in previous studies. The intuition behind our approach is that the profile image of a user should be a good indicator for the classification of accounts. Our main contributions include:
We propose a simple Language-Independent Individual Classification approach (Section 3), named LIIC, to classify social media accounts into individuals and non-individuals, with the focus on leveraging their profile images.
We evaluate our approach with several state-of-the-art approaches using two datasets with ground truth labels in Section 4, and show that LIIC can achieve competitive performance in classifying Twitter accounts compared to language-dependent approaches.
Through an ablation study in Section 5, we further reveal that profile images are indeed an important indicator for classifying individual user accounts, which have not been explored in previous studies.
2. Related Work
In this section, we review related works which are classified into language-dependent and language-independent ones.
Language-dependent approaches. This line of approaches exploits textual content such as social posts or profile descriptions of users for feature engineering or learning latent representations via deep learning approaches [5, 20, 21]. For example, Oentaryo et al. used content, social, and temporal features and investigated several machine learning approaches such as random forests and gradient boosting [6], and showed that the gradient boosting classifier provides the best performance [13]. Wood-Doughty et al. proposed using a character-based Convolutional Neural Network (CNN) [9] to learn the representation of a user’s name and incorporated profile features such as the ratio of followers to friends together for classifying individual accounts [20].
Language-independent approaches. This line of approaches uses statistical features such as the social network structure and the posting frequency of a user without relying on textual content for classifying individual accounts [3, 4, 18, 19]. For example, Tavares et al. used a naive Bayes classifier with features related to the time distribution between social posts [18]. More recently, Daouadi et al. proposed a set of comprehensive features incorporating profile- and activity-related ones, and used gradient boosting regression trees and random forests for classifying individual accounts on Twitter [3, 4].
Despite the appreciable body of previous studies, the profile image of an account, which might be a critical indicator for the classification, has not been explored. In this work, we close the gap and focus on leveraging profile images for classifying individuals and use other profile-related features only if those images are not retrievable or are default ones. Our approach can be considered as one of the language-independent approaches as this approach does not require textual content such as social posts or profile descriptions.
3. LIIC: Language Independent Individual Classifier
In this section, we introduce our Language Independent Individual Classifier (LIIC) and its components in Section 3.1, and provide the training details of LIIC in Section 3.2.
3.1. LIIC Architecture
Figure 1 illustrates an overview of the LIIC architecture. LIIC uses three types of input such as the profile image, screen name, and profile features of a user, and leverages three different types of neural networks to learn the representations of each input for the binary classification of the target account (individual or non-individual).
The main assumption of LIIC is that the profile image of an account should be a good indicator for distinguishing individuals. For example, an individual user tends to use his/her face photo while non-individual accounts for companies or conferences tend to use their organization logos instead. However, using profile images only can limit the capacity for classifying accounts, e.g., when the retrieved profile images are default ones from Twitter, or a profile image can not be retrieved using its URL obtained from the Twitter API. Therefore, we also use screen names and profile features in addition to profile images.
Profile image representation. Given profile images as input, it is natural to use a CNN to extract image features. Instead of training a CNN from scratch, we use a pre-trained CNN model – VGG16 [17] to extract the image features, where the last layer of VGG16 is used for generating an image representation vector $vi in mathbb{R}^{20548}$. Afterwards, $vi$ is fed into two dense layers to output the final representation $vp in mathbb{R}^{256}$:
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime