Leveraging GPT for the Generation of Multi-Platform Social Media Datasets for Research
Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs and platform regulations. As such, acquiring datasets that span multiple platforms which are crucial for a comprehensive understanding of the digital ecosystem is particularly challenging. This paper explores the potential of large language models to create lexically and semantically relevant social media datasets across multiple platforms, aiming to match the quality of real datasets. We employ ChatGPT to generate synthetic data from a real dataset consisting of posts from three different social media platforms. We assess the lexical and semantic properties of the synthetic data and compare them with those of the real data. Our empirical findings suggest that using large language models to generate synthetic multi-platform social media data is promising. However, further enhancements are necessary to improve the fidelity of the outputs.

Leveraging GPT for the Generation of Multi-Platform Social Media Datasets for Research

Henry Tari, Maastricht University, Netherlands, h.tari@student.maastrichtuniversity.nl

M. Danial Khan, Maastricht University, Netherlands, md.khan@student.maastrichtuniversity.nl

Darian Othman, Maastricht University, Netherlands, do.othman@student.maastrichtuniversity.nl

Thales Bertaglia, Utrecht University, Netherlands, t.f.costabertaglia@uu.nl

Rishabh Kaushal, Maastricht University, Netherlands, rishabh.kaushal@maastrichtuniversity.nl

Adriana Iamnitchi, Maastricht University, Netherlands, a.iamnitchi@maastrichtuniversity.nl

Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs and platform regulations. As such, acquiring datasets that span multiple platforms which are crucial for a comprehensive understanding of the digital ecosystem is particularly challenging. This paper explores the potential of large language models to create lexically and semantically relevant social media datasets across multiple platforms, aiming to match the quality of real datasets. We employ ChatGPT to generate synthetic data from a real dataset consisting of posts from three different social media platforms. We assess the lexical and semantic properties of the synthetic data and compare them with those of the real data. Our empirical findings suggest that using large language models to generate synthetic multi-platform social media data is promising. However, further enhancements are necessary to improve the fidelity of the outputs.

CCS Concepts: • Information systems → Social networking sites;

Keywords: Social Media Research, LLMs, Synthetic Data

ACM Reference Format: Henry Tari, M. Danial Khan, Justus Rutten, Darian Othman, Thales Bertaglia, Rishabh Kaushal, and Adriana Iamnitchi. 2024. Leveraging GPT for the Generation of Multi-Platform Social Media Datasets for Research. In 35th ACM Conference on Hypertext and Social Media (HT '24), September 10--13, 2024, Poznan, Poland. ACM, New York, NY, USA 7 Pages. https://doi.org/10.1145/3648188.3675153

1 INTRODUCTION

Social media platforms help us study user behaviour, their interaction with content and with each other. Social media data is never perfect [19] and is often costly to collect and curate [10]. In addition, it cannot be shared with the research community due to privacy and legal concerns even when social media posts are public. Many social platforms have emerged with different types of content and audiences. A single platform is not an accurate reflection of society. Consequently, research has shifted towards multi-platform studies, including topics such as election discussions [16, 18], coordinated information operations [13], and cyberbullying [20]. However, datasets from multiple social media platforms are even more challenging to (re)collect and curate.

This paper poses the following question: Given a dataset comprising text messages from various social media platforms on a specific topic within a certain time frame, how accurately can we generate a corresponding synthetic dataset? Specifically, we examine a multi-platform social media dataset that includes messages pertaining to certain topics over a specified period. Our long-term goal is to create synthetic multi-platform data thereby facilitating reproducibility in social media research, as outlined in [17], without contravening legal or platform-specific regulations. We take the first steps towards this vision by making the following contributions. First, we propose and evaluate two prompt engineering strategies appropriate for our objectives. Second, we quantify the fidelity of the synthetic datasets for three different social media platforms: Twitter, Facebook, and Reddit. We discovered that, while convincing at the individual social media post level in terms of appropriate use of lexical features such as hashtags and emojis, at the collection level the (re)use of tags is unrealistically low. Third, we show through analyses using a multi-platform dataset related to the US 2022 midterm elections that chatGPT-generated data semantically reproduces the topics in the real data.

2 RELATED WORK

Moller et al. [14] found that synthetic data generated from GPT-4 and Llama-2 improves rare class performance in multi-class classification tasks related to computational social science. Bertaglia et al. [3] generated synthetic Instagram posts using four prompting strategies to augment training data for the sponsored content detection task. They observed that the fidelity of the generated synthetic dataset and its utility for the downstream task can conflict. Ghanadian et al. [6] investigated the feasibility of generating synthetic Reddit posts for suicidal ideation. LLMs were prompted to generate suicidal text related to different factors like depression, anxiety, anger, hopelessness, etc. Hartvigsen et al. [8] created ToxiGen, a dataset of toxic and benign statements synthetically generated using GPT-3 by giving examples taken from Reddit in their prompt. Similarly, Das et al. [4] proposed a dataset comprising implicit offensive speech targeted against 38 groups. Veselovsky et al. [22] focused on the sarcasm detection task using self-disclosed sarcastic tweets.

These works focus mainly on the utility of the resulting synthetic dataset to solve a downstream task in a single-platform setting. As a consequence, their prompting approach is often biased towards a specific downstream task. In contrast, our approach in this paper focuses on generating high-fidelity synthetic datasets for multi-platform social media data without biasing it into a particular downstream task.

3 DATASET

For this study, we selected the multi-platform dataset collected by Aiyappa et al. [1] on the topic of the US elections 2022 from three platforms: Twitter, Facebook and Reddit. Posts from these platforms are highly distinct in length and usage of textual features such as hashtags and URLs. For instance, hashtags are hardly used on Reddit but are common on Twitter. Because our focus is on textual data only and on evaluating the efficacy of LLM's generative capability with few shot grounding samples, we randomly selected 1k posts from each platform in this work.

4 SYNTHETIC DATA GENERATION VIA CHATGPT

Figure 1

Pipeline of our methodology for generating and evaluating synthetic social media datasets.

Figure 1: Pipeline of our methodology for generating and evaluating synthetic social media datasets.

To generate data, we opted for OpenAI's model GPT-3.5-turbo, which is a text completion model capable of retaining context from previous interactions. It is also the most cost-effective option among OpenAI's model offerings, making it preferable for generating larger datasets. Another factor that influenced our choice of this model is its compact size compared to other options, which helps reduce the carbon footprint of our approach, especially since we are generating large multi-platform datasets.

We employed various strategies to generate text resembling realistic social media posts. We began with a basic zero-shot prompt instructing the generation of a specific social media post such as, “Generate a Twitter post based on US elections 2022”, but it proved ineffective in producing diverse results: Upon receiving multiple iterations of the same prompt, chatGPT starts reproducing the same outputs with only minor variations such as adding an emoji or using a different hashtag at the end. As a consequence, we tried few-shot prompting by providing a variety of contextual examples in prompts. We experimented with prompts ranging from two to five input examples. Generally, using more examples enhanced the diversity and variation in the output, which seemed particularly effective for platforms with shorter messages, such as Twitter. After experimenting with different numbers of grounding examples, across all the platform, we obtained the best results when providing three randomly chosen posts from our real dataset and prompting chatGPT to generate three output posts. Using more than three examples becomes problematic when dealing with platforms like Reddit, where the long text messages are challenging for the token limit of a single API call and for keeping track of context.

Figure 1 depicts the pipeline of our prompt approach. We implemented our prompting approach using the OpenAI API1. Each prompt specified a desired JSON output format to facilitate the extraction of the generated posts. Outputs that deviated from the specified JSON format were discarded, as the model occasionally produced invalid results. During our initial prompt experience, we observed that chatGPT often ignored the context of the three input examples provided. Therefore, we refined the prompt instruction and included phrases like ‘Using following examples...’. This modification improved performance within the given context, aligning results more closely with our real dataset. However, after some runs, the chatGPT model often produced results that were mere reproductions of the examples. Henceforth, we further rephrased our prompt instruction by adding phrases like ‘generate new posts keeping them true to their content and writing style.’ This prompt instruction encouraged the generation of diverse and novel social media posts, resembling the examples but distinct enough to be considered unique.

Ultimately, we converged to two different types of prompt, ‘platform aware’ and ‘platform agnostic’. In platform-aware prompting, we explicitly ask in the prompt instructions to generate text for the target social media platform. This prompting strategy was adopted to leverage the model's capability to generate posts with styling characteristics that are typical of the target social media platform. In platform-agnostic prompting, the chatGPT model was not informed about the target social media platform, but it was asked to keep the generated output similar in content and length to the examples provided.

We experimented with different values for the two primary parameters of the model, temperature (T) and $top\p$ (P). T values exceeding 1 yielded highly varied yet extremely artificial results, while those below 0.7 mostly replicated the given examples. T values at 0.7 and 1, however, produced realistic variations, which were then selected for further exploration. Likewise, the P values exhibited a comparable trend. Consequently, values of 1 and 0.7 for both T and P were chosen for further experimentation. When both parameters were set to 0.7, sub-optimal results were obtained; hence, this combination was discarded from further experiments. We perform experiments on three combinations of T and P values, namely, <T=0.7, P=1>, <T=1, P=0.7>, and <T=1, P=1>.

5 EVALUATION OF SYNTHETIC SOCIAL MEDIA POSTS

We evaluate the textual fidelity of the synthetic posts compared to real posts. We measure fidelity along four dimensions. First, we compare the occurrence of platform-specific textual features typical to social media posts, namely hashtags, user tags, URLs, and emojis. Second, we compare the sentiment of synthetic to real posts. Third, we assess the diversity and coherence of topics found in synthetic posts compared to real posts. Finally, we measure text similarity between real and synthetic posts in the embedding space.

5.1 Hashtags, User Tags, URLs and Emojis in Synthetic Data

Table 1 presents the average number of hashtags, user tags, URLs, and emojis, respectively, per real and synthetic post, along with distinct occurrences of these tokens. On average, chatGPT (particularly platform agnostic) tends to create more hashtags for Twitter, Facebook, and even Reddit, where the use of hashtags in reality is rather low. However, when chatGPT was made aware of the platform (say Reddit), it creates less hashtags. In terms of hashtag diversity, we observe a larger number of distinct hashtags in the synthetic data than in the real (shown between parentheses in Table 1), indicating lower hashtag reuse among posts in the synthetic datasets than in reality. On social media platforms, hashtags are meant to connect posts with similar topics, yet, understandably, chatGPT misses this point. This difference was also observed in [3] and likely relates to the sequential manner in which LLMs work, without trying to make connections between individual outputs. However, this may also suggest that prompting strategies that emphasize shared hashtags among the examples given in prompts could lead to more realistic generated datasets.

Table 1: The average number of lexical features per message in the real and synthetic social media datasets. The number of distinct tokens appears within parentheses. Each data collection includes 1k social media posts from each platform.

Platform

Token

Real

Platform Agnostic



Platform Aware






T=0.7 P=1

T=1 P=0.7

T=1 P=1

T=0.7 P=1

T=1 P=0.7

T=1 P=1

Twitter

Hashtag

1.19 (964)

1.77 (1402)

1.74 (1403)

1.84 (1458)

1.73 (1378)

1.69 (1304)

1.81 (1405)


Tag

1.02 (930)

0.72 (680)

0.70 (654)

0.56 (516)

0.65 (608)

0.60 (555)

0.54 (510)


URL

0.35 (356)

0.13 (126)

0.12 (123)

0.07 (75)

0.06 (62)

0.07 (69)

0.05 (50)


Emoji

0.30(175)

0.34 (183)

0.32 (186)

0.43 (225)

0.31 (163)

0.26 (143)

0.37 (193)

Facebook

Hashtag

0.92 (830)

1.74 (1420)

1.78 (1420)

1.75 (1381)

1.28 (1080)

1.35 (1114)

1.38 (1118)


Tag

0.05 (48)

0.01 (16)

0.19 (22)

0.03 (36)

0.01 (16)

0.02 (19)

0.01 (9)


URL

0.17 (190)

0.04 (44)

0.05 (56)

0.04 (38)

0.04 (39)

0.04 (44)

0.03 (32)


Emoji

0.00 (3)

0.10 (78)

0.08 (55)

0.15 (99)

0.11 (22)

0.07 (58)

0.17 (95)

Reddit

Hashtag

0.14 (72)

1.24 (963)

1.15 (830)

1.15 (908)

0.13 (118)

0.10 (94)

0.12 (128)


Tag

0.00 (2)

0.00 (2)

0.00 (1)

0.03 (4)

0.00 (1)

0.00 (1)

0.00 (1)


URL

0.09 (110)

0.01 (11)

0.01 (16)

0.01 (14)

0.01 (14)

0.02 (19)

0.01 (14)


Emoji

0.01 (3)

0.08 (45)

0.04 (31)

0.11 (63)

0.02 (22)

0.01 (14)

0.02 (16)

Tagging other platform users is common on Twitter but almost non-existent on Reddit. In our experiments, chatGPT underestimates the use of user tags in Twitter and Facebook, except for Reddit, where it correctly estimates virtually no use of such tags. The number of URLs included in posts is also underrepresented in synthetic data across all platforms. It is hard for chatGPT to generate URLs in the synthetic posts. Finally, emojis are represented in similar proportion in in the synthetic datasets for Twitter. However, in case of Facebook and Reddit, emojis are over represented in the synthetic data. Moreover, synthetically generated posts include a much larger variety of emojis than in our real data, especially for the more restrained platforms, such as Facebook and Reddit. Finally, we may conclude that, while we experimented with two prompting strategies and three combinations of parameters each, we could not see a consistent benefit of any one combination in these lexical metrics.

5.2 Sentiment Analysis

Moving beyond lexical tokens, we compare sentiments expressed in generated content with real data. We employed the cardiffnlp/twitter-roBERTA-base-sentiment-latest model proposed by Loureiro et al. [12] available at Hugging Face2 to classify posts as positive, negative, or neutral. The model is based on the RoBERTa architecture [11] and is pre-trained for language modelling on  124M tweets and then fine-tuned for sentiment analysis on the TweetEval benchmark [2]. Table 2 presents the percentage of posts belonging to a particular sentiment category for each platform. We observe that generated content is more positive and less negative than reality for all platforms. Neutral sentiment is also under-expressed in synthetic content across all platforms except for Reddit. Recent work [4, 5, 8, 24] has shown the use of LLMs for generating toxic content. Therefore, our observed outputs from the chatGPT suggest that efforts are being made to decrease negativity in the generated content. However, from the standpoint of improving fidelity, prompting techniques that use positive intent as presented in [4] could generate content that better matches the negative sentiment in real posts. In the rest of our evaluation section we will only include the default parameters (temperature T = 1, probabilistic sampling P = 1) for the two prompting strategies.

Table 2: Percentage of posts out of a collection of 1k for each platform and setting that are classified based on sentiment analysis.

Platform

Sentiment

Real

Platform Agnostic



Platform Aware






T=0.7 P=1

T=1 P=0.7

T=1 P=1

T=0.7 P=1

T=1 P=0.7

T=1 P=1

Twitter

Negative

46.76

32.87

33.30

26.75

31.86

31.85

29.50


Neutral

21.06

17.93

15.91

15.50

15.48

16.37

13.50


Positive

32.18

49.20

50.79

57.75

52.66

51.78

57.00


Negative

34.19

22.74

24.44

19.77

21.98

23.47

18.97

Facebook

Neutral

31.80

29.73

29.77

27.28

26.57

26.95

27.51


Positive

34.01

47.53

45.79

52.95

51.45

49.58

53.52


Negative

81.91

57.54

56.54

56.03

65.10

62.93

60.95

Reddit

Neutral

7.95

15.25

15.94

14.56

15.29

17.45

18.43


Positive

10.14

27.21

27.52

29.41

19.61

19.62

20.62

Figure 2

Topic overlap among platforms in the real and synthetic datasets. Platform-agnostic/aware prompts, P = 1, T = 1.

Figure 2: Topic overlap among platforms in the real and synthetic datasets. Platform-agnostic/aware prompts, P = 1, T = 1.

Table 3: Comparison of topics (representative names of topics corresponding to topic words are generated through GPT4 prompting) in US elections datasets. Fb: Facebook, Rd: Reddit, Tw: Twitter.

Elections Dataset






Fb

Rd

Tw

Real

Platform Agnostic

Platform Aware

✓

✓

✓

"Voter Engagement".

"Voter Engagement".

"Election Process and Voting Methods".

✓

✓


"Mail-in Voting and Absentee" and "Women Reproduction Rights"

"Mail-in Voting", "Legislative Control", "Social Media Misinformation" and "Women Reproduction Rights".

"Women Reproduction Rights".

5.3 Topic Generation and Overlap

Considering that social media posting behavior revolves around topics, we analyze topics in the synthetic data and compare them with topics in real posts. Intuitively, the generated data should be representative of the midterm US 2022 elections, yet new topics in these general subjects are welcome for diversity and potential utility of synthetic data. For topic extraction we used BERTopic [7] which is being increasingly used for social media data [9, 21]. For sentence embedding required by BERTopic, we used all-MiniLM-L6-v23 [23] sentence transformer that works with sentences as well as short paragraphs, which is useful in our case because of the length variety from long Reddit posts to short Twitter tweets. We performed minimal text cleaning for optimization purposes. Salient features such as hashtags and emojis are useful in text analysis. Therefore, only URLs and stop words were removed during this process. We focused only on the topics that appeared in at least 10 social media posts on the corresponding platform. Figure 2 shows the number of shared and disjoint topics across different platforms in real and synthetic datasets. For topic overlap, we applied Open AI's embedding model [15] (text-embedding-3-large) on top-10 most representative words for each topic and used cosine similarity for topic comparison with a threshold obtained empirically.

Figure 3

Comparison of topics between real and synthetic data generated using both prompt strategies on the US elections (a) Topic overlap count, (b) Word clouds of unique topics in the real dataset, (c) Word clouds of common topics, and (d) Word clouds of unique topics in the synthetic dataset.

Figure 3: Comparison of topics between real and synthetic data generated using both prompt strategies on the US elections (a) Topic overlap count, (b) Word clouds of unique topics in the real dataset, (c) Word clouds of common topics, and (d) Word clouds of unique topics in the synthetic dataset.

In the elections dataset of real posts, there is more topic diversity on Facebook and Reddit than on Twitter. Both prompting approaches can replicate the topic diversity, though not as much in the case of Facebook. Platform-aware prompts generate more topics in Reddit, perhaps because Reddit post lengths are the highest among all platforms, giving chatGPT more grounding. The topic ‘Voter Engagement’ appears common in all platforms, both in real and synthetic (platform agnostic) data. ‘Women Reproduction Rights’ appears as a common topic across Facebook and Reddit. Platform agnostic prompting generates four topics of which two (‘Mail-in Voting’ and ‘Women Reproduction Rights’) are the same as those present in the real posts. In the influencer dataset of real posts, we observe that topics related to health, wellness, and fitness occur across all platforms.

Figures 3a and  3b present topic overlap between real and synthetic (both prompting) datasets along with word clouds of topic unique and common across datasets. There are fewer topics in generated content in elections. The word clouds are produced from the text generated by GPT 4 as descriptions of the top 10 words that describe a topic extracted with BERT. We observe that discussions in the synthetic data remain on the same general subjects as in the real datasets: for example, the synthetic US election datasets include state names (e.g., Arizona, Florida), typical election terminology (candidate, vote, ballot, absentee), politicians active in this context (Stefanik, Desantis, Feterman, Biden), and timely topics (border security, the opioid crisis, the release of Brittney Griner). Different topics and keywords have become more popular in the synthetic dataset, thus appearing in the word cloud: for example, “Arabia” in the word cloud representing topics unique to the synthetic data in Figure 3b appears in the original dataset (in the context of oil prices) but it is emphasized significantly in the generated data.

5.4 Embedding Similarity

Most downstream tasks are solved by converting text into embedding vectors. Therefore, assessing the similarity of generated content with real content in embedding space becomes important. To evaluate the similarity between real and synthetic social media posts, we used OpenAI's text-embedding-3-large model [15] to convert all text into embedding vectors. We then calculated pairwise cosine similarity for all possible pairs drawn from real and synthetic posts. To measure similarity, we selected the top 1k highest similarity scores from each set and computed their average. This average represents the central tendency of the most similar pairs, providing a focused measure of similarity that highlights the best matches between the datasets. Additionally, we computed the average of all similarity scores across each dataset to establish a baseline, giving an overall sense of similarity regardless of the top matches. Table 4 displays the similarity matrix for platform agnostic and platform aware prompts with default settings in the election dataset. In our experiment, platform-agnostic prompts yielded content with slightly higher similarity scores than the platform-aware prompt. Overall, both prompts generate content similar to the real dataset. Among the platforms, Facebook exhibited the highest similarity among top 1k highest similarity scores for both prompting strategies followed by Twitter and Reddit.

Table 4: Similarity matrix computed using cosine distance for Platform Agnostic and Platform Aware Prompt and T= 1, P= 1 parameters for Election and Influencer dataset.

Platform

Platform Agnostic


Platform Aware



Top 1k

Average

Top 1k

Average

Facebook

0.824

0.234

0.815

0.235

Twitter

0.802

0.229

0.804

0.233

Reddit

0.762

0.263

0.754

0.261

Figure 4

t-SNE plots of embedding vectors (1k data points are clustered into 50 clusters, and cluster centroids are plotted) of real, and synthetic (platform agnostic and platform aware).

Figure 4: t-SNE plots of embedding vectors (1k data points are clustered into 50 clusters, and cluster centroids are plotted) of real, and synthetic (platform agnostic and platform aware).

To visualize embedding vectors, we created t-SNE plots using the embeddings of both real and synthetic datasets with the default settings (T=1, P=1). For each 1k data points across all scenarios (platform aware, platform agnostic, and real), we apply K-means clustering (K=50) and represent cluster centroids as dots in Figure 4. We experimented with different perplexity values and found consistent results, the plots are drawn on default perplexity value. While t-SNE has limitations, it presents a good estimate for the preservation of local structure in data. In the election dataset, synthetic data shows high similarity with real data for Facebook. In the case of Reddit and Twitter, synthetic datasets (platform aware and platform agnostic) show similarities along only one of the reduced dimensions. Real Reddit posts appear near to their synthetic counterparts, but for Twitter, they are quite far apart.

6 SUMMARY AND DISCUSSIONS

This paper investigates the use of chatGPT in generating social media datasets across various platforms and topics, focusing on two English-language datasets. The dataset includes posts from Facebook, Reddit, and Twitter related to the US midterm 2022 elections.

We employed two prompting strategies: one where chatGPT was prompted to generate posts similar to given examples without specifying the platform, and another where it was prompted to generate platform-specific posts. We also varied two parameters: temperature and probabilistic sampling. We evaluated the quality of the generated datasets based on lexical features (hashtags, URLs, emojis, and user tags), sentiment, topics, and embedding similarity. Our findings indicate that the synthetic dataset generally display high fidelity across various metrics. The synthetic posts notably preserve key social media lexical features such as emojis and hashtags and maintain semantic similarities with the original datasets, demonstrated through topic analysis and embedding similarity. However, neither prompting strategy consistently outperformed the other, and adjustments in parameters temperature and probabilistic sampling within our set limits did not significantly impact our evaluations.

Notable differences were observed in the utility of synthetic data for certain downstream tasks. For instance, chatGPT-generated datasets underutilized tags and URLs compared to real datasets, with synthetic tags often not aligning with practices observed in human-generated content. Additionally, synthetic posts exhibited a disproportionately positive sentiment, particularly noticeable on Twitter, suggesting that other LLMs or alternative prompting strategies might better suit tasks requiring a broader range of sentiments. The study has limitations that will be explored in future research. First, our results are based solely on the output of GPT-3.5-turbo as of early 2024, setting a benchmark for future studies that might replicate our approach with different LLMs. Second, as our datasets cover non-overlapping sets of platforms, it remains unclear whether the observed results are platform-specific or data-specific. Future studies will aim to explore various topics across the same platforms. Third, our focus was limited to English due to dataset availability and our linguistic proficiency, but future efforts will aim to include more diverse languages. Lastly, we did not address privacy and legal concerns about releasing synthetic datasets.

This research demonstrates that generating synthetic social media datasets spanning multiple platforms is feasible with current technology. Improved prompting techniques and post-processing could improve the fidelity of the data set. Realistic multi-platform synthetic datasets can promote reproducibility in social media research and offer resilience against restrictive platform policies, facilitating access to data for researchers globally.

REFERENCES

    Rachith Aiyappa, Matthew R. DeVerna, Manita Pote, Bao Tran Truong, Wanying Zhao, David Axelrod, Aria Pessianzadeh, Zoher Kachwala, Munjung Kim, Ozgur Can Seckin, Minsuk Kim, Sunny Gandhi, Amrutha Manikonda, Francesco Pierri, Filippo Menczer, and Kai-Cheng Yang. 2023. A Multi-Platform Collection of Social Media Posts about the 2022 U.S. Midterm Elections. Proceedings of the International AAAI Conference on Web and Social Media 17 (June 2023), 981–989.

    Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification. In Findings of the Association for Computational Linguistics: EMNLP 2020. ACL Anthology, online, 1644–1650.

    Thales Bertaglia, Lily Heisig, Rishabh Kaushal, and Adriana Iamnitchi. 2024. InstaSynth: Opportunities and Challenges in Generating Synthetic Instagram Data with ChatGPT for Sponsored Content Detection. Proceedings of the International AAAI Conference on Web and Social Media 18, 1 (May 2024), 139–151.

    Amit Das, Mostafa Rahgouy, Dongji Feng, Zheng Zhang, Tathagata Bhattacharya, Nilanjana Raychawdhary, Mary Sandage, Lauramarie Pope, Gerry Dozier, and Cheryl Seals. 2024. OffLanDat: A Community Based Implicit Offensive Language Dataset Generated by Large Language Model Through Prompt Engineering. arXiv preprint arXiv:2403.02472 3, 02472 (2024), 1–15.

    Arka Dutta, Adel Khorramrouz, Sujan Dutta, and Ashiqur R KhudaBukhsh. 2023. Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models. arXiv e-prints 1, 2309 (2023), arXiv–2309.

    Hamideh Ghanadian, Isar Nejadgholi, and Hussein Al Osman. 2024. Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models. IEEE Access 12, 1 (2024), 1–14.

    Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794 1, 1 (2022), 1–10.

    Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509 03, 09509 (2022), 1–13.

    Yining Hua, Hang Jiang, Shixu Lin, Jie Yang, Joseph M Plasek, David W Bates, and Li Zhou. 2022. Using Twitter data to understand public perceptions of approved versus off-label use for COVID-19-related medications. Journal of the American Medical Informatics Association 29, 10 (2022), 1668–1678.

    David M. J. Lazer, Alex Pentland, Duncan J. Watts, Sinan Aral, Susan Athey, Noshir Contractor, Deen Freelon, Sandra Gonzalez-Bailon, Gary King, Helen Margetts, Alondra Nelson, Matthew J. Salganik, Markus Strohmaier, Alessandro Vespignani, and Claudia Wagner. 2020. Computational social science: Obstacles and opportunities. Science 369, 6507 (2020), 1060–1062.

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. ArXiv abs/1907.11692 (2019), 1–13.

    Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, and Jose Camacho-Collados. 2022. TimeLMs: Diachronic Language Models from Twitter. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. Association for Computational Linguistics, Dublin, 251–260.

    Josephine Lukito. 2020. Coordinating a Multi-Platform Disinformation Campaign: Internet Research Agency Activity on Three U.S. Social Media Platforms, 2015 to 2017. Political Communication 37, 2 (2020), 238–255.

    Anders Giovanni Møller, Jacob Aarup Dalsgaard, Arianna Pera, and Luca Maria Aiello. 2023. Is a prompt and a few samples all you need? Using GPT-4 for data augmentation in low-resource classification tasks. arXiv preprint arXiv:2304.13861 04, 13861 (2023), 1–12.

    Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005 01, 10005 (2022), 1–13.

    Francesco Pierri, Geng Liu, and Stefano Ceri. 2023. ITA-ELECTION-2022: A Multi-Platform Dataset of Social Media Conversations Around the 2022 Italian General Election. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management(CIKM ’23). Association for Computing Machinery, Birmingham, 5386–5390.

    David Schoch, Chung hong Chan, Claudia Wagner, and Arnim Bleier. 2023. Computational Reproducibility in Computational Social Science. arxiv:2307.01918 [cs.CY]

    Juan Carlos Medina Serrano, Morteza Shahrezaye, Orestis Papakyriakopoulos, and Simon Hegelich. 2019. The Rise of Germany's AfD: A Social Media Analysis. In Proceedings of the 10th International Conference on Social Media and Society (Toronto, ON, Canada). Association for Computing Machinery, Toronto, 214–223.

    Zeynep Tufekci. 2014. Big Questions for Social Media Big Data: Representativeness, Validity and Other Methodological Pitfalls. Proceedings of the International AAAI Conference on Web and Social Media 8, 1 (May 2014), 505–514.

    David Van Bruwaene, Qianjia Huang, and Diana Inkpen. 2020. A multi-platform dataset for detecting cyberbullying in social media. Lang. Resour. Eval. 54, 4 (dec 2020), 851–874.

    Tim Verbeij, Ine Beyens, Damian Trilling, and Patti M Valkenburg. 2024. Happiness and Sadness in Adolescents’ Instagram Direct Messaging: A Neural Topic Modeling Approach. Social Media+ Society 10, 1 (2024), 20563051241229655.

    Veniamin Veselovsky, Manoel Horta Ribeiro, Akhil Arora, Martin Josifoski, Ashton Anderson, and Robert West. 2023. Generating faithful synthetic data with large language models: A case study in computational social science. arXiv preprint arXiv:2305.15041 05, 15041 (2023), 1–8.

    Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems 33 (2020), 5776–5788.

    Boyang Zhang, Xinyue Shen, Wai Man Si, Zeyang Sha, Zeyuan Chen, Ahmed Salem, Yun Shen, Michael Backes, and Yang Zhang. 2023. Comprehensive Assessment of Toxicity in ChatGPT. arXiv preprint arXiv:2311.14685 11, 14685 (2023), 1–11.

FOOTNOTE

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org.

HT '24, September 10–13, 2024, Poznan, Poland

© 2024 Copyright held by the owner/author(s). Publication rights licensed to ACM.

ACM ISBN 979-8-4007-0595-3/24/09.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime