Seeing Through AI's Lens: Enhancing Human Skepticism Towards LLM-Generated Fake News
Converted the HT'24 paper on the ESAS metric for detecting LLM-generated fake news (via arXiv 2406.14012) into faithful index.md (full text, 16 figures, 4 tables, 34 references) plus a no-CC stub.md, all under the 250KB per-asset limit.

Seeing Through AI's Lens: Enhancing Human Skepticism Towards LLM-Generated Fake News

Navid Ayoobi (University of Houston, Houston, Texas, USA), Sadat Shahriar (University of Houston, Houston, Texas, USA), Arjun Mukherjee (University of Houston, Houston, Texas, USA)

Published in HT '24: 35th ACM Conference on Hypertext and Social Media, Poznan, Poland, September 10-13, 2024 · DOI: 10.1145/3648188.3675136 · License: © Copyright held by the owner/author(s). Publication rights licensed to ACM.

Keywords: ChatGPT, Fake news, LLM-generated news, Large language models, Llama2, Mistral

Session: AI Reports

Pages: 1–11

Abstract

LLMs offer valuable capabilities, yet they can be utilized by malicious users to disseminate deceptive information and generate fake news. The growing prevalence of LLMs poses difficulties in crafting detection approaches that remain effective across various text domains. Additionally, the absence of precautionary measures for AI-generated news on online social platforms is concerning. Therefore, there is an urgent need to improve people’s ability to differentiate between news articles written by humans and those produced by LLMs. By providing cues in human-written and LLMgenerated news, we can help individuals increase their skepticism towards fake LLM-generated news. This paper aims to elucidate simple markers that help individuals distinguish between articles penned by humans and those created by LLMs. To achieve this, we initially collected a dataset comprising 39k news articles authored by humans or generated by four distinct LLMs with varying degrees of fake. We then devise a metric named Entropy-Shift Authorship Signature (ESAS) based on the information theory and entropy principles. The proposed ESAS ranks terms or entities, like POS tagging, within news articles based on their relevance in discerning article authorship. We demonstrate the effectiveness of our metric by showing the high accuracy attained by a basic method, i.e., TFIDF combined with logistic regression classifier, using a small set of terms with the highest ESAS score. Consequently, we introduce and scrutinize these top ESAS-ranked terms to aid individuals in strengthening their skepticism towards LLM-generated fake news.

CCS Concepts

• Information systems →Social networks; • Security and privacy →Social network security and privacy.

Keywords

Fake news, Large language models, LLM-generated news, ChatGPT, Llama2, Mistral

ACM Reference Format: Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee. 2024. Seeing Through AI’s Lens: Enhancing Human Skepticism Towards LLM-Generated Fake News. In 35th ACM Conference on Hypertext and Social Media (HT ’24),

Sadat Shahriar sshahria@cougarnet.uh.edu University of Houston Houston, Texas, USA

Arjun Mukherjee arjun@cs.uh.edu University of Houston Houston, Texas, USA

September 10–13, 2024, Poznan, Poland. ACM, New York, NY, USA, 11 pages. https://doi.org/10.1145/3648188.3675136

1 INTRODUCTION

The proliferation of large language models (LLMs) has facilitated the generation of text that closely resembles human writing in terms of quality. As a result, LLMs are utilized across various domains like chatbots [34], code generation [3, 18], translation [31], and text summarization [22]. Despite the beneficial roles they play, LLMs present opportunities for misuse by malicious users, enabling the dissemination of misinformation or deception. Consequently, this misuse can exacerbate the prevalence of fake reviews, disruptions in educational systems, and the erosion of trust among users of online social networks [4, 23]. Furthermore, the generative and reasoning capabilities of LLMs render them effective tools for the generation of fake news, typically with the intent of promoting specific political agendas, influence electoral outcomes, and undermining democratic principles. The pursuit of developing precise detection methods emerges as an essential solution in combating the misuse of LLMs and alleviating their consequential impacts. Currently, there is increasing research dedicated to exploring techniques for improving the accuracy of such detection systems. Literature indicates that in scenarios where there is consistency in the topic of the articles present in the training data and the testing data, and where the same LLM generates the text for both training and testing articles, detection methods exhibit notably high levels of accuracy. This phenomenon primarily stems from the decoding strategies inherent to LLMs, along with the statistical patterns they impart to detection systems [20]. Conversely, human evaluators often exhibit diminished accuracy (about 70%) in discerning AI-generated text [12] compared to detectors trained on extensive datasets. This discrepancy arises from humans’ inability to track common and shared patterns among all generated sample texts. Consequently, the virtually indistinguishable nature of generated text at the surface level presents significant challenges for human identification. On the other hand, altering the text domain or employing multiple LLMs for text generation can deviate the distribution that the LLM relies on for predicting subsequent tokens, thus potentially misleading detectors trained on specific text domains and resulting in a substantial decline in their performance [5]. These unreliable detectors may unintentionally deceive individuals, hindering their skepticism towards fake news. Watermarking was introduced as a method to enable the detection of any text generated by an LLM, irrespective of its domain [16, 17]. Watermarking involves embedding a concealed pattern within the syntactic text, facilitating its algorithmic identification.

HT ’24, September 10–13, 2024, Poznan, Poland Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee

The drawback associated with watermarking resides in its susceptibility to removal through paraphrasing. Moreover, if the watermarking procedure is open-sourced, imposters retain the ability to disrupt the watermark following the generation of fabricated content. On the other hand, if the watermarking strategy remains closed-source, it becomes imperceptible to regular users, with only model developers possessing awareness of its existence. Hence, this situation remains unchanged for individuals exposed to fake news on the platforms other than the one where the watermarking strategy is known. The rise of abundant number of LLMs from multiple research groups [33] makes it further challenging to devise a detection approach that accommodates the majority of these LLMs. Further aggravating the situation is the absence of automatic detection systems in online social networks to flag potential LLM-generated articles. Hence, there is an urgent need to enhance individuals’ proficiency in distinguishing between articles penned by humans and those generated by LLMs. By offering some hints in humanauthored news and news generated by LLMs, we can assist people in fostering a more critical perspective on the source of synthetic fake news. In this paper, we focus on elucidating straightforward indicators that enable individuals to discern whether an article is authored by a human or an LLM. Our methodology comprises four key stages. Initially, we present our collected dataset, comprising 3000 news articles sourced from trustworthy outlets and encompassing diverse topics such as sports, celebrities, history and religion, politics and government, social culture and civil rights, science and information technology. Utilizing four LLMs, we produce fake versions of the genuine news at three levels of fake, resulting in a dataset of 39k news articles. To the best of our knowledge, this is the first large dataset that can be used for analyzing fake news generated by different LLMs at varying degrees of fake. We make this dataset publicly available to facilitate further research in this area. In the second stage, we devise a metric termed as Entropy-Shift Authorship Signature (ESAS), leveraging information theory and entropy principles. This metric ranks the terms or other entities, like Part-of-Speech (POS) tagging, within the news article based on their significance in identifying the authorship of an article as either human-written or LLM-generated. In the third stage, we demonstrate the efficacy of the proposed ESAS metric in identifying significant cues within news articles. This is evidenced by the high accuracy achieved by a basic approach, specifically TF-IDF [24] combined with logistic regression classifier, when fed with a small set of terms with the highest ESAS score. The resulting accuracy exceeds current human capabilities in detecting AI-generated content. We believe that incorporating these top ESAS-ranked terms can heighten human skepticism towards LLM-generated fake news. Consequently, in the fourth stage, we delve into and highlight the most crucial terms under diverse conditions, taking into account various LLMs and fake levels. The primary contributions of this paper can be summarized as follows:

• We collect a reasonably large dataset comprising 39k news articles. These articles are either authored by humans or

generated by four different LLMs, exhibiting varying degrees of fake news.

• We introduce the ESAS metric, which ranks terms or entities, like POS tagging, within news articles based on their relevance to identifying article authorship.

• We offer news readers cues to identify the authorship of articles, enhancing their ability to detect fake news crafted by LLMs.

2 RELATED WORK

Numerous studies have concentrated on identifying fake news by examining linguistic patterns [6, 25, 26], fact-checking methodologies [11, 21], understanding writer intentions [8], and sentiment analysis [2, 7]. Choudhary et al. [6] designed a neural network model for fake news detection, utilizing syntactic, grammatical, sentimental, and readability features from news articles. Their findings revealed that a model trained exclusively on syntax, sentiment, and grammar-based features outperformed one trained on readabilitybased features. Authors in [26] proposed a set of knowledge-based features, introduced as fact-verification features, which include the credibility of the news website, the count of sources publishing the news, and the views of well-known fact-checking websites. By incorporating these with linguistic features, they achieved a 94% accuracy rate on the Buzzfeed Political News dataset. SAME, an end-to-end deep learning framework, was proposed by Cui et al. [7] to assess the latent sentiments present in user comments on social media, potentially aiding in the discrimination of fake news from reliable information. Giachanou et al. [8] approached the issue from a psycholinguistic perspective. The authors utilized the Convolutional Neural Network (CNN) and introduced the CheckerOrSpreader model. This model aims to distinguish between users who propagate fake news (spreaders) and those who challenge it (checkers). Their research outcomes revealed that checkers generally use more positive language and a richer vocabulary, whereas spreaders use more informal language, such as slang and swear words. The emergence of language models has opened up new avenues for imposters to produce vast amounts of fake news across various platforms, attempting to sway public opinion and persuade them with their LLM-generated misleading content [10, 14, 15, 20, 27, 28, 30, 32]. Recognizing the limitations of current detection systems can assist in identifying the risks posed by different imposters using LLMs. Pagnoni et al. [20] explored three threat scenarios based on an adversary’s budget and expertise, affecting the text generation style. These scenarios include text generated using available LM APIs, utilizing a pre-trained model with potential parameter modifications, and an LM model fine-tuned on specific data. The effectiveness of different detection systems was evaluated across these scenarios. Wu and Hooi [30] proposed SheepDog as a solution to counteract the deterioration of existing detection methods after LLM Style Attacks, where imposters mimic the style of trusted news sources using LLMs. SheepDog is a style-independent fake news detector that leverages LLMs to produce different reframings for each news article, introducing a range of styles. Subsequently, it maximizes the consistency in the model’s predictions across the original article and its diverse reframings. In contrast to conventional detection methods that rely on machine learning and deep

Seeing Through AI’s Lens HT ’24, September 10–13, 2024, Poznan, Poland

learning, Jiang et al. [14] explored the potential of LLMs to detect the fake content they produce. The findings indicated that ChatGPT struggles to effectively detect the disinformation it generates. However, by crafting a tailored prompt, its detection accuracy can be moderately improved. While the emphasis of most studies has been on enhancing detector accuracy, the absence of warning systems for AI-generated news on online platforms remains a reality. With the rapid introduction of more LLMs over time, this could not have been more timely to educate people about potential indicators in news articles to heighten their skepticism towards fake LLM-generated news. Therefore, in this paper, we propose simple cues that individuals can employ to distinguish between human-authored and LLM-generated news.

3 DATASET

For the purpose of this study, a total of 200k news articles sourced from reputable outlets, including ABC News, Aljazeera, American Press, Associated Press News, CBS News, CNN, NBC News, Reuters, and The Guardian, were gathered. The BERTopic [9] model was employed to filter and select 3000 articles, ensuring their primary focus aligned with one of the following six topics: sports, celebrities, history and religion, politics and government, social culture and civil rights, science and information technology. To ensure dataset balance, each of the aforementioned topics is represented by an equal sample size of 500 articles. Four models, specifically GPT 3.5, Mistral-7B [13], Llama2-7B, and Llama2-13B [29] were employed to generate LLM counterparts of the original articles. To obtain different level of fake, three strategies were devised for generating LLM samples: 1. Article Rephrasing, 2. Article Extending, and 3. Summary Expanding. The first strategy, article rephrasing, involved prompting the LLM with “Assume you are a journalist in a news agency. Rephrase following news article”, and subsequently inputting the original article. The second strategy, article extending, entailed determining the total word count of the original article and using the prompt “Assume you are a journalist in a news agency. Extend and complete following news article until it reaches a length of [N] words”. The placeholder [N] was substituted with the total word count of the original article, and the first third of the original article was then fed as input to be expanded by the LLM. The third strategy, summary expanding, involved a combination of summary and expansion. Initially, the prompt “Summarize following news article” was used to condense the original news article. Subsequently, the total word count of the original article was calculated. The resulting summary, along with the prompt “Assume you are a journalist in a news agency. Write a news article that comprises [N] words based on the following summarized news article”, was used to guide the LM in generating a final sample. To maintain length consistency, the placeholder [N] was substituted with the total word count of the original article prior to summarization. The dataset collected will be made available to the research community, enabling researchers to pursue further investigations on this field1.

1https://github.com/navid-aub/News-Dataset

4 METHODOLOGY

In this section, we employ principles derived from information theory [1], specifically utilizing mutual information, to derive a metric of term significance in news articles. This metric serves to quantify the level of uncertainty between the terms utilized within a news article and its attributed authorship, distinguishing between human-authored news and that produced by an LLM. Through the establishment of this metric, we ascertain the systematic ranking of terms predicated upon their discriminatory efficacy. Such systematic ranking underscores the pivotal terms that aid individuals in cultivating skepticism regarding the provenance of news articles they read, particularly in discerning whether they originate from an LLM or human sources. Initially, we demonstrate that a straightforward approach, namely term frequency-inverse document frequency (TF-IDF) [24] coupled with a basic classifier, like logistic regression, exhibits a remarkable capability to discern LLMgenerated news articles when it utilizes only a limited number of terms ranked based on the proposed metric. Therefore, by introducing and analyzing these terms, we anticipate a reduction in human susceptibility to falling into the trap of fake news propagated by LLMs. By familiarizing themselves with the characteristic terms present in LLM-generated news articles, individuals can enhance their discernment and accuracy in distinguishing between authentic and fabricated news.

4.1 Term awareness and uncertainty reduction

The quantification of information pertaining to an event with probability 𝑃(X=𝑥) is denoted by log( 1

𝑃(X=𝑥) ). Self-entropy, 𝐻(X), representing the extent of uncertainty pertaining to a random variable X, is determined by the expected amount of information associated to X,

𝐻(X) = ∑︁

𝑥∈X 𝑃(X=𝑥) log( 1

𝑃(X=𝑥) ) (1)

Let A be a random variable indicating the categorization of news article authorship, wherein two discrete outcomes are delineated: 𝑎1 representing human-authored news, and 𝑎2, denoting LLM authorship. Let W denote a random variable representing the constituent words within a news article. The permissible outcomes of W encompass the set of distinct lexical entities, 𝑤𝑖’s, drawn from the aggregate vocabulary derived from both human-written news articles and those generated by LLMs. The pairwise mutual information, 𝑀(𝑎𝑖,𝑤𝑗),concerning authorship 𝑎𝑖and word 𝑤𝑗is expressed as the discrepancy in information content between the joint probability 𝑃(A=𝑎𝑖, W=𝑤𝑗) and the product of the probabilities 𝑃(A=𝑎𝑖) and 𝑃(W=𝑤𝑗) under the assumption of independence between authorship 𝑎𝑖and word 𝑤𝑗,

𝑀(𝑎𝑖,𝑤𝑗) = log 𝑃(A=𝑎𝑖, W=𝑤𝑗)

𝑃(A=𝑎𝑖)𝑃(W=𝑤𝑗) (2)

The expected mutual information, denoted by 𝐼(A; W), serves as a metric quantifying the reduction in uncertainty with respect to authorship A consequent to the awareness of the words comprising

HT ’24, September 10–13, 2024, Poznan, Poland Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee

the news article,

𝐼(A; W) = ∑︁

∑︁

𝑎∈A 𝑃(𝑎,𝑤) log( 𝑃(𝑎,𝑤)

𝑃(𝑎)𝑃(𝑤) )

𝑤∈W

= ∑︁

∑︁

𝑎∈A 𝑃(𝑎)𝑃(𝑤|𝑎) log( 1

𝑃(𝑎) ) −𝑃(𝑤)𝑃(𝑎|𝑤) log( 1

𝑃(𝑎|𝑤) )

𝑤∈W

𝑃(𝑦) ) − ∑︁

𝑤∈W 𝑃(𝑤) ∑︁

= ∑︁

𝑎∈A 𝑃(𝑎|𝑤) log( 1

𝑎∈A 𝑃(𝑎) log( 1

𝑃(𝑎|𝑤) )

= 𝐻(A) −𝐻(A|W) (3)

Equation 3 validates that 𝐼(A; W) encapsulates the difference in entropy pertaining to authorship before and after focusing on attributes exclusively provided by the words contained within a news article. Consequently, the capacity of each word within the vocabulary to serve as an indicator of the authorship of a news article can be quantified and ranked accordingly. To achieve this objective, we represent the mutual information as the summation over all words,

𝑤𝑖∈W 𝑃(𝑤𝑖)  𝐻(A) −𝐻(A|W = 𝑤𝑖) 

𝐼(A; W) = ∑︁

def = ∑︁

𝑤𝑖∈W 𝐼𝑖(A; W = 𝑤𝑖) (4)

We term 𝐼𝑖as the Entropy-Shift Authorship Signature (ESAS). Utilizing the ESAS metric enables the prioritization of words in the vocabulary of news articles, highlighting the most crucial ones for human readers in distinguishing between human-written news and news generated by LLMs.

4.2 Estimating probabilities in ESAS metric equations

Let D=D𝐿∪D𝐻denote the collection of all news articles composed by both LLMs and humans, where D𝐿represents news articles generated by LLMs and D𝐻is the set of news articles authored by humans. The total word count within D is assumed to be N. The set of distinct words within D is represented as W, and the frequency of the 𝑖𝑡ℎword in W, denoted as 𝑤𝑖, within D is computed as 𝑁𝑖=𝑁𝐿,𝑖+𝑁𝐻,𝑖, where 𝑁𝐿,𝑖and 𝑁𝐻,𝑖denote the frequency of 𝑤𝑖in news articles generated by LLMs and those written by humans, respectively. The probability 𝑃(𝑤𝑖) represents the likelihood of occurrence of 𝑤𝑖in a news article and can thus be approximated as the ratio 𝑁𝑖

𝑁. To compute 𝐻(A), we assume a uniform probability distribution over the random variable A. Consequently, a news article is considered either human or LLM-generated with equal probability, 𝑃(A=𝑎1) = 𝑃(A=𝑎2)= 1

2, and 𝐻(A) is computed as log 2. In order to calculate the conditional entropy 𝐻(A|W=𝑤𝑖), it is necessary to compute the conditional probability 𝑃(𝑎𝑗|𝑤𝑖). This probability is determined by dividing the number of occurrences of 𝑤𝑖in authorship category 𝑎𝑗(𝑗= 1, 2) by the total occurrences of𝑤𝑖 in both authorship categories (LLM-generated and human-authored news articles).

𝐻(A|W=𝑤𝑖)= −𝑃(𝑎1|𝑤𝑖) log(𝑃(𝑎1|𝑤𝑖))−𝑃(𝑎2|𝑤𝑖) log(𝑃(𝑎2|𝑤𝑖))

=𝑁𝐿,𝑖

𝑁𝐿,𝑖 ) + 𝑁𝐻,𝑖

𝑁𝑖 log( 𝑁𝑖

𝑁𝑖 log( 𝑁𝑖

𝑁𝐻,𝑖 )

(5)

Thus, ESAS is derived as follows,

 1 + 𝑁𝐿,𝑖

𝑁𝑖 )  (6)

𝑁𝑖 log( 𝑁𝐿,𝑖

𝑁𝑖 ) + 𝑁𝐻,𝑖

𝑁𝑖 log( 𝑁𝐻,𝑖

𝐼𝑖(A; W = 𝑤𝑖) = 𝑁𝑖

𝑁

4.3 Evaluating the effectiveness of ESAS metric

To evaluate the effectiveness of ESAS metric, we initially divide the data into a 70% training set and a 30% testing set. Subsequently, we tokenize the news articles at the word level for both humanauthored news articles (HANA) and LLM-generated news articles (LGNA). The ESAS metric is then computed for each word within the combined vocabulary, followed by ranking them. By selecting the 𝑚words with the highest ESAS scores as the term vocabulary for the TF-IDF method, we train a binary classifier using logistic regression. Fig. 1 displays the results of this trained classifier on the testing set’s news articles across various LLMs and fake levels. The results suggest that the classifier achieves high accuracy consistently when only a limited number of words are present in the TF-IDF vocabulary. For 𝑚=10, an average accuracy of 92%, 84.9%, 84.6%, and 89.5% are achieved for ChatGPT, Llama2-7b, Llama2-13b, and Mistral-7b, respectively. Table 1 outlines the vocabulary sizes from the training data for every LLM and fake level. The “Common” column indicates the number of words shared between the vocabularies of HANA and LGNA. Conversely, the “HANA uncommon” and “LGNA uncommon” columns display the counts of words exclusive to the HANA and LGNA vocabularies, respectively. Based on the data represented in the table, a size of 𝑚=10 constitutes merely 0.022%, 0.022%, 0.022%, and 0.021% of the unique vocabulary sizes for ChatGPT, Llama2-7b, Llama2-13b, and Mistral-7b, respectively. Despite these minuscule percentages, the high accuracy achieved through the rudimentary classifier method confirm the ESAS metric’s effectiveness in prioritizing and selecting significant words for discerning the authorship of news articles. In addition, Fig. 1 illustrates varying levels of difficulty in predicting the labels for the three prompt strategies employed. The “Expanded summary” strategy proves to be the easiest to detect as it achieves higher accuracy for a specific number of selected words, 𝑚. Conversely, the “Rephrased” strategy emerges as the most challenging. Across all LLMs, the “Extended” strategy curve lies between the other two. This implies that our designed prompt strategies effectively generate distinct levels of fake.

5 RESULTS AND DISCUSSION

In this section, we extract cues from news articles based on unigrams, bigrams, and POS tagging across different LLMs in various scenarios. These cues are derived from the 10 most significant entities (unigram, bigram, or POS tagging) selected using the ESAS metric from the entire training set of news articles. We believe that by keeping these cues in mind while reading news articles on online social platforms, individuals can attain accuracy levels similar to the rudimentary classifier for identifying news articles which are likely generated by LLMs. In all scenarios in this section, the data is divided into training and testing sets with a ratio of 70% and 30%, respectively.

Seeing Through AI’s Lens HT ’24, September 10–13, 2024, Poznan, Poland

100

95

90

85

Accuracy

80

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 5 10 15 20 25 30 35 40 45

Number of selected words

(a) ChatGPT

100

95

90

85

80

Accuracy

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 5 10 15 20 25 30 35 40 45

Number of selected words

(c) Llama2-13b

100

95

90

85

80

Accuracy

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 5 10 15 20 25 30 35 40 45

Number of selected words

(b) Llama2-7b

100

95

90

85

Accuracy

80

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 5 10 15 20 25 30 35 40 45

Number of selected words

(d) Mistral-7b

Figure 1: The accuracy of TF-IDF in classifying the authorship of news articles based on the number of words selected for its vocabulary using ESAS metric.

Figure 1: The accuracy of TF-IDF in classifying the authorship of news articles based on the number of words selected for its vocabulary using ESAS metric.

5.1 Cues based on 10 most significant unigrams for different level of fake

We tokenize the news articles in the training set at the unigram level and apply the ESAS metric to all unique unigrams in the aggregated vocabularies of HANA and LGNA. The unigrams are sorted based on their ESAS scores, and the top 10 words with the highest scores are selected. Table 2 presents selected words in descending order with respect to ESAS score for each LLM and prompt strategy. For each word among these 10 selected words, we calculated its frequency ratio in HANA or LGNA relative to the word with the highest frequency among the words selected. The first number in parentheses represents this ratio for HANA, while the second number indicates the ratio for LGNA.

The table reveals that in 10 out of 12 instances, the first chosen word is “said”. The frequency of “said” in HANA is significantly higher than its frequency in LGNA. This indicates that the LLMs under examination tend to use the word “said” less frequently in their generated news articles, while human authors are more inclined to use it more often in reporting news. Comparing the results with the curves shown in Fig. 1 reveals that the presence or absence of the word “said” in a news article (𝑚=1) can achieve an accuracy of over 85% for news generated by ChatGPT across all prompt strategies. The ratios of 0, 0.02, and 0 for the three strategies in ChatGPT’s results for the word “said” confirm that ChatGPT is unlikely to use this word in its news articles. Similarly, the word “told” is more prevalent in HANA, featuring in 7 out of 12 instances, while its relative frequency ratio in LGNA is nearly zero. Additionally, the frequency of most selected words is greater

HT ’24, September 10–13, 2024, Poznan, Poland Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee

Table 1: The size of common and uncommon vocabulary for HANA and LGNA across different LLMs and fake levels

ChatGPT Rephrased 34414 9921 2168 Extended 23692 20643 2569 Expanded summary 25874 18461 2150

Llama2-7b Rephrased 27931 16181 615 Extended 25097 19015 2350 Expanded summary 23282 20830 1585

Llama2-13b Rephrased 28537 15798 683 Extended 26659 17676 2522 Expanded summary 22404 21931 1700

Mistral-7b Rephrased 33373 10962 2481 Extended 23412 20923 4753 Expanded summary 26827 17508 3708

LLM Prompt Common HANA uncommon LGNA uncommon

Table 2: Top 10 words by ESAS metric. The left and right numbers in parentheses indicate frequency ratios to the most frequent word among the 10 selected for HANA and LGNA, respectively.

LLM Prompt Words with highest ESAS value

Rephrased said(0.13,0.0), it(0.13,0.06), we(0.06,0.02), is(0.14,0.07), was(0.11,0.05), but(0.06,0.02), told(0.02,0.0), and(0.42,0.31), the(1.0,0.83), be(0.07,0.03)

ChatGPT

Extended said(1.0,0.04), he(0.76,0.25), we(0.49,0.14), at(0.66,0.24), after(0.26,0.05), told(0.14,0.0), you(0.29,0.06), was(0.83,0.4), these(0.09,0.31), because(0.14,0.01)

Rephrased the(1.0,0.58), said(0.13,0.01), of(0.44,0.24), to(0.46,0.25), in(0.36,0.18), it(0.13,0.04), we(0.06,0.01), and(0.42,0.26), was(0.11,0.04), that(0.2,0.11)

Llama2-7b

Extended said(0.23,0.12), and(0.77,1.0), significant(0.0,0.03), at(0.15,0.07), you(0.07,0.02), has(0.14,0.24), conclusion(0.0,0.02), he(0.18,0.1), an(0.11,0.06), address(0.01,0.03)

Expanded summary said(1.0,0.28), he(0.76,0.25), you(0.28,0.03), was(0.83,0.34), significant(0.02,0.18), at(0.64,0.27), conclusion(0.0,0.13), told(0.14,0.01), because(0.14,0.01), what(0.23,0.05)

Llama2-13b

Extended and(0.36,0.51), has(0.07,0.13), said(0.11,0.06), significant(0.0,0.02), he(0.08,0.04), you(0.03,0.01), the(0.86,1.0), at(0.07,0.04), however(0.0,0.02), com(0.01,0.0)

Rephrased said(0.13,0.01), the(1.0,0.67), it(0.13,0.04), to(0.46,0.29), of(0.43,0.27), in(0.36,0.22), that(0.2,0.1), is(0.14,0.06), he(0.1,0.04), we(0.06,0.02)

Mistral-7b

Extended said(0.23,0.06), at(0.15,0.06), he(0.18,0.08), you(0.07,0.02), and(0.77,1.0), was(0.19,0.1), because(0.03,0.0), told(0.03,0.0), after(0.06,0.02), year(0.06,0.02)

Expanded summary said(0.64,0.09), he(0.48,0.15), you(0.18,0.02), it(0.65,0.28), was(0.53,0.24), because(0.09,0.0), told(0.09,0.0), that(1.0,0.62), so(0.12,0.02), significant(0.01,0.1)

Expanded summary said(0.64,0.0), we(0.31,0.03), he(0.48,0.13), you(0.18,0.01), it(0.65,0.24), was(0.53,0.17), that(1.0,0.56), told(0.09,0.0), so(0.12,0.01), because(0.09,0.0)

Rephrased said(0.13,0.02), the(1.0,0.63), in(0.36,0.17), to(0.46,0.26), it(0.13,0.04), of(0.43,0.25), we(0.06,0.01), he(0.1,0.03), was(0.11,0.04), on(0.15,0.07)

Expanded summary said(0.36,0.07), he(0.27,0.07), at(0.23,0.08), has(0.21,0.42), you(0.1,0.02), it(0.36,0.18), was(0.3,0.14), told(0.05,0.0), in(1.0,0.7), significant(0.01,0.06)

in HANA compared to LGNA. However, there are some terms with notably higher frequencies in LGNA that can assist individuals in recognizing articles produced by LLMs, like “significant”, “address”, “conclusion”, and “however”. A noteworthy observation from the table is the abundance of stop words selected as the most significant words. While crafting straightforward rules for news readers based on stop words is challenging, their role in providing discriminative features for such

detection should not be overlooked. Therefore, detection systems that exclude stop words in preprocessing stages should assess this impact on their performance.

5.2 Cues based on 10 most significant unigrams for different news topics

Research has shown that the performance of AI detection systems deteriorates when the field or topic of the text changes. In this

Seeing Through AI’s Lens HT ’24, September 10–13, 2024, Poznan, Poland

Table 3: Top 10 words by ESAS metric for different news topics. The left and right numbers in parentheses indicate frequency ratios to the most frequent word among the 10 selected for HANA and LGNA, respectively.

LLM Topic Words with highest ESAS value

History said(0.65,0.0), he(0.67,0.2), we(0.24,0.02), was(0.51,0.16), that(1.0,0.53), within(0.03,0.2), told(0.09,0.0), you(0.1,0.0), these(0.05,0.23), ongoing(0.0,0.09)

Science said(0.66,0.0), we(0.28,0.03), it(0.69,0.25), that(1.0,0.46), you(0.13,0.0), was(0.3,0.08), so(0.11,0.0), potential(0.04,0.2), would(0.17,0.03), told(0.08,0.0)

ChatGPT

Society said(1.0,0.01), we(0.51,0.05), people(0.5,0.07), was(0.84,0.26), you(0.25,0.0), he(0.59,0.14), told(0.15,0.0), individuals(0.02,0.22), because(0.14,0.0), it(0.92,0.43)

Sports said(0.68,0.0), we(0.43,0.03), you(0.19,0.0), it(0.65,0.23), was(0.48,0.16), he(0.37,0.1), com(0.15,0.01), ap(0.13,0.01), that(1.0,0.53), because(0.11,0.0)

History said(0.98,0.27), he(1.0,0.32), was(0.77,0.31), significant(0.01,0.18), ongoing(0.0,0.13), towards(0.01,0.16), told(0.14,0.01), title(0.01,0.13), you(0.15,0.01), conclusion(0.0,0.11)

Politics said(0.61,0.17), he(0.59,0.17), was(0.51,0.15), significant(0.01,0.15), has(0.57,1.0), would(0.21,0.04), told(0.1,0.0), concerns(0.02,0.14), had(0.2,0.04), were(0.2,0.04)

Science said(1.0,0.21), was(0.45,0.11), potential(0.06,0.3), he(0.31,0.06), conclusion(0.0,0.14), you(0.2,0.02), significant(0.02,0.21), title(0.0,0.12), concerns(0.07,0.29), told(0.12,0.0)

Llama2-7b

Society said(1.0,0.31), he(0.59,0.15), was(0.84,0.34), you(0.25,0.02), title(0.0,0.14), told(0.15,0.01), conclusion(0.0,0.12), had(0.28,0.07), she(0.33,0.09), towards(0.02,0.16)

Sports said(1.0,0.32), com(0.22,0.0), you(0.28,0.02), ap(0.2,0.0), was(0.71,0.23), https(0.18,0.0), he(0.55,0.18), at(0.7,0.28), twitter(0.15,0.0), significant(0.01,0.17)

History said(0.37,0.07), he(0.38,0.1), has(0.19,0.39), in(1.0,0.67), was(0.29,0.13), at(0.21,0.08), significant(0.01,0.07), ongoing(0.0,0.05), told(0.05,0.0), what(0.07,0.01)

Politics said(0.55,0.1), he(0.53,0.13), was(0.46,0.15), has(0.51,1.0), significant(0.01,0.14), after(0.22,0.04), told(0.09,0.0), at(0.29,0.09), it(0.46,0.21), future(0.03,0.15)

Science said(0.97,0.11), he(0.3,0.05), was(0.43,0.13), at(0.49,0.17), significant(0.02,0.2), it(1.0,0.53), potential(0.05,0.26), what(0.2,0.03), told(0.11,0.0), has(0.49,0.91)

Llama2-13b

Society said(0.34,0.08), he(0.2,0.05), you(0.09,0.0), has(0.18,0.39), was(0.29,0.14), told(0.05,0.0), in(1.0,0.69), what(0.07,0.01), people(0.17,0.06), at(0.19,0.08)

Sports said(0.94,0.2), com(0.2,0.0), https(0.17,0.0), at(0.66,0.22), ap(0.19,0.0), he(0.51,0.14), has(0.45,1.0), you(0.26,0.03), was(0.67,0.25), because(0.15,0.01)

Politics said(1.0,0.14), he(0.97,0.26), was(0.85,0.34), significant(0.02,0.23), title(0.0,0.17), it(0.84,0.37), after(0.39,0.1), told(0.16,0.01), concerns(0.03,0.21), because(0.12,0.0)

Science said(0.66,0.07), it(0.69,0.29), potential(0.04,0.22), you(0.13,0.01), that(1.0,0.56), title(0.0,0.1), significant(0.02,0.13), he(0.21,0.05), told(0.08,0.0), so(0.11,0.01)

Mistral-13b

Society said(1.0,0.14), he(0.59,0.17), title(0.0,0.17), you(0.25,0.03), was(0.84,0.36), people(0.5,0.17), told(0.15,0.01), it(0.92,0.46), because(0.14,0.01), they(0.51,0.21)

Sports said(1.0,0.13), you(0.28,0.02), it(0.96,0.4), he(0.55,0.15), com(0.22,0.01), ap(0.2,0.01), was(0.71,0.27), https(0.18,0.01), because(0.16,0.0), that(1.0,0.57)

section, we investigate the impact of concentrating on a specific news topic on the selected unigrams for distinguishing between HANAs and LGNAs. To achieve this, at each step, we filter one topic and construct a training set comprising only that particular topic. Following this, we tokenize the news articles in the training set at the unigram level. The ESAS metric is applied to the aggregated vocabulary, and the unigrams are sorted according to their ESAS scores. We subsequently choose the 10 most significant unigrams. Due to limited space, we focus solely on the LGNA that are produced from the “Expanded Summary” strategy in this part. We opted for the “Expanded Summary” strategy over the other two strategies

Politics said(1.0,0.0), he(0.97,0.21), was(0.85,0.2), we(0.37,0.02), it(0.84,0.29), concerns(0.03,0.3), significant(0.02,0.26), landscape(0.01,0.22), after(0.39,0.07), told(0.16,0.0)

History said(0.65,0.1), he(0.67,0.2), was(0.51,0.22), title(0.01,0.09), told(0.09,0.01), ongoing(0.0,0.08), that(1.0,0.63), because(0.08,0.01), you(0.1,0.01), significant(0.01,0.09)

as we attained higher accuracy with TF-IDF using this particular prompt strategy. Hence, we have greater confidence in the terms selected using the ESAS metric in this scenario. Table 3 presents the top 10 words selected for each news topic across each LLM. In all cases, the word “said” achieves the highest ESAS score. Among the LLMs, Llama2-7b displays a stronger inclination towards using the word “said”, particularly in sports news. However, its frequency is still approximately three to four times less frequent than in HANA. Consequently, the presence of the word “said” remains a pivotal indicator that the news article is human-written. Similarly, as in Section 5.1, the words “told” and “significant” offer

(a) ChatGPT said the he said

HT ’24, September 10–13, 2024, Poznan, Poland Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee

she said

the ongoing

within the

said in said it he said

said the

going to

commitment to

at thehas been

in conclusion

said the he said

the ongoing

the need

has also

conclusion the

to address

to address

conclusion the

in conclusion

the need

has sparked

the ongoing

continues to the potential

said it

(b) Llama2-7b

importance of

commitment to

it was

(c) Llama2-13b said he

the ongoing

the potential the importance

said thehe said

(d) Mistral-7b

Figure 2: The word cloud of 10 most significant bigrams, with font size representing the relative ESAS score. Terms highlighted in green and red indicate higher relative frequencies in HANA and LGNA, respectively.

Figure 2: The word cloud of 10 most significant bigrams, with font size representing the relative ESAS score. Terms highlighted in green and red indicate higher relative frequencies in HANA and LGNA, respectively.

said he

valuable insights into the news source regardless of the topic. For the former, the HANA frequency surpasses that of LGNA, while for the latter, the frequency in LGNA exceeds that of HANA. From the table, it is evident that particular words emerge when focusing on individual topics. For instance, in sports-related news articles, tokens like “com” and “https” can achieve high ESAS scores. Their HANA frequency ratio significantly surpasses their LGNA frequency ratio. This suggests that the presence of a website link may indicate that the news article is authored by a human. In history-related news, the word “ongoing” is found in LGNA about five times more often than in HANA. In the context of science news, LLMs favor the use of “potential” over human authors, with a frequency approximately five times higher. Meanwhile, human authors use the word “people” more frequently in science articles. In political news, new terms do not show consistent patterns. For ChatGPT, Llama2-13b, and both Llama2-7b and Mistral-7b, the terms “landscape”, “future” and “concern” appear, respectively.

5.3 Cues based on 10 most significant bigrams

We tokenize the news articles in the training set at the bigram level and apply the ESAS metric to all unique bigrams in the aggregated vocabularies of HANA and LGNA. We sort the bigrams based on their ESAS scores and select the top 10 with the highest scores. Similar to Section 5.2, we focus on the “Expanded Summary” strategy. Fig. 2 presents the word cloud of selected bigrams for each LLM, with font size indicating their relative ESAS scores. Bigrams more

frequent in HANA are shown in green, while those more frequent in LGNA are displayed in red. Across all LLMs, bigrams containing the word “said” make the most significant contribution to identifying the authorship of a news article. Specifically, bigrams like “he said”, “said he”, “said it”, “said the”, “she said”, and “said in” are common in human-authored news writing. The only bigram consistently present across all LLMs for discerning LLM-generated news is “the ongoing”. Additionally, each LLM presents its unique key bigram cues for identification as LLM-generated. ChatGPT frequently employs the bigram “within the”, and “commitment to” which is also found in Mistral-7b. Additionally, Mistral-7b uses bigrams featuring the word “importance”, including “the importance” and “importance of”. While Llama2-7b and Llama2-13b share several common bigrams, such as “to address”, “the need” and “conclusion the” they also reveal unique bigrams not shared between them. In LGNA produced by Llama2-7b, bigrams like “continues to” and “has sparked” serve as strong indicators for identifying the news source as an LLM. Conversely, in LGNA generated by Llama2-13b, bigrams such as “has been”, and “has also” are crucial for distinguishing LLM-generated articles. The findings demonstrate that the complexity of developing a detection method is influenced by both the model and its parameter count. Consequently, to construct a more effective detection system, training should encompass diverse samples from multiple LLMs, and samples from each LLM should be drawn from various settings.

Seeing Through AI’s Lens HT ’24, September 10–13, 2024, Poznan, Poland

100

95

90

85

Accuracy

80

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 10 20 30 40 50 60 70 80 90 100

Number of selected POS bigrams

(a) ChatGPT

100

95

90

85

80

Accuracy

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 10 20 30 40 50 60 70 80 90 100

Number of selected POS bigrams

(c) Llama2-13b

100

95

90

85

80

Accuracy

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 10 20 30 40 50 60 70 80 90 100

Number of selected POS bigrams

(b) Llama2-7b

100

95

90

85

80

Accuracy

75

70

65

60

Rephrased Extended Expanded Summary

55

50

1 10 20 30 40 50 60 70 80 90 100

Number of selected POS bigrams

(d) Mistral-7b

Figure 3: The accuracy of TF-IDF in classifying the authorship of news articles based on the number of POS bigrams selected for its vocabulary using ESAS metric.

Figure 3: The accuracy of TF-IDF in classifying the authorship of news articles based on the number of POS bigrams selected for its vocabulary using ESAS metric.

5.4 Cues based on 10 most significant bigrams in POS tagging

Initially, we extract the POS tagging of sentences in each news article from our training set using Natural Language Toolkit (NLTK) POS tagging [19] in Python. Subsequently, we employ a sliding window that combines two consecutive tags into a single entity, referred to as POS bigrams. We apply the ESAS metric to all unique POS bigrams in the aggregated vocabularies of HANA and LGNA. The POS bigrams are then sorted based on their ESAS scores, and the top 10 POS bigrams with the highest scores are selected. Fig. 3 shows the accuracy of the TF-IDF combined with logistic regression classifier on the testing set’s news articles across different LLMs and levels of fake. While the classifier attains reasonably high accuracy, it is lower than the accuracy explored in Section 5.1.

In Fig. 4, the 10 most significant POS bigrams for the “Extended summary” prompt strategy are shown across different LLMs. The definitions for each tag are provided in Table 4. For almost all POS bigrams, human writers use each entity more often than LLMgenerated articles. The exception arises when the “POS” tag, representing a genitive marker, appears in the POS bigrams. All LLMs tend to use “POS NN” more frequently than human writers. The ChatGPT and Mistral-7b models also utilize the POS bigram “NNP POS” alongside “POS NN”. The genitive case usually denotes ownership or possession and is represented by an apostrophe followed by an “s” or simply an apostrophe (’). Hence, identifying the apostrophe used for possession can serve as a useful cue to be more skeptical about the possibility of the news article being generated by an LLM.

HT ’24, September 10–13, 2024, Poznan, Poland Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee

MD VB Human LLM ESAS value

PRP MD

NN NNP

VBD NNP

NNP NN

POS NN

NNP POS

PRP VBP

NNP VBD

PRP VBD

0.0 0.2 0.4 0.6 0.8 1.0

(a) ChatGPT

VBD IN Human LLM ESAS value

NN VBD

VBD DT POS NN

NN NN

NNP NN NN NNP

PRP VBD

NNP NNP NNP VBD

0.0 0.2 0.4 0.6 0.8 1.0

(c) Llama2-13b

NN NNP Human LLM ESAS value

VBD IN

NN VBD

POS NN

NN NN

VBD DT

NNP NN

PRP VBD

NNP NNP NNP VBD

0.0 0.2 0.4 0.6 0.8 1.0

(b) Llama2-7b

NN NN Human LLM ESAS value

NN NNP

PRP VBP

NNP NNP VBD NNP

POS NN

NNP NN

NNP VBD

NNP POS

PRP VBD

0.0 0.2 0.4 0.6 0.8 1.0

(d) Mistral-7b

Figure 4: The barplots of 10 most significant POS bigrams across different LLMs for “Extended summary” prompt strategy.

Figure 4: The barplots of 10 most significant POS bigrams across different LLMs for “Extended summary” prompt strategy.

Table 4: POS tag set

POS Tag Definition POS Tag Definition

DT determiner POS genitive marker IN preposition or conjunction, subordinating PRP pronoun, personal MD modal auxiliary VB verb, base form NN noun, common, singular or mass VBD verb, past tense NNP noun, proper, singular VBP verb, present tense, not 3rd person singular

6 CONCLUSION AND FUTURE WORK

In this study, we introduced the Entropy-Shift Authorship Signature (ESAS), a metric designed to rank terms and entities, such as POS tagging, within news articles based on their importance in distinguishing between human-written and LLM-generated news. Furthermore, we presented our collected news dataset comprising 39k news articles authored by humans or produced using four LLMs, i.e. ChatGPT, Llama2-7b, Llama2-13b, and Mistral-7b, across three levels of fake. We showcased the effectiveness of the proposed ESAS metric in in identifying significant indicators within the news articles. This was demonstrated by the impressive accuracy attained

by a basic method, namely TF-IDF combined with a logistic regression classifier, when provided with a limited set of terms with the highest ESAS scores. We analyzed and presented these top-ranked ESAS terms as straightforward cues that individuals can utilize to increase their skepticism towards LLM-generated fake news. One significant area yet to be explored in our research is the situations where imposters utilize LLMs to create fake news and then manipulate them before publishing. This poses a question about the potential consequences of manipulating LLM-generated fake news. We defer this matter to future studies.

Seeing Through AI’s Lens HT ’24, September 10–13, 2024, Poznan, Poland

Acknowledgments

Research was supported in part by grant ARO W911NF-20-1-0254. The views and conclusions contained in this document are those of the authors and not of the sponsors. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein.

References

[1] Akiko Aizawa. 2003. An information-theoretic perspective of tf–idf measures. Information Processing & Management 39, 1 (2003), 45–65. https://doi.org/10. 1016/S0306-4573(02)00021-3

[2] Miguel A Alonso, David Vilares, Carlos Gómez-Rodríguez, and Jesús Vilares. 2021. Sentiment analysis for fake news detection. Electronics 10, 11 (2021), 1348. https://doi.org/10.3390/electronics10111348

[3] Matin Amoozadeh, David Daniels, Daye Nam, Aayush Kumar, Stella Chen, Michael Hilton, Sruti Srinivasa Ragavan, and Mohammad Amin Alipour. 2024. Trust in Generative AI among students: An exploratory study. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (Portland, OR, USA). 67–73. https://doi.org/10.1145/3626252.3630842

[4] Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee. 2023. The looming threat of fake and llm-generated linkedin profiles: Challenges and opportunities for detection and prevention. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy). 1–10. https://doi.org/10.1145/3603163.3609064 — ACM HyperText copy

[5] Amrita Bhattacharjee, Raha Moraffah, Joshua Garland, and Huan Liu. 2024. EAGLE: A Domain Generalization Framework for AI-generated Text Detection. arXiv preprint arXiv:2403.15690 (2024).

[6] Anshika Choudhary and Anuja Arora. 2021. Linguistic feature based learning model for fake news detection and classification. Expert Systems with Applications 169 (2021), 114171. https://doi.org/10.1016/j.eswa.2020.114171

[7] Limeng Cui, Suhang Wang, and Dongwon Lee. 2019. Same: sentiment-aware multi-modal embedding for detecting fake news. In Proceedings of the 2019 IEEE/ACM international conference on advances in social networks analysis and mining (Vancouver, British Columbia, Canada). 41–48. https://doi.org/10.1145/ 3341161.3342894

[8] Anastasia Giachanou, Bilal Ghanem, Esteban A Ríssola, Paolo Rosso, Fabio Crestani, and Daniel Oberski. 2022. The impact of psycholinguistic patterns in discriminating between fake news spreaders and fact checkers. Data & knowledge engineering 138 (2022), 101960. https://doi.org/10.1016/j.datak.2021.101960

[9] Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794 (2022).

[10] Beizhe Hu, Qiang Sheng, Juan Cao, Yuhui Shi, Yang Li, Danding Wang, and Peng Qi. 2024. Bad actor, good advisor: Exploring the role of large language models in fake news detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 22105–22113. https://doi.org/10.1609/aaai.v38i20.30214

[11] Linmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi, Nan Duan, and Ming Zhou. 2021. Compare to the knowledge: Graph neural fake news detection with external knowledge. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 754–763. https://doi.org/10.18653/v1/2021.acl-long.62

[12] Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2019. Automatic detection of generated text is easiest when humans are fooled. arXiv preprint arXiv:1911.00650 (2019).

[13] Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023).

[14] Bohan Jiang, Zhen Tan, Ayushi Nirmal, and Huan Liu. 2024. Disinformation detection: An evolving challenge in the age of llms. In Proceedings of the 2024 SIAM International Conference on Data Mining (SDM). SIAM, 427–435. https: //doi.org/10.1137/1.9781611978032.50

[15] Waleed Kareem and Noorhan Abbas. 2023. Fighting lies with intelligence: Using large language models and chain of thoughts technique to combat fake news. In International Conference on Innovative Techniques and Applications of Artificial Intelligence. Springer, 253–258. https://doi.org/10.1007/978-3-031-47994-624

[16] John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A watermark for large language models. In International Conference on Machine Learning. PMLR, 17061–17084.

[17] Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang. 2023. Robust distortion-free watermarks for language models. arXiv preprint arXiv:2307.15593 (2023).

[18] Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2024. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems

36 (2024).

[19] Edward Loper and Steven Bird. 2002. Nltk: The natural language toolkit. arXiv preprint cs/0205028 (2002).

[20] Artidoro Pagnoni, Martin Graciarena, and Yulia Tsvetkov. 2022. Threat scenarios and best practices to detect neural fake news. In Proceedings of the 29th International Conference on Computational Linguistics. 1233–1249.

[21] Jeff Z Pan, Siyana Pavlova, Chenxi Li, Ningxi Li, Yangmei Li, and Jinshuo Liu. 2018. Content based fake news detection using knowledge graphs. In The Semantic Web–ISWC 2018: 17th International Semantic Web Conference, Monterey, CA, USA, October 8–12, 2018, Proceedings, Part I 17. Springer, 669–683. https://doi.org/10. 1007/978-3-030-00671-639

[22] Xiao Pu, Mingqi Gao, and Xiaojun Wan. 2023. Summarization is (almost) dead. arXiv preprint arXiv:2309.09558 (2023).

[23] Kristina Radivojevic, Nicholas Clark, and Paul Brenner. 2024. LLMs Among Us: Generative AI Participating in Digital Discourse. arXiv preprint arXiv:2402.07940 (2024).

[24] Juan Ramos et al. 2003. Using tf-idf to determine word relevance in document queries. In Proceedings of the first instructional conference on machine learning, Vol. 242. Citeseer, 29–48.

[25] Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. 2017. Truth of varying shades: Analyzing language in fake news and political fact-checking. In Proceedings of the 2017 conference on empirical methods in natural language processing. 2931–2937. https://doi.org/10.18653/v1/D17-1317

[26] Noureddine Seddari, Abdelouahid Derhab, Mohamed Belaoued, Waleed Halboob, Jalal Al-Muhtadi, and Abdelghani Bouras. 2022. A hybrid linguistic and knowledge-based analysis approach for fake news detection on social media. IEEE Access 10 (2022), 62097–62109. https://doi.org/10.1109/ACCESS.2022.3181184

[27] Jinyan Su, Terry Yue Zhuo, Jonibek Mansurov, Di Wang, and Preslav Nakov. 2023. Fake news detectors are biased against texts generated by large language models. arXiv preprint arXiv:2309.08674 (2023).

[28] Yanshen Sun, Jianfeng He, Limeng Cui, Shuo Lei, and Chang-Tien Lu. 2024. Exploring the Deceptive Power of LLM-Generated Fake News: A Study of RealWorld Detection Challenges. arXiv preprint arXiv:2403.18249 (2024).

[29] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023).

[30] Jiaying Wu and Bryan Hooi. 2023. Fake News in Sheep’s Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks. arXiv preprint arXiv:2310.10830 (2023).

[31] Binwei Yao, Ming Jiang, Diyi Yang, and Junjie Hu. 2023. Empowering LLM-based machine translation with cultural awareness. arXiv preprint arXiv:2305.14328 (2023).

[32] Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. Advances in neural information processing systems 32 (2019).

[33] Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023).

[34] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2024).

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime