Disambiguation of Implicit Scientific References on X
Authors
Salim Hafid — médialab, Sciences Po, Paris, Paris, France and LIRMM, CNRS, University of Montpellier, Montpellier, France salim.hafid@sciencespo.fr; ORCID: 0000-0002-1775-8542
Yavuz Selim Kartal — GESIS - Leibniz Institute for the Social Sciences, Cologne, Germany YavuzSelim.Kartal@gesis.org; ORCID: 0000-0002-2146-2680
Sebastian Schellhammer — GESIS - Leibniz Institute for the Social Sciences, Cologne, Germany sebastian.schellhammer@gesis.org; ORCID: 0009-0001-6413-5823
Vetea Jacot — LIRMM, CNRS, University of Montpellier, Montpellier, France vetea.jacot@lirmm.fr; ORCID: 0009-0009-0194-6061
Sandra Bringay — LIRMM, CNRS, University of Montpellier, Montpellier, France and University of Paul-Valéry Montpellier 3, Montpellier, France sandra.bringay@lirmm.fr; ORCID: 0000-0002-2830-3666
Stefan Dietze — GESIS - Leibniz Institute for the Social Sciences, Heinrich-Heine-University, Cologne, Germany Stefan.Dietze@gesis.org; ORCID: 0009-0001-4364-9243
Konstantin Todorov — LIRMM, CNRS, University of Montpellier, Montpellier, France konstantin.todorov@lirmm.fr; ORCID: 0000-0002-9116-6692
Abstract
Original PDF: Download source PDF
Scientific discourse on the social web has been shown to compromise the accuracy of scientific findings. Complex scientific claims are uttered in the form of short snippets with "implicit references" (seen as references to scientific publications where the URLs to the actual studies are never cited). This has led to uninformed online scientific debates on topics such as health pandemics or climate. To enhance social media content, we introduce in this paper the novel task of disambiguation of implicit scientific references, where the goal is to retrieve the original scientific publications implicitly referred to by social media users. We contribute the first formalization, ground-truth corpus, and baselines for the task. With this work, we aim at shaping an understanding of implicit references on social media, and at laying a foundation for developing and evaluating methods for the disambiguation of implicit references.
Introduction
Scientific web discourse, i.e. science-related discourse on the social web, can involve various harmful phenomena such as polarization, mis-, dis- and malinformation [[17], [18]]. It has thus been increasingly studied across a range of disciplines [[2], [4], [13]]. Implicit scientific references, i.e. references made without explicit links to studies (see Figure [1]), are commonly used in scientific web discourse. While scientific references are often observed in online debates, formal citation standards are not widely adopted, in contrast to academic writing. Thus, informal and implicit references which do not enable access to the publications from which claims and findings originate are common practice. However, access to the original scientific publications cited by social media users is crucial to keeping scientific web discourse informed, and thus to avoiding the aforementioned harmful phenomena. This makes the task of disambiguating implicit scientific references challenging yet necessary.
Existing work in this context has focused on either detecting texts which ought to cite a publication [[9], [21]], or on retrieving original publications cited by news articles [[12], [16], [19]]. To our knowledge, no work exists on disambiguating implicit scientific references from the social web. In order to address this challenge, we contribute: (1) The first task formalization, ground-truth corpus (TweetCite), and baselines for implicit scientific reference disambiguation on X (previously Twitter);[<sup>1</sup>] (2) Insights on the prevalence and distribution of implicit and explicit scientific references on X.
Figure 1. Illustration of the task of disambiguation of implicit references, seen as a retrieval task. Compared to explicit references (left) which contain links to the actual publications, implicit references (right) contain no such links, thus making their disambiguation both challenging and crucial to keeping online discussions accurate and informed.
Definitions and Task Formalization
We define two types of scientific references (illustrated in Figure [1]).
(1) Explicit reference: defined as a micropost (e.g., a tweet) which contains a direct URL to a research publication.[<sup>2</sup>]
(2) Implicit reference: defined as a micropost (e.g., a tweet) which mentions a research publication but contains no URLs.
Based on these definitions, we formalize the task of implicit reference disambiguation as a specific instantiation of the entity linking problem,[<sup>3</sup>] where the surface mention is an implicit reference and the mentioned entity is a particular scientific publication. Given an implicit reference to a scientific publication, we aim at retrieving the mentioned publication from a pool of candidate publications.
Preliminary Data Analysis
In this section, we aim at motivating the relevance of implicit references by showing their prevalence on social media. We choose X for its popularity for science dissemination, as existing work has shown that more than one-third of recent publications get at least one citation on X [[6]]. We retrieve explicit and implicit references on X of COVID-19-related topics and compare their temporal distributions. We choose COVID-19 as our overarching topic because of its shown impact on scientific online discourse during the pandemic [[1], [11], [15]].
An initial pool of explicit and implicit references is retrieved as follows: For explicit references, we use Altmetric data[<sup>4</sup>] to retrieve the tweet IDs for all tweets explicitly referencing publications from the CORD-19 corpus, a corpus of over one million COVID-19-related academic publications [[20]]. Using these IDs, we extract the corresponding tweets from the 1% sample tweet archive underlying TweetsKB [[5]]. This process yields an initial pool of 138,702 explicit references, covering the period between 2020-01-01 (around the start of the pandemic) and 2022-09-23 (the most recent date in the Altmetric data at hand).
For implicit references, we first train a Twitter-RoBERTa-based model[<sup>5</sup>] on the SciTweets dataset [[10]], which contains both explicit and implicit scientific references from X (category 2). We then apply the model to predict all tweets from the 1% sample tweet archive in the same time frame as for the explicit references. We keep tweets that contain scientific references (category 2 score ≥ 0.5) but do not contain URLs in their text to ensure that the references contained in tweets are implicit, resulting in an initial pool of 393,327 implicit references.[<sup>6</sup>]
To identify key topics, we embed publication titles from the CORD-19 corpus using the SPECTER model [[3]] and extract topic clusters using HDBSCAN[<sup>7</sup>]. For the 20 largest topic clusters, based on the number of tweets containing explicit references, we use BERTopic [[8]] to extract the top keyword from the corresponding tweets. Using the top keyword per topic, we perform keyword-based matching to extract topic-specific tweets from our initial pools of explicit and implicit references, and qualitatively compare their temporal distributions by plotting a 5-day moving average (see Figure [2] for an example of such plot on the topics "masks" and "ivermectin"). The figure shows that both reference types have similar temporal distributions and that implicit references are, on average, more numerous than explicit ones on a given topic. We quantitatively verify this by computing Pearson correlations between the two distributions for each of the top 5 topics (see Table [1]). Results show moderate to strong correlations and confirm that, across topics, implicit references are, on average, more numerous than explicit ones.[<sup>8</sup>]
These preliminary results motivate our task of implicit reference disambiguation by showing that these references are prevalent on X. Further, the similarity between the temporal distributions of explicit and implicit references leads to the assumption that they refer to the same publications on similar dates. Since we know the specific publications being explicitly referenced, we follow this assumption to retrieve publication candidates for implicit references to label these pairs in the next section.
Figure 2. Temporal distributions of two different types of scientific references (explicit vs. implicit) on two example topics from the CORD-19 corpus.
Table 1:. Top-5 Pearson correlations (5-day moving average) between explicit and implicit references on specific topics. On average, implicit references of a given topic are 1.79 times more numerous than explicit references of the same topic.
Topic-keyword | #implicit references | #explicit references | correlation |
|---|---|---|---|
coronavirus | 20,402 | 7,268 | 0.64 |
remdesivir | 1,213 | 741 | 0.47 |
masks | 5,754 | 5,257 | 0.42 |
ivermectin | 3,249 | 2,114 | 0.38 |
vaccine | 30,452 | 16,440 | 0.33 |
avg-ratio | 1.79 | ||
max-ratio | 2.81 |
Constructing the TweetCite Corpus
In this section, we detail how we built a ground truth dataset containing tweet-publication pairs. TweetCite comprises both explicit references from X (where the URLs are deleted to approximate implicit references), and implicit references from X, each linked to publications from CORD-19.
We collect scientific publications from CORD-19, where each data point contains the title, abstract, and metadata of a publication related to COVID-19 research. The initial corpus comprises 1,056,660 academic publications. We keep only metadata records that contain both a title and an abstract to have enough context to learn a link between a scientific reference and its source, resulting in a collection of 821,005 publications, which we refer to as C.
Sampling Explicit References
To collect explicit references, we use the Altmetric corpus combined with the TweetsKB archive. We collect tweets that contain explicit publication references to C, where URLs will be removed from the query set. Next, we applied various filtering mechanisms. First, we select tweets containing a scientific claim using the SciTweets classifier [[10]] (category 1 score ≥ 0.9). This step ensures that the IR task remains realistic by filtering out tweets whose text is irrelevant to the referenced publications. For instance, we want to keep tweets such as: "Stanford study published in Cell shows covid vaccines work for children under ten: https://pubmed...", but filter out tweets such as: "Check out this paper https://arxiv.org/891..". Second, we drop duplicates, i.e., tweets with the same text, keeping only one of them. Third, we select tweets containing only one URL to ensure the claim is unambiguously based on the referenced academic publication. Finally, we keep all tweets in which the URL is at the end of the text so that removing the URL does not interfere with the syntax of the remaining text. Our final collection of explicit references comprises 14,253 tweet-publication pairs, where tweets contain scientific claims and point to 7,718 academic publications from C. We divide the explicit references into a train set and a dev set as shown in Table [2].
Sampling Implicit References
To collect implicit references, we target specific popular publications from C. First, we select 34 specific publications based on (i) topic (we pick top-topics from our previous list, e.g., ivermectin, remdesivir, masks, vaccine), and (ii) popularity (we pick popular publications in C based on Altmetric’s attention score). Then, for each publication, we sample tweets from TweetsKB that were published in a time window of 5 days around peak dates of explicit references of that publication (a peak date is seen as the top 3 days with the highest numbers of explicit references of a given publication). We assume that implicit references of a given publication are more likely to occur on dates where explicit references of that same publication are at their highest (see Section [3]), i.e., dates where a given publication is most discussed in online discourse. We use the SciTweets classifier to select tweets which contain references (category 2, predicts both explicit and implicit references), and we filter out tweets which contain URLs to ensure that sampled tweets contain implicit references. This results in a pool of 46,034 tweets with implicit references. To ensure that sampled tweets are likely to be implicit references of our targeted publications, we perform keyword extraction followed by keyword matching, where we select one keyword per publication based on its title and abstract, and filter candidate tweets which contain that keyword. Keyword extraction was performed manually. The list of keywords is made available on Github. This results in 1409 candidate tweet-publication pairs, where tweets implicitly reference 34 publications from C. To ensure that tweets reference the actual publications, we manually annotate them using a protocol which we describe next.
Data Annotation
An annotation protocol was constructed and used by five annotators in total to link implicit references to specific publications. The protocol is made publicly available. Annotators include three authors from this paper, as well as two Computer Science students (PhD/Bachelor). Annotators were presented with candidate tweet-publication pairs, where each pair included the publication’s title, abstract, publication date, author affiliations and publication venue, and the candidate tweet’s text and publication date. For each candidate pair, annotators assessed whether the implicit reference corresponds to the specific publication in the pair, responding with either "Yes", "No", or "I don’t know". Three separate annotation rounds were held: the first involved 99 candidate pairs (3 unique publications) annotated by 3 annotators, the second included 335 candidate pairs (8 unique publications) annotated by 4 annotators, and the third involved 840 candidate pairs (23 unique publications) annotated by 3 annotators. We measured the inter-annotator agreement by computing Fleiss Kappa κ [[7]] with the labels of all annotators to evaluate the annotation quality, resulting in an average agreement score of 0.57 across all rounds (seen as close to substantial agreement [[14]]). This score is in line with similar tasks such as detecting whether a social media post contains a scientific reference [[10]] (κ =0.63) or a scientific claim [[10]] (κ =0.61). Note that our task is more difficult given that it goes beyond detecting claims/references by asking the annotator to determine the referenced publication. We thus estimate our results to be encouraging. Moreover, given that our task is a retrieval task, the final TweetCite corpus only contains the subset of annotated pairs where the outcome was "Yes" by majority vote. In TweetCite, 43.6% of the pairs were answered with full agreement, indicating the high quality of our annotations. Finally, labels for hard cases (majority vote) and easy cases (full agreement) are included in our data to enable experiments on even higher quality data subsets.
Corpus Statistics
We show the obtained corpus in Table [2]. Our corpus consists of tweet-publication pairs, where tweets reference publications from CORD-19. Tweets contain two distinct reference types: explicit references, and manually annotated implicit references. The data is made available on Github.
Table 2:. TweetCite corpus statistics[<sup>9</sup>]
Data set | Reference-type | #tweets | #referenced publications |
|---|---|---|---|
train | Explicit | 12,853 | 6,946 |
dev | Explicit | 1,400 | 772 |
test | Implicit | 146 | 24 |
Total | 14,399 | 7,718 |
Baselines & results
We evaluate the novel IR task of reference disambiguation on social media on two query sets: explicit references (dev set), and implicit references (test set). The collection set is C reduced to the publications who are referenced by at least one tweet from either of the train/dev/test sets, resulting in a pool of 7,718 scientific publications (see Table [2]). The explicit URLs are deleted from all explicit references. The explicit references set represents an approximation of implicit references at a large scale (see Section [3]), while the implicit references set represents the real-world scenario. We evaluate using the Mean Reciprocal Rank (MRR@5) score. Our baselines include both zero-shot and fine-tuned models. The first zero-shot baseline is BM25-based, where the system ranks CORD-19 publications by term-frequency. The publications’ titles and abstracts are used to construct the BM25 corpus. The second zero-shot baseline relies on a Sentence-Transformer model (SBERT),[<sup>10</sup>] which measures pairwise cosine similarities between embeddings of the tweet text and a pool of candidate scientific publications (titles), and predicts the top-5 publications based on similarity. A random baseline was also computed, but given the size of the collection set, the score was MRR@5=0.00, further underlying the difficulty of the task. For fine-tuned baselines, we evaluated several sentence-transformer models. Similar to the zero-shot version, we used cosine similarity to compare embeddings of tweets’ texts to publications, this time using both the publications’ title and the last 10 sentences of the abstracts. Fine-tuning parameters and implementation details are provided in Appendix [A]. Fine-tuned baselines were first trained on the train set, then evaluated on the dev and test sets (separately), while zero-shot baselines were directly evaluated on the dev and test sets. Finally, large language models (LLM) baselines were also considered (LLaMa3 and ChatGPT) but did not yield evaluable results (they either hallucinated the target publications to be retrieved, or did not process the text at all due to too large context-windows. We discuss the LLM setting further in Appendix [C]).
Results in Table [3] show that the BM25 baseline outperforms the SBERT baseline on both evaluation sets, reaching MRR@5=0.55 on the explicit set (where the harmonic mean-rank of the correct retrieved publication=1.81) and MRR@5=0.25 on the implicit set (harmonic mean-rank=4). For the fine-tuned baseline, we show the performance of the best performing model e5, which is a sentence-transformer pre-trained using weakly-supervised contrastive learning.[<sup>11</sup>]. The complete list of evaluated models can be found in Appendix [B]. Results show that the fine-tuned e5 model largely outperforms both zero-shot baselines, with scores of MRR@5=0.66 (harmonic mean-rank=1.52) and MRR@5=0.47 (harmonic mean-rank=2.13) on the explicit and implicit sets (respectively). Considering the difficulty of the task, we consider these baseline results to be encouraging. Results also show that overall, scores on the implicit set are lower than on the explicit set. We hypothesize that the difference in performance is due to explicit references more often containing parts of the publication’s title in the tweet’s text, thus making the retrieval task easier. We found that 26% of explicit references (dev set) contained part of the full publication title, while only 3% of implicit references (test set) verified the same condition.[<sup>12</sup>] Overall, results show that, both when using explicit references as an approximation (with deleted URLs) or when using implicit references, the task is challenging and can benefit from the development of more robust methods and models.
Table 3:. MRR@5 baseline results on two sets of tweets: explicit references (dev set) and implicit references (test set).
Baseline | Training | Explicit | Implicit |
|---|---|---|---|
BM25 | None | 0.55 | 0.25 |
SBERT | None | 0.46 | 0.17 |
e5 | Fine-tuned | 0.66 | 0.47 |
#tweets | 1,400 | 146 |
Conclusion and Future Work
Scientific web discourse, as seen on social media, has been shown to compromise the accuracy of scientific findings. This phenomenon has led to online scientific debates being uninformed and has conduced to controversy and polarization. Examples include online debates about health pandemics or climate change. Access to original studies implicitly referenced by social media users is crucial to keeping such debates accurate and informed. In this work, we contributed the first task formalization, ground-truth corpus (TweetCite), and baselines for implicit scientific reference disambiguation on X. We also provided insights on the prevalence and distribution of both implicit and explicit references on X. In future work, we plan to extend this corpus to other platforms where scientific research is discussed (e.g., Mastodon or Bluesky). We also plan on upscaling the number of implicit references in our data, by automating the identification and sampling of implicit references, where a promising but understudied research direction is the retrieval of "study-identifiers", seen as specific n-grams used by social media users to refer to specific popular publications.
Fine-tuning parameters
We ran all computations on a Kaggle Notebook using a Tesla P100-PCIE-16GB GPU. All models were fine-tuned with the same GPU for 10 epochs. We used the MultipleNegativesRankingLoss loss function with a batch size of 16 and all other parameters at their default value: the Adamw optimizer, a learning rate of 5e+5, no warmup steps, and a weight decay of 0. For a given tweet from the query set, the positive examples for the loss function are the corresponding referenced scientific publications, while the negatives examples are scientific publications that are not referenced by the tweet. Fine-tuning took 49min for the best performing model (e5). Once fine-tuned, generating embeddings on the dev and test sets took 2ms per query, and measuring the cosine similarity between a (tweet, publication) pair took 20ms. Generating embeddings for all test queries took about 35s.
List of evaluated fine-tuned baselines
For the fine-tuned baselines, the best performing model was e5, whose results are reported in Section [5]. Remaining models who were evaluated in either zero-shot or fine-tuned but performed less well are the following: ModernBERT, Sentence-T5, SciBERT-MultiNLI, Specter, Twitter4SSE, SciBERT, Longformer, BERTweet, BGE, e5, PubMedBERT, MPNet, LLaMa-embedding, DistilRoBERTa, MiniLM, DistilBERT-covid, PubMed-embeddings, BioLORD.
LLM prompts and limitations
We considered both LLaMa3 and ChatGPT as LLM-baselines in both zero-shot and few-shot-learning settings. The prompt template we used was the following: "PUBLICATIONS: [publicationtitles] LABELS: [labels] INSTRUCTION: Analyze the following text: [querypost] and identify the top 5 LABELS from the list above that are most likely mentioned in the text. Present the results in a descending order list: [top1, top2, top3, top4, top5]. Generate only a LABEL list and do not modify LABELS."
In both cases, models did not yield evaluable results: they either hallucinated the target publications to be retrieved, or did not process the text at all due to too large context-windows. To reduce the size of the context window, we evaluated an alternative hybrid method where, for each query post, the pool of candidate publications is first reduced to the top 100 candidates via BM25, then only these top 100 are served to the LLM as candidates. This hybrid method is resource-intensive as it requires iterative manual formatting of the prompts served to the LLM. It was thus only evaluated using ChatGPT 4o-mini on the test set (146 tweets), where it performed MRR@5=0.32. This performance is better than the strictly BM25-based baseline (MRR@5=0.25) but still performs lower than the best fine-tuned model e5 (MRR@5=0.47, harmonic mean-rank=1.52).
Distribution of attention scores in the CORD-19 corpus
We show in Figure [3] the distribution of scientific publications by attention score (as given by Altmetric) in the full CORD-19 corpus (1,056,660 academic publications). The plot shows that the vast majority of publications receive very little attention (i.e., posts and reactions on social media), while a minority of publications are extremely popular and are highly debated on social media. We found that the least popular studies are referenced by 1.61 tweets (per study) on average, while the most popular studies are referenced by 218.63 tweets (per study) on average, further underlying the discrepancy between popular and unpopular studies on Twitter. We hypothesize that, specifically for such highly popular (i.e, highly debated/controversial/societally relevant) studies, the use of informal references is high. In this scenario, informal references would be likely to be used as a placeholder for studies that matter societally. This provides additional motivation, both for our empirical task presented in this paper, as well as for further data analysis towards a better understanding of the distribution of informal references on social media, specifically for highly popular publications.
Figure 3. Distribution of attention scores in the CORD-19 corpus.
References
[1]Mohammad Al-Ramahi, Ahmed Elnoshokaty, Omar El-Gayar, Tareq Nasralah, and Abdullah Wahbeh. 2021. Public discourse against masks in the COVID-19 era: Infodemiology study of Twitter data. JMIR Public Health and Surveillance (2021).
[2]Michael Brüggemann, Ines Lörcher, and Stefanie Walter. 2020. Post-normal science communication: exploring the blurring boundaries of science and journalism. Journal of Science Communication (2020).
[3]Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020. SPECTER: Document-level Representation Learning using Citation-informed Transformers. In ACL. 2270–2282.
[4]Sharon Dunwoody. 2021. Science journalism: Prospects in the digital age. In Routledge handbook of public communication of science and technology. Routledge.
[5]Pavlos Fafalios, Vasileios Iosifidis, Eirini Ntoutsi, and Stefan Dietze. 2018. Tweetskb: A public and large-scale rdf corpus of annotated tweets. In ESWC. Springer, 177–190.
[6]Zhichao Fang, Rodrigo Costas, Wencan Tian, Xianwen Wang, and Paul Wouters. 2020. An extensive analysis of the presence of altmetric data for Web of Science publications across subject fields and research topics. Scientometrics 124, 3 (2020), 2519–2549.
[7]Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin (1971).
[8]Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:https://arXiv.org/abs/2203.05794 (2022).
[9]Salim Hafid, Wassim Ammar, Sandra Bringay, and Konstantin Todorov. 2024. Cite-worthiness Detection on Social Media: A Preliminary Study. In NSLP. Springer Nature.
[10]Salim Hafid, Sebastian Schellhammer, Sandra Bringay, Konstantin Todorov, and Stefan Dietze. 2022. SciTweets-A Dataset and Annotation Framework for Detecting Scientific Online Discourse. In CIKM. 3988–3992.
[11]Michael Robert Haupt, Jiawei Li, and Tim K Mackey. 2021. Identifying and characterizing scientific authority-related misinformation discourse about hydroxychloroquine on twitter using unsupervised machine learning. Big Data & Society (2021).
[12]Kayvan Kousha and Mike Thelwall. 2019. An automatic method to identify citations to journals in news stories: A case study of uk newspapers citing web of science journals. Journal of Data and Information Science (2019).
[13]SE Kreps and DL Kriner. 2020. Model uncertainty, political contestation, and public trust in science: evidence from the COVID-19 pandemic. Sci. Adv. 6, eabd4563.
[14]J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977), 159–174.
[15]Enrique Prada, Andrea Langbecker, and Daniel Catalan-Matamoros. 2023. Public discourse and debate about vaccines in the midst of the covid-19 pandemic: A qualitative content analysis of Twitter. Vaccine (2023).
[16]James Ravenscroft, Amanda Clare, and Maria Liakata. 2018. HarriGT: Linking news articles to scientific literature. In ACL. 19.
[17]Yasmim Mendes Rocha, Gabriel Acácio de Moura, Gabriel Alves Desidério, Carlos Henrique de Oliveira, Francisco Dantas Lourenço, and Larissa Deadame de Figueiredo Nicolete. 2021. The impact of fake news on social media and its influence on health during the COVID-19 pandemic: A systematic review. Journal of Public Health (2021).
[18]Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science (2018).
[19]Jun Wang and Bei Yu. 2021. News2PubMed: A Browser Extension for Linking Health News to Medical Literature. In SIGIR. 2605–2609.
[20]Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, Yunyao Li, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Nancy Xin Ru Wang, Christopher Wilhelm, Boya Xie, Douglas M. Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020. CORD-19: The COVID-19 Open Research Dataset. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020. ACL.
[21]Dustin Wright and Isabelle Augenstein. 2021. CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding. In ACL-IJCNLP.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime