Publishing, linking and translating news in multilingual communities: a mirror of cultural differences?
Giuseppe Carrino (University of Bologna, Bologna, Italy) · Angelo Di Iorio (University of Bologna, Bologna, Italy) · Davide Picca (University of Lausanne, Lausanne, Switzerland)
All authors contributed equally to this research.
Published in HT '24: 35th ACM Conference on Hypertext and Social Media · DOI: 10.1145/3648188.3675143 · License: CC BY 4.0
Authors: Giuseppe Carrino, Angelo Di Iorio, Davide Picca
Keywords: Multiculture, News analysis, TARO
Session: Critical Reading
Pages: 106–112
Conference: HT'24
Abstract
Multilingual countries are made of communities with their own history and traditions that maintain their identity but, at the same time, have strong social interactions and exchanges. This paper proposes a particular perspective to look at these communities: news circulation and translation. The hypothesis is that understanding how the same news is treated by different communities, and how each community focuses on some news instead of others, can be a proxy for their cultural similarities and differences. The paper presents a pilot analysis run on Switzerland news and exploits a fully quantitative approach. The experiments confirmed the feasibility of such an automatic approach and gave valuable insights about possible improvements and applications.
CCS CONCEPTS
• Information systems → Content analysis and feature selection; • Applied computing → Document analysis; Publishing; • Computing methodologies → Information extraction.
KEYWORDS
News analysis, TARO, Multiculture
ACM Reference Format:
Giuseppe Carrino, Angelo Di Iorio, and Davide Picca. 2024. Publishing, linking and translating news in multilingual communities: a mirror of cultural differences?. In 35th ACM Conference on Hypertext and Social Media (HT '24), September 10–13, 2024, Poznan, Poland. ACM, New York, NY, USA, Article 4, 7 pages. https://doi.org/10.1145/3648188.3675143
1 INTRODUCTION
Multilingualism is a major concern in Natural Language Processing (NLP) [6, 8] and cultural differences [2]. Thus, the case of Switzerland with four official languages, namely German, French, Italian, and Romansh, is ideal to examine the dynamics in the countries with multiple languages. Switzerland's approach of multilingualism policy is in contrast to other nations with single language and uphold cultural diversity that fosters the cohesiveness of the society.
As it is the case with other multilingual countries like Canada, India, or South Africa, there are multiple official languages, which means that people can interact in various ways and have different cultural affiliations. These settings are many times characterized by a high degree of bilingualism or multilingualism, thus offering a background that is ideal for capturing language and culture.
Our research utilizes SwissInfo¹. Based on this hypothesis, the present study will analyze the linguistic communities in Switzerland and the flow of information through the official multilingual online news platform: SwissInfo.ch, to determine its effectiveness. SwissInfo offers its news in all the four official languages of the country thus linking various communities and presenting a rich perspective on how news is distributed. Indeed this paper aims to uncover news circulation among the linguistic communities in Switzerland. Based on this, one can assume that the flow of news is similar or different in these groups. Our goal is to quantify the amount of translated news articles, estimate the degree of multilingual articles, and track the translation flow between languages.
Our approach adopts a fully quantitative methodology, using the TARO framework (See section 1.1). The first observations reported here, including the case of selective article translation in SwissInfo, demonstrate the integration and segregation of the communities according to language. This investigation is necessary for comprehending how language affects the sharing of information and the circulation of culture in a context that is multilingual.
1.1 TARO
The TARO (Tons of Articles Ready to Outline) framework is a methodological approach for collecting, translating, and analyzing online newspaper articles. As detailed by Carrino et al. [4], TARO aims to create an end-to-end modular system for examining the flow of news across different online outlets. Central to this framework is the concept of Snapshot Extension, which involves capturing regular snapshots of an outlet's homepage to compare how various outlets publish news and to identify exclusive content.
The Snapshot Extension allows for long-term analysis, which is particularly useful for studying publication trends in outlets like SwissInfo from both quantitative and qualitative perspectives. Data collection in TARO is facilitated by the Python Scrapy library, while translation employs the OpenNMT-based Neural Network Argos.
Another key concept in TARO is the definition of Equivalence between news items, which refers to articles that report the same news. Equivalence is calculated using document embeddings and cosine similarity, although for the SwissInfo analysis, a labeled system for article equivalence was utilized, explained in Section 3.
The motivation for using TARO in this analysis includes:
The need for long-term analyses using Snapshot Extension rather than isolated time points.
The theoretical utility of distinguishing between pieces of news and news items.
The ability to measure the overlap of news across homepages using the definition of identity and commonality.
In this research, the computation of identity will be adapted methodologically from the original TARO framework, avoiding the use of the less-tested similarity measure, as detailed in Section 3.
2 RELATED WORKS
Several studies investigate news circulation and translation in multilingual countries, including Switzerland. Vogler and Udris [12] examined how Swiss news media cover regions and bridge language barriers. They analyzed 25,035 articles from 47 outlets, detecting 48,861 unique mentions of place names, and found substantial regional segmentation despite some transregional coverage, aligning with our findings.
Vogler and Schneider [11] focused on Swiss expatriates and their news consumption, concluding that stronger connections to Switzerland increase the use of heritage country news. This highlights the peculiar role expatriates play in Swiss news consumption, corroborating our experiments. Furst and Grin [5] explored the correlation between multilingualism and creative thinking in Switzerland, showing a positive impact on cognitive and creativity processes.
Regarding news translation, a study [7] emphasized cross-cultural collaboration and highlighted biases and limitations in translating global news. Another extensive review [10] categorized 80 relevant research articles into linguistics, sociology, and communication studies, offering a tool for similar research with quantitative data on translated news. Caimotto and Gaspari [3] examined issues in translated news, noting that articles are often mediated and edited, revealing biases and patterns in journalistic translation research.
3 METHODS AND MATERIAL
Our research is carried on using the TARO framework tools, adapted for SwissInfo. SwissInfo is a Swiss online newspaper whose interesting characteristic for us is the co-presence of multiple versions, distinguished by language. Many linguistic versions of the articles and pages of SwissInfo exist, but the main focus of our research will be the English, German, French, and Italian versions. The last three versions are interesting since they are three of the four official Swiss spoken languages, while English is useful since it can be seen as a version for the general public i.e. expat or readers from other nations. Even if Romansch is a Swiss national language, it won't be considered in our analysis since it's not present in the selected outlet. The homepages of these four versions have been scraped, collecting all the news on them, considering that the structure of those is the same, except for the number of news items present.
The homepages of these versions can present a varying number of articles - from 40 to more than 100 - and this mainly depends on the number of carousel news, that vary from one language to another. For this, carousels have been excluded from our analyses, considering only the other articles present in the home pages. In this scenario, every homepage has an average of 21 articles, with a variance of |3| articles for each day. On SwissInfo, each news item not only has a title, content, and meta-data information, but one of the main features of every article is the translations section. Here, a list of pairs {language} - {URL} is provided, referring to the same piece of news translated in a different language on SwissInfo (example in Figure 1). This information is extremely valuable for the process of news items pairing in TARO. The SwissInfo linking system will be enough to identify similar news, as explained more in depth in the following Section. In the translations section of the articles, we can also see a label - visible in Figure 1 - that identifies the original version of the article, i.e. the language in which the article has been originally written, from which every other translation has been obtained. This information can also be very interesting for this specific analysis, and give some deeper insights on the editorial behavior of SwissInfo. Given this, the data collection step is performed by simply using the ScraPy library and collecting the same information mentioned in [4], with the additional information about translations and the original version. For this paper, not all the snapshots we collected have been used. The experiments have been focused on different small sets of days: for the commonality analysis, the used Snapshot Extension goes from 19-10-2023 to 23-10-2023, including so 5 full days. The same Snapshot Extension has been used also for the analyses on the originality - i.e. on the origin of news and their translation tendencies. For the content analysis, instead, single snapshots are used, because of the long computational time of LDA on big amount of data. The considered days are 23-10-2023 and between the 15th of November and the 21st of December 2023, as final computation to assess the reliability of previous analyses. The used JSON of SwissInfo articles can be found here ².
Figure 1: Example of articles translations list
3.1 Equivalence through Linking
The most important technical aspect of the usage of TARO for this research is the paradigm for the identification of news items that refer to the same piece of news. In the referenced paper, a subject-based news comparison was performed to pair news items and mark them as similar. For this work, however, the presence of translation links allows us to have certainty of these pairings by design, and so perform the desired quantitative analysis correctly assuming that correspondent news is paired. The first approach for equivalence checking relied on the translation list links. Given two homepages, $A$ and $B$, the algorithm for Equivalence through linking works as follows. We define homepage $A$ as cited and homepage $B$ as citing. What we do is looping on all news items of citing: for each of them, we loop on the translations URLs list. Given so a translation URL, we check if it is present in the list of news items of cited. Finally, a list of the found matches is returned. This algorithm, as described, should be symmetric, i.e., swapping $A$ and $B$ as cited and citing should output the same results. However, in reality, not always the translation lists of the news items are complete, creating situations in which the URL of news item $A$ can be found in the translations list of article $B$, but not vice versa. To reduce the occurrence of these asymmetries, an identificator-based comparison has been used too. More in-depth, formally speaking, we refer to the definition of Equivalence through identity in [4]: in the SwissInfo environment, given two articles $X$ and $Y$, the identity function $id$ retrieves the URL of the original version of a given news item. So, $id(X)$ will check the retrieved URL marked as original from the translations list of X. $id(Y)$ does the same for $Y$: if $id(X) == id(Y)$, then $X$ and $Y$ talk about the same piece of news. A visual representation of the id extraction is shown in Figure 2.
Figure 2: Example of articles ID extraction and mapping.
Starting from the definition of the same news, in this research we will focus on the commonality of news between the different SwissInfo versions. Our focus is on the overlapping of news items between the different homepages taken into account. To do so, we say that homepage A and homepage B have $N$ news in common: this means that $N$ news ids got from articles of $A$ have been found equal to $N$ ids got from articles of $B$. Also in this case, not always the original article is present in the translations list, leading to false negatives, but not in a significant number in our analyses.
3.2 Content and topic analysis
A major addition to the current TARO framework is a content analysis step, performed using the library Gensim³. Thanks to its function of Latent Dirichlet Allocation, news items can be analyzed and their themes extracted, simply clustering the semantically closely used terms. Its usage for topic extraction is widely spread in the literature [9], and was exploited for several applications on news coverage [13] and topic analysis [1].
In our work, LDA is used through the Python library Gensim to get a list of the most salient terms for different days and each analyzed homepage of SwissInfo, firstly trying to understand what different versions talk about and then exploiting the percentages of overlapping terms between the languages, over time. Is important to understand the difference between the analyses performed about commonality, in the first part of the next section, and the LDA intersection data showing. The first is performed using the URL linking in order to measure the amount of common news between the homepages, which is obtained by translations performed by SwissInfo itself. The LDA-based test, instead, is based on our translations - performed via Argos library - and topic extraction, focusing only on the most salient terms, and so giving a similarity measure computed mostly on the longest and most present articles - i.e. the most salient pieces of news of each homepage. In the next section, the most relevant results of both these analyses are reported.
4 RESULTS
The results of our experiments can be summarized in three different macro-categories: (1) Commonality Analysis, (2) Originality Analysis and (3) Topic Analysis.
4.1 Commonality Analysis
To check the intersections between news published in the four languages, a matrix has been plotted. On the x-axis we have the cited version and the y-axis represents the citing version. So, if we consider ITA as $x$ and ENG as $y$, the value shown is the number of news on the italian homepage found also on the english homepage. In the chart below, one snapshot - 2023-10-23 at 23:45 - is considered, with all three canton languages having 26 non-carousel articles on their homepage, whilst the English homepage only shows 21 items.
Figure 3: Commonality matrix of a single snapshot
The right-low square - including the three canton languages - portrays a relatively higher number of overlapping news, in particular considering the Italian and French versions. On the other hand, there is an intersection of less than 50% with the English homepage. To further investigate the overlapping between the four versions, the number of articles - for each of them - also published in different languages has been computed. The bar plot below (Figure 4) portrays the amount of news for each version grouped by the number of homepages they are present on (in all other languages). These values are then normalized by the number of total news present on the considered homepage, to have a ratio for each of them.
Figure 4: Normalized number of news found on one or more sections on a timespan of 5 days
For this experiment, a snapshot of 5 days have been considered, from 19-10 to 23-10, also to take into account the publication misalignment due to translations. It is visible how the three cantonal languages - French, German, and Italian - have very similar numbers of news shared between more than one version. On the other hand, English articles are way less translated, mostly present only on the English homepage⁴.
To confirm that there is a concrete overlapping between published news, the next table takes a closer look at the couples and the triplets of versions publishing the same news, counting how many articles are present in all the the considered homepages and then normalizing the number on the total number of analyzed news. Also, in this case, the 5 aforementioned days are taken into analysis.
| GE-IT | FR-IT | FR-GE | EN-IT | EN-GE | EN-FR |
| --- | --- | --- | --- | --- | --- |
| 0.18% | 0.21% | **0.22%** | 0.13% | 0.14% | 0.14% |Table 1: Normalized number of same news published in couples of languages with a 5-day timespan
| FR-GE-IT | EN-GE-IT | EN-FR-IT | EN-FR-GE |
| --- | --- | --- | --- |
| **0.33%** | 0.20% | 0.23% | 0.23% |Table 2: Normalized number of same news published in triplets of languages with a 5 days timespan
As expected, the number of non-carousel articles shared by the three cantonal languages is, in proportion, way higher than the other possible triplets, portraying a situation of relative exclusion of the English version and a cluster of articles including these three specific versions.
4.2 Originality Analysis
Now that an overview on the news sharing has been provided, and that a strong and interesting overlapping between cantonal languages has been outlined, the focus of the experiment is the understanding of the flow of news between the considered homepages. To do so, the concept of originality - already explained in Section 3 - has been exploited in order to show how the articles are mainly originally written and how they are then translated to the other versions. The questions we are trying to address are so: (i) which are the languages articles are originally written in? (ii) Is there a prominent one? (iii) Are there patterns of translations starting from different original languages?
Table 3 tries to answer to the first two questions, showing the distribution of languages of original news, considering the 5 days timespan between 19-10 and 23-10.
| ENG | FRE | GER | ITA |
| --- | --- | --- | --- |
| 15.1% | 18.7% | **57.8%** | 8.4% |Table 3: Distribution of original articles
We see how the vast majority of news is originally written in German, whilst only a small part is originally published in French and a even smaller one in Italian. We now take a look following which distribution these original articles are translated in the other analyzed languages (Figure 5).
Figure 5: Distribution of translations
Considering the same timespan mentioned before, it's visible how only a smaller part of the articles is translated into English - consistently with what is pictured in the previous plots. On the other hand, articles originally written in French or Italian are mostly translated into German, portraying, jointly with Table 3, a German-centric news production situation.
4.3 Topic Analysis
The last type of experiment focuses on the article's content, using strategies defined in Section 3. These analyses have been made to not only answer to how the publication of the article is defined and what are the relations between the other languages' news, but also to what is important to tell for each of them in a given time range. For this reason, mainly original news will be analyzed to compare the diverse editorial lines between the different linguistic boards. The presented tables will show top-5 salient terms for each of the four homepages taken into the exam, not focusing in particular on their order of, but mostly on their composition. The table below summarizes the topics on the day of 23-10-2023: due to the nearness to election day (22nd of the same month), expected topics are mostly candidate parties and the election themselves.
| | ENG | FRE | GER | ITA |
| --- | --- | --- | --- | --- |
| 1° | Collardi | **party** | Switzerland | Switzerland |
| 2° | Joyce | witchcraft | **party** | people |
| 3° | bank | witch | language | **vote** |
| 4° | Nora | hotel | canton | border |
| 5° | film | **campaign** | hydrogen | control |Table 4: Top-5 salient terms on the original news of 23-10-2023. Highlighted terms refer to elections.
Every set of original articles - except for the English one - presents at least one term directly related to elections. The English version, on the other hand, seems to mainly focus on editorials. In order to check the impact of translated news on this behavior, the same experiment has been conducted including all the articles of each homepage, not only the one originally written in the considered language.
| | ENG | FRE | GER | ITA |
| --- | --- | --- | --- | --- |
| 1° | **party** | Switzerland | swiss | swiss |
| 2° | collardi | hydrogen | **party** | Switzerland |
| 3° | swiss | **party** | **election** | **party** |
| 4° | Switzerland | abroad | foreign | right |
| 5° | **election** | **election** | neutrality | national |Table 5: Top-5 salient terms on the 23-10-2023 news. Highlighted terms refer to elections.
The inclusion of translated news implies a higher overlapping between the four versions, this time including also the English one. Even though the English homepage behavior in this specific day has shown as interesting, since the elections proximity could lead to a high bias in the topics distribution of the newspaper, other experiments have been conducted on different days. In the next table, most salient terms on the original news online during the 15th of November are pictured.
| | ENG | FRE | GER | ITA |
| --- | --- | --- | --- | --- |
| 1° | gold | party | hydrogen | swiss |
| 2° | UN | UDC | Kosovo | hydrogen |
| 3° | cannabis | sale | Russia | Switzerland |
| 4° | drone | watch | russian | wolf |
| 5° | Lancet | brand | component | phone |Table 6: Top-5 salient terms on the original news of 15-11-2023.
Table 6 is significantly different from the previous ones: there is almost no common term between the four versions, except for hydrogen both cited in Italian and German articles. However, this single common term, since it's extracted from original articles, implies a similar editorial line between these two SwissInfo versions, not only due to translation - as pictured in charts of Figure 5. In the next table, as previously done, all the articles from the same day are analyzed, extracting their topics, to look for differences including translated articles.
| | ENG | FRE | GER | ITA |
| --- | --- | --- | --- | --- |
| 1° | gold | hydrogen | Kosovo | hydrogen |
| 2° | UN | construction | drone | Switzerland |
| 3° | hydrogen | election | hydrogen | Russian |
| 4° | cannabis | war | Russia | Russia |
| 5° | exhibition | Ukraine | neutrality | Kosovo |Table 7: Top-5 salient terms on the 15-11-2023 news.
Including translated articles, also in this case, a way higher overlap is visible: the three cantonal languages versions seem to share the Russian-Ukranian war topic - whilst it's not in the top salient terms for the English version. Also, they all now share the hydrogen term, referring to green policies. Still Italian and German versions are presented as the most similar ones, with almost a 100% terms overlapping, being the only two versions citing Kosovo.
This comparative analysis has been performed automatically on a range of 6 days (all days between 15-11-2023 and 20-11-2023, both included) to check the percentage of intersection between the 10-top salient terms of each version. The analysis has been performed both on original articles and complete homepages, as manually done in the previous paragraphs. The computed values for only original articles are summarized in the next table, showing the percentages of intersecting terms over the total number of them - i.e. 10 - with pairs as rows and considered days as columns. Percentages higher than 20% are highlighted.
| Pairs | 15 | 16 | 17 | 18 | 19 | 20 | Avg |
| --- | --- | --- | --- | --- | --- | --- | --- |
| **EN-FR** | 0% | 0% | 0% | 0% | 0% | 10% | 1.67% |
| **EN-GE** | **20%** | 10% | **20%** | 10% | 0% | 10% | 11.67% |
| **EN-IT** | 0% | **20%** | 10% | 10% | **20%** | 10% | 11.67% |
| **FR-GE** | 0% | 0% | 0% | 10% | 0% | **40%** | 8.33% |
| **FR-IT** | 0% | 10% | 10% | 0% | 0% | 10% | 5.00% |
| **GE-IT** | 10% | 10% | 10% | 10% | 10% | 10% | 10.00% |Table 8: Top-10 salient terms overlapping between 15-11-2023 and 20-11-2023 (only originals).
We see here how the overlapping is not very high and surely not constant: for instance, the French and German versions have a average intersection rate of 8.33%, but on the 20th of November they had 40% of the most salient terms in common. This seems to confirm also what was visible in the previous tables: not a high overlapping in general, considering only the original articles - each version usually has its topics, usually.
If we consider triplets, also, we have an average value of intersections all equal to 1.67%, showing no difference between the three cantonal languages and the English version of SwissInfo. The situation is slightly different when looking at all articles. A table formatted as Table 8, but considering the whole homepages is following.
| Pairs | 15 | 16 | 17 | 18 | 19 | 20 | Avg |
| --- | --- | --- | --- | --- | --- | --- | --- |
| **EN-FR** | 10% | 10% | 10% | 0% | 0% | 10% | 6.67% |
| **EN-GE** | 10% | 10% | 0% | 0% | 10% | 0% | 5.00% |
| **EN-IT** | 10% | **20%** | **20%** | **20%** | **20%** | **20%** | 18.33% |
| **FR-GE** | **20%** | **50%** | **70%** | **20%** | **70%** | **60%** | **48.33%** |
| **FR-IT** | **20%** | **40%** | **30%** | **30%** | **60%** | **50%** | **38.33%** |
| **GE-IT** | **70%** | **40%** | **40%** | **50%** | **50%** | **20%** | **45.00%** |Table 9: Top-10 salient terms overlapping between 15-11-2023 and 20-11-2023 (all articles).
Now the overlapping values are way higher, with intersection average percentages higher than 35% for all the couples involving cantonal versions of SwissInfo. The French and German versions share almost half of the terms, showing so similar rates of overlapping news to the one shown in Figure 3.
This behavior reflects also on average intersections of triplets: now, the three cantonal versions have an average overlapping of 26.33% of the 10-top salient terms, whilst all other possible triples - i.e. the ones including also the English version - result in circa 3% of terms intersection.
The same experiment has been performed by using a single Snapshot Extension including all 6 days instead of one computation per day, showing no remarkable differences. To further validate our analysis, the same experiments have been performed for a whole month, between 21-11-2023 and 21-12-2023. The intersection percentages are subsequently presented. The charts presented are relative to all the articles, not just the original ones for each version.
Figure 6: Couples intersection over one month [1]
Figure 7: Couples intersection over one month [2]
As shown, no visible pattern is shown in couples intersections, except for the fact that all of the couples seem to follow a similar trend over days. Furthermore, the average intersection rate is slightly higher in couples not including English, as visible in the next Table.
| GE-IT | FR-IT | FR-GE | EN-IT | EN-GE | EN-FR |
| --- | --- | --- | --- | --- | --- |
| 23.2% | 25.8% | 26.4% | 20.0% | 19.0% | 15.8% |Table 10: Average couple's intersection percentages over one month
5 DISCUSSION AND CONCLUSIONS
The paper tries to demonstrate a comprehensive analysis that delineates a nuanced view regarding the inter-relationship of language, culture, and media in a multilingual society like Switzerland.
The linguistic hierarchy of Swiss newspaper production, as shown in Table 3, has a high proportion of originality in German-origin articles. This links the translation inherently with the politics of publishing in Swiss newspapers, because the majority demographic residing in Switzerland speaks in German. The distribution thus underlines domestic market orientation for national languages but also outlines the specialized niche that English occupies and is aimed, mainly at expatriates and international readers. The dynamic interaction of linguistic communities exemplified by the heat map in Figure 3 makes for a very interesting dialogue of intensified exchanges between the German, French, and Italian linguistic versions within Switzerland. The limited translation engagement of English articles underlines its targeted appeal to certain demographics, possibly those that are less anchored in the immediacy of Swiss national discourse. This further, therefore, underscores the strategic role of English towards global or expatriate audiences in light of reflecting Switzerland's international stance and commitment to multicultural inclusivity.
Combining graphical data analysis with a study of salient news terms is rather interesting. One looks here as an attractive opportunity to plunge deeper into a subtle correlation between language, cultural identity, and public speech in Switzerland. From the dominance of German to the predominance of English in the strategy against an international audience and very balanced French and Italian translations, it seems to be a complex entanglement of Swiss linguistic diversity and media practice.
Thus, this landscape reflects the linguistic and cultural diversity of the nation, whereas it points to the complex links of language, culture, and media in developing collective public discourse.
Before concluding, it is helpful to summarize the main threats to validity of our experiments. First of all, our study considers only one source of data - Swissinfo.ch. On the one hand, Swissinfo includes articles written and translated in the four languages used in Switzerland and can be a good proxy for the publication and consumption of news in these languages. On the other hand, the publication and translation of an article might depend on the editorial choices of Swissinfo and cannot be generalized for all Swiss news sources. We consider these limitations acceptable since our goal is not to get an exhaustive answer on Swiss linguistic communities, rather to verify the feasibility of a TARO-based quantitative approach for such an analysis.
The same considerations can be extended to the temporal dimension: we considered a quite large interval in order to be as general as possible but more experiments would produce more accurate results. Indeed, external events might have affected the choice of news that Swissinfo has published and translated. Again, a larger set of experiments is helpful to mitigate this bias.
There is another technical issue worth highlighting: our experiments assume that the links between the original articles and their translations are all correct and that there is no link missing. Such a validity cannot be verified externally but obviously affects all our results.
Notes
https://www.swissinfo.ch/
https://doi.org/10.5281/zenodo.10882717
https://radimrehurek.com/gensim/
Other experiments not reported here portrayed a very similar situation for other day(s)
References
Watanabe Arisa and Chakraborty Basabi. 2021. Time-series Analysis of Newspaper Articles for Automatic Event Detection using LDA. In 2021 IEEE 4th International Conference on Knowledge Innovation and Invention (ICKII). 166–169. https://doi.org/10.1109/ICKII51822.2021.9574704
Larissa Aronin and Muiris Ó Laoire. 2013. The material culture of multilingualism: moving beyond the linguistic landscape. International Journal of Multilingualism 10 (2013), 225 – 235. https://doi.org/10.1080/14790718.2012.679734
Maria Cristina Caimotto and Federico Gaspari. 2018. Corpus-based study of news translation: challenges and possibilities. Across Languages and Cultures Across Languages and Cultures 19, 2 (2018), 205 – 220. https://doi.org/10.1556/084.2018.19.2.4
Giuseppe Carrino, Angelo Di lorio, and Gioele Barabucci. 2023. Comparison of news commonality and churn in international news outlets with TARO. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT '23). Association for Computing Machinery, New York, NY, USA, Article 26, 10 pages. https://doi.org/10.1145/3603163.3609062
Guillaume Fürst and François Grin. 2023. Multilingualism, multicultural experience, cognition, and creativity. Frontiers in Psychology 14 (2023). https://doi.org/10.3389/fpsyg.2023.1155158
D. Jain, Yamila García-Martínez Eyre, Akshi Kumar, B. Gupta, and K. Kotecha. 2023. Knowledge-based Data Processing for Multilingual Natural Language Analysis. ACM Transactions on Asian and Low-Resource Language Information Processing (2023). https://doi.org/10.1145/3583686
Kayo Matsushita and Christina Schäffner. 2018. Multilingual collaboration for news translation analysis: possibilities and limitations. Across Languages and Cultures 19, 2 (2018), 165 – 184. https://doi.org/10.1556/084.2018.19.2.2
Davide Picca, Alfio Gliozzo, and Simone Campora. 2009. Bridging languages by SuperSense entity tagging. In Proceedings of the 2009 Named Entities Workshop: Shared Task on Transliteration (NEWS 2009). 136–142.
Zhou Tong and Haiyi Zhang. 2016. A Text Mining Research Based on LDA Topic Modelling. Computer Science & Information Technology 6, 201–210. https://doi.org/10.5121/csit.2016.60616
Roberto A. Valdeón. 2020. Journalistic translation research goes global: theoretical and methodological considerations five years on. Perspectives 28, 3 (2020), 325–338. https://doi.org/10.1080/0907676X.2020.1723273
Daniel Vogler and Jörg Schneider. 2023. Analysing the media repertoires that Swiss expatriates use to inform themselves about their heritage country. The Journal of International Communication 29, 2 (2023), 175–195. https://doi.org/10.1080/13216597.2022.2162948
Daniel Vogler and Linards Udris. 2021. Transregional News Media Coverage in Multilingual Countries: The Impact of Market Size, Source, and Media Type in Switzerland. Journalism Studies 22, 13 (2021), 1793–1813. https://doi.org/10.1080/1461670X.2021.1965909
Batool Zehra, Naeem Mahoto, and Vijdan Khalique. 2018. Estimating News Coverage Patterns using Latent Dirichlet Allocation (LDA). Sukkur IBA Journal of Emerging Technologies 1 (06 2018). https://doi.org/10.30537/sjet.v1i1.142
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime