Abstract
In this paper, we analyze how friend recommendation algorithms on social networks promote echo chambers. We analyze both link bias and content bias using a real social graph from X (Twitter). We extract a follow graph from X, repeatedly add new edges selected by a recommendation algorithm, and observe how the degree of bias in the graph changes. Our findings include: (1) the follow graph of X is sufficiently homophilic for recommendation algorithms to produce link bias, (2) iterated recommendations do not accelerate increase of content bias, (3) even when an algorithm recommends no user from the target user’s community, it sometimes produces link bias by recommending users from a few other communities, (4) but no similar phenomenon is observed for content bias.
CCS Concepts: Information systems → Social networking sites.
Keywords: social network, Twitter, filter bubble, social division, polarization
1 Introduction
Individuals in social networks tend to form communities with others having similar attributes and thoughts [17], [24]. It results in a closed environment called “echo chamber” [20], in which opinions that are different from one’s own become impossible to obtain.
Online social network services (SNS) have been criticized as one of the causes of social division [12], [19]. In particular, it has been discussed whether friend recommendation on SNS promotes echo chambers. Some studies have rejected that hypothesis [3], [7], [18], [26], and some studies have shown that echo chambers can be caused by other factors such as cognitive bias of individual users and information sharing functions of SNS [14], [29]. However, recent studies reported that friend recommendation algorithms can promote echo chambers [5], [9], [11], [30]. Following these studies, further analysis has become an important research topic.
In this paper, we analyze the extent to which friend recommendation algorithms promote echo chambers on X (Twitter). While most existing studies only analyze link bias [11], [30] and/or only use synthetic data [5], [11], [30], we analyze content bias using real X data. Content bias is important because structural and homogeneity-based echo chambers do not necessarily coincide [9], and content bias may be a more direct cause of bias in user thoughts. In addition, several studies reported that item recommendation does not promote content bias [7], [18], [26], and even friend recommendation on X does not promote political content bias [9], while other studies reported that friend recommendation can promote content bias [5] on synthetic data. We need further research on content bias on X.
We also analyze both global bias throughout the graph and local bias for each user. In addition, we observe whether the increase in bias is accelerated through the iteration of recommendation.
We compare the following recommendation algorithms: content-based method [16], most-common-neighbor method [25], a method based on the Adamic-Adar index [1], and Personalized SALSA [15].
The outline of the analysis procedure is as follows. First, we extract a follow graph from X, and divide it into communities using a community detection algorithm. We then repeatedly add new edges selected by each recommendation algorithm, and observe how the degree of bias in the graph changes. An alternative approach would be to collect multiple snapshots from X and observe the change in the degree of bias over a period, but it would include changes owing to other factors. To isolate the influence of the recommendation from the other factors, we simulate the situation where all users choose the next user to follow solely based on the recommendation.
Another important factor in the formation of echo chambers is the change of user’s opinions under the influence from the neighbors. Ge et al. [13] reported that recommendation on e-commerce sites increases the bias in users’ interests, and the bias in users’ interests increases the bias of the recommendation. However, Cinus et al. [5] reported that friend recommendation can produce echo chambers even when users do not change their opinions. In this research, we examine whether recommendation algorithms promote echo chambers on X even when users do not change opinions.
Our findings include: (1) the follow graph of X is sufficiently homophilic for recommendation algorithms to further produce link bias, (2) the generation of link bias is accelerated through the iteration by some algorithms, but not by all algorithms, (3) the generation of content bias is accelerated through iteration by none of the four algorithms, (4) even when an algorithm recommends no user from the target user’s community, it sometimes produces link bias by recommending users from a few other communities, (5) but no similar phenomenon is observed for content bias: No algorithm recommended users exclusively from a small number of content topics that are not the topic of the target user’s current interest.
2 Related Work
Some studies have argued that recommendation systems do not reduce the variety of content consumed, or have a small impact. Nguyen et al. [26] analyzed the impact of collaborative filtering-based recommendations in movie rating websites. Their result showed that the diversity of content consumed by users who did not act in line with recommendations decreased compared to that of users who frequently selected recommended content. Hosanagar et al. [18] analyzed the effect of recommendation on music platforms and showed that users who acted in line with recommendations expanded their range of interests through the recommendation.
There are also studies that showed that factors other than the recommendation system can cause echo chambers. Geschke et al. [14] showed that echo chambers can occur even when information filtering was not performed at the system’s recommendation level but was performed only at the individual user’s level. Sasahara et al. [29] showed that the post sharing functions in SNS, such as repost (retweet) in X, is also one of the factors that cause echo chambers.
However, Cinus et al. [5] reported that their simulation of friend recommendation algorithms on synthetic networks produces echo chambers if the initial graph exhibits sufficient homophily and is not fully modularized. We conduct a simulation starting from the real network from X to examine if it produces echo chambers.
Fabbri et al. [10] also conducted simulations using synthetic networks, and reported that sufficiently homophilic minority groups get a disproportionate advantage in exposure to other users. It also underscores the importance of research using real network from X.
3 Analysis Procedure
We use the following notations in this paper. G(V, E) denotes a follow graph where (u, v) ∈ E iff a user u ∈ V follows v ∈ V. We divide G into k communities, which are denoted by C<sub>1</sub>, …, C<sub>k</sub>. Γ<sup>+</sup>(u) and Γ<sup>−</sup>(u) denote a set of followees and followers of u, respectively. Γ(u) denotes the neighbors of u, i.e., Γ(u) = Γ<sup>+</sup>(u) ∪ Γ<sup>−</sup>(u).
3.1 Compared Recommendation algorithms
We compare the following four friend recommendation algorithms.
The content-based method (CM), as the name implies, recommends a user based on the content similarity. We use cosine similarity between feature vectors of users. The feature vector of a user u is defined by the centroid of the TF-IDF vectors produced from the posts by the neighbors of u. We chose this method because Hannon et al. [16] reported that the posts by neighbors well represent a user. When |Γ(u)| > 1, 000, we randomly sampled 1,000 users in total from Γ<sup>+</sup>(u) and Γ<sup>−</sup>(u) proportionally to their sizes.
Most-common-neighbor method (MCN) assumes that the probability of an edge between a pair of nodes is proportional to the number of their common neighbors. The original model [25] is defined for undirected graphs, and we modify it for directed graphs. We assume that u is likely to follow x if x is followed by many neighbors of u, therefore, we define MCN of u and x as follows:
For each u, we recommend x with the largest MCN(u, x) value.
We also define a method that uses the Adamic-Adar index [1] instead of MCN. Because it is also defined for undirected graphs, we modify its definitions for directed graphs as follows:
Personalized SALSA (pSALSA) [15] is a random walk-based algorithm proposed for friend recommendation on X. Similarly to PageRank [27], SALSA is based on a random walk, but it uses a bipartite graph similar to those used in HITS [21]. Its random walk goes back and forth between the two components of the bipartite graph. We used the variation of pSALSA used in [28], where u is used as the starting point. The random jump probability α was set to 0.95 and the random walk step number N was set to 1,000.
3.2 Metrics for global and local bias
We use two metrics for evaluating the global bias: modularity and Pseudo F statistic. Modularity [6] measures how well communities are separated in a graph. It is defined by the deviation of its intra-community edge density from the expected one in a random graph consisting of the same set of nodes preserving the node degrees. Large modularity means high link structure bias in the graph.
Pseudo F statistic [4] measures how well a set of data points are clustered. It is defined by the ratio of the sum of the distance between the centroid of each cluster and the centroid of the whole data set to the sum of the distances between each data point and the centroid of the cluster it belongs to. If its value is large for the feature vectors produced from the contents consumed by each user, it means that there exists high content bias in the social network.
We also defined four metrics for local bias for each user.
The recommendation link commonality (LC) is the ratio of the recommended users belonging to the same community as the recommendation-target. If most users newly followed by u belong to the same community as u, u is likely to be in an echo chamber.
The recommendation link bias (LB) is the ratio of communities including at least one of the recommended users. If all recommended users belong to a few communities, even if they are not the community of the target user u, u is likely to be in an echo chamber. That is, low (not high) LB values imply a high degree of bias.
Table 1:. Example of LC and LB
community of newly followed users | LC | LB | |
|---|---|---|---|
u<sub>1</sub> ∈ C<sub>1</sub> | C<sub>1</sub>, C<sub>1</sub>, C<sub>2</sub>, C<sub>3</sub> | 0.5 | 0.75 |
u<sub>2</sub> ∈ C<sub>1</sub> | C<sub>1</sub>, C<sub>1</sub>, C<sub>2</sub>, C<sub>2</sub> | 0.5 | 0.5 |
u<sub>3</sub> ∈ C<sub>1</sub> | C<sub>3</sub>, C<sub>4</sub>, C<sub>4</sub>, C<sub>4</sub> | 0 | 0.5 |
LC and LB are complementary to each other as exemplified by cases listed in Table 1. Let u<sub>1</sub>, u<sub>2</sub>, and u<sub>3</sub> belong to the community C<sub>1</sub>, and have newly followed four recommended users. Table 1 lists the communities of these four users. Let the number of communities in the graph be four. LC values for u<sub>1</sub> and u<sub>2</sub> are the same, but u<sub>2</sub> is provided with information from a smaller world, which is a property of echo chambers. There is also a possibility that u<sub>2</sub> is in a middle-sized echo chamber including C<sub>1</sub> and C<sub>2</sub>. LB can distinguish the cases of u<sub>1</sub> and u<sub>2</sub>. When comparing u<sub>2</sub> and u<sub>3</sub>, their LB values are the same, but u<sub>2</sub> is provided with less information from outside of its community C<sub>1</sub>, and is more likely to be in an echo chamber solely consisting of C<sub>1</sub>. LC can distinguish the cases of u<sub>2</sub> and u<sub>3</sub>.
The recommendation content commonality (CC) is the average content similarity between the newly followed users and the recommendation-target user u. If CC is large, u consumes contents biased toward the characteristics that u already has.
The recommendation content bias (CB) is the average content similarity between the newly followed users. If CB is large, the contents consumed by the recommendation-target user u is biased toward a narrow range, which is a property of echo chambers. Similar to LC and LB, CC and CB are complementary to each other.
3.3 Details of the procedure
As explained before, we use the modularity of the graph for measuring global link structure bias. We divide the acquired follow graph into communities before adding edges through the recommendations. We used the standard Louvain method [2] for it because previous studies [8], [28] have shown that the choice of the community detection algorithms did not significantly affect the result of comparative evaluation of recommendation systems. This community structure was fixed, that is, we did not re-compute the communities when we add new edges.
The number of edges added in each iteration is the same as the total number of users in the graph, i.e., |V|. The following are the two possible procedures for adding |V| new edges.
•One to One (OtO): Adding one out-edge to each user.•Depend on Follows (DoF): The number of outgoing edges added to each user is proportional to the number of outgoing edges it currently has. For example, when |V| = 100 and |E| = 1, 000, a user with 20 follows will newly follow 10020/1, 000 = 2 users per iteration.
In the former strategy, we can analyze the change of local bias for all users in every iteration. In the latter strategy, some users may be given no new edge. However, the latter is closer to what happens on the real SNS. We compared the two procedures in a preliminary experiment with a smaller dataset. Because the result showed that these two procedures do not largely change the result, we used OtO in our main experiment. The details will be explained later.
Every time we add edges, we recompute the recommendation. However, if we every time recompute the ranking of all candidates x for all target u, its computational cost is too high. To reduce it, we only recompute the rankings of the top 100 users in the previous ranking. We only use the top one user for recommendation, so re-ranking the top 100 in the previous ranking is likely to suffice.
4 Experiments
We collected users (and their posts) from X by using the rejection-controlled Metropolis-Hastings algorithm [23]. A random walk-based algorithm was chosen because it is superior in preserving the characteristics of the original graph [22]. Users that met any of the following conditions were excluded during the sampling:
•private accounts,•users with less than 20 posts, and•users with the most recent post more than six months old.
When retrieving their posts, we excluded reposts. For quoted reposts, only the text added by the users was extracted.
We collected two graphs, one for the preliminary experiment, and one for the main experiment. For the former, we collected latest 20 posts for each user. For the latter, latest 50 posts (or between 20 and 50 if the user has less than 50 posts). Table 2 lists some statistics.
Table 2:. Dataset Statistics
# nodes | # edges | # communities | # tweets | |
|---|---|---|---|---|
pre | 14,805 | 184,951 | 5 | 296,100 |
main | 48,857 | 709,143 | 12 | 1,630,615 |
4.1 Preliminary experiment
Change in modularity: (left) OtO (right) DoF
Change in pseudo F statistic: (left) OtO (right) DoF
In the preliminary experiment, we only compared the modularity and pseudo F, and used only CM, MCN, and pSALSA. We iterated the addition of new edges 10 times.
Fig. 1 shows the change in the modularity through the iteration of the new edge addition by the OtO (left) and DoF (right). The blue, red, and green lines represent the results for CM, MCN, and pSALSA, respectively. Although the absolute values of modularity are different in these two graphs, there is no significant qualitative difference between them. Fig. 2 shows the change in the pseudo F statistic. Similarly, there is no significant qualitative difference. Based on these results, we conducted the main experiment with the OtO procedure, which is better for the observation of local bias.
4.2 Main experiment
Change in modularity (left) and pseudo F (right)
In the main experiment, we iterated the addition of new edges 12 times, which nearly doubles |E| and must be enough. The change in modularity through these iterations is shown in Fig. 3 (left). The blue, red, yellow, and green lines represent the results for CM, MCN, ADA, and pSALSA, respectively. All four algorithms increased the modularity, i.e., produced link bias. Note that the modularity is defined by the deviation from random graphs and is not expected to increase if random edges are added. MCN increases it the most, closely followed by ADA. Another link-based method pSALSA follows next. This is an expected result because MCN and ADA use more local graph structure than pSALSA does, and ADA reduces the influence of densely connected neighbors compared to MCN. This result also suggests that the follow graph of X is sufficiently homophilic, which is the necessary condition for recommendation algorithms to produce link bias [5]. Another finding here is that the generation of link bias is accelerated by MCN and ADA through the iteration, but not by pSALSA and CM.
Fig. 3 (right) shows the change in pseudo F statistic. The content-based method CM increases the content bias the most, as expected. This result also shows that even link-based methods produce content bias when we start from a follow graph of X. The order among the three link-based methods is the same as that for modularity. Another finding is that the generation of content bias is accelerated through iteration by none of these four methods. CM even slowed down probably because of saturation. This suggests that content bias has no reinforcement factor when user opinions do not change.
Figs. 4 and 5 show the results for the metrics for local bias. They show a histogram, i.e., the number of users that have the metric values within a specific range after all iterations.
Fig. 4 (left) shows the result for LC. The blue bars are shorter than the others on the right side, which means that CM produces high link bias for fewer users. On the right side, the order among the three link-based methods is MCN, ADA, pSALSA, which is consistent with the result for modularity in Fig. 3 (left).
Fig. 4 (right) shows the result for LB. Note that smaller LB values mean more bias. There is a difference between the results for LB and LC. On the left side of the result for LB, where bias is larger, we can observe the same order, MCN, ADA, pSALSA, and CM, as in the right side (larger LC means larger bias) of the result for LC. However, in the result for LB, CM is more concentrated in the middle ranges, and on the right side (where bias is small), pSALSA has more users than CM. That is, LC shows that CM is less likely to recommend users from the target user’s own community, but LB shows that CM often produces modest link bias by recommending users exclusively from a small number of other communities.
In Fig. 4 (left), LC=0 means that no user in the community of the target user was recommended. MCN and ADA take larger values at LC=0 than at LC=1/12. That is, MCN and ADA are relatively likely to recommend users entirely from other communities. In Fig. 4 (right), LB=1/12 means that all the recommended users belong to the same community. The height of the red and yellow bars at LB=1/12 is almost the same as their height at LC=0. This is because MCN and ADA recommend users entirely from other communities when they choose all those users from some other single community.
Fig. 5 (left) shows the result for CC. On the right side, where the content bias is large, CM dominates. pSALSA is taller than the others in the middle ranges. We can observe similar results in Fig. 5 (right) showing the result for CB. There was a difference between the results for LC and LB as explained above. By contrast, no big difference was observed in the results for CC and CB. That is, no algorithm tends to produce content bias by recommending only a few topics that are not the topic of the target user’s current interest.
Distribution of LC (left) and LB (right)
Distribution of CC (left) and CB (right)
5 Conclusion
In this paper, we analyzed the formation of echo chambers on a follow graph of X by four major friend recommendation algorithms: CM, MCN, ADA, and pSALSA. The important findings are:
•The follow graph of X is sufficiently homophilic for recommendation algorithms to produce further link bias [5].•The generation of link bias is accelerated by MCN and ADA through the iteration, but not by pSALSA and CM.•The generation of content bias is not accelerated.•Even when CM recommends no user from the target user’s community, it may still produce link bias by recommending users exclusively from a small number of communities.•A similar phenomenon was not observed for content bias.
We expect that these findings and further analysis of these phenomena will provide insights for designing better friend recommendation algorithms that do not lead to echo chambers.
Acknowledgments
This work was supported by JSPS KAKENHI JP23K28095.
References
[1]Lada A Adamic and Eytan Adar. 2003. Friends and neighbors on the web. Social Networks 25, 3 (2003), 211–230.
[2]Vincent D Blondel, Jean-Loup Guillaume, Renaud Lambiotte, and Etienne Lefebvre. 2008. Fast unfolding of communities in large networks. Journal of Statistical Mechanics: Theory and Experiment 2008, 10 (2008), P10008.
[3]Axel Bruns. 2019. Filter bubble. Internet Policy Review 8, 4 (2019). https://policyreview.info/concepts/filter-bubble.
[4]Tadeusz Caliński and Jerzy Harabasz. 1974. A dendrite method for cluster analysis. Communications in Statistics-Theory and Methods 3, 1 (1974), 1–27.
[5]Federico Cinus, Marco Minici, Corrado Monti, and Francesco Bonchi. 2022. The effect of people recommenders on echo chambers and polarization. In Proceedings of AAAI ICWSM , Vol. 16. 90–101.
[6]Aaron Clauset, Mark EJ Newman, and Cristopher Moore. 2004. Finding community structure in very large networks. Physical Review E 70, 6 (2004), 066111.
[7]Giordano De Marzo, Pietro Gravino, and Vittorio Loreto. 2024. Recommender systems may enhance the discovery of novelties. Journal of Physics: Complexity 5, 4 (2024), 045008.
[8]Pasquale De Meo, Emilio Ferrara, Giacomo Fiumara, and Alessandro Provetti. 2014. On Facebook, most ties are weak. Commun. ACM 57, 11 (2014), 78–84.
[9]Kayla Duskin, Joseph S Schafer, Jevin D West, and Emma S Spiro. 2024. Echo chambers in the age of algorithms: an audit of twitter’s friend recommender system. In Proceedings of ACM WebSci. ACM, New York, NY, USA, 11–21.
[10]Francesco Fabbri, Maria Luisa Croci, Francesco Bonchi, and Carlos Castillo. 2022. Exposure inequality in people recommender systems: The long-term effects. In Proceedings of AAAI ICWSM , Vol. 16. 194–204.
[11]Antonio Ferrara, Lisette Espín-Noboa, Fariba Karimi, and Claudia Wagner. 2022. Link recommendations: Their impact on network structure and minorities. In Proceedings of ACM WebSci. ACM, New York, NY, USA, 228–238.
[12]Venkata Rama Kiran Garimella and Ingmar Weber. 2017. A Long-Term Analysis of Polarization on Twitter. Proceedings of AAAI ICWSM, 528–531.
[13]Yingqiang Ge, Shuya Zhao, Honglu Zhou, Changhua Pei, Fei Sun, Wenwu Ou, and Yongfeng Zhang. 2020. Understanding echo chambers in e-commerce recommender systems. In Proceedings of ACM SIGIR. ACM, New York, NY, USA, 2261–2270.
[14]Daniel Geschke, Jan Lorenz, and Peter Holtz. 2019. The triple-filter bubble: Using agent-based modelling to test a meta-theoretical framework for the emergence of filter bubbles and echo chambers. British Journal of Social Psychology 58, 1 (2019), 129–149.
[15]Pankaj Gupta, Ashish Goel, Jimmy Lin, Aneesh Sharma, Dong Wang, and Reza Zadeh. 2013. Wtf: The who to follow service at twitter. In Proceedings of WWW. 505–514.
[16]John Hannon, Mike Bennett, and Barry Smyth. 2010. Recommending twitter users to follow using content and collaborative filtering approaches. In Proceedings of ACM RecSys. ACM, New York, NY, USA, 199–206.
[17]Bas Hofstra, Rense Corten, Frank Van Tubergen, and Nicole B Ellison. 2017. Sources of segregation in social networks: A novel approach using Facebook. American Sociological Review 82, 3 (2017), 625–656.
[18]Kartik Hosanagar, Daniel Fleder, Dokyun Lee, and Andreas Buja. 2014. Will the global village fracture into tribes? Recommender systems and their effects on consumer fragmentation. Management Science 60, 4 (2014), 805–823.
[19]Jasper Jackson. 2017. Eli Pariser: Activist whose filter bubble warnings presaged Trump and Brexit. The Guardian (Jan. 2017).
[20]Kathleen Hall Jamieson and Joseph N Cappella. 2008. Echo chamber: Rush Limbaugh and the conservative media establishment. Oxford University Press.
[21]Jon M Kleinberg et al. 1998. Authoritative sources in a hyperlinked environment.. In Proceedings of ACM-SIAM SODA , Vol. 98. 668–677.
[22]Jure Leskovec and Christos Faloutsos. 2006. Sampling from large graphs. In Proceedings of ACM SIGKDD. ACM, New York, NY, USA, 631–636.
[23]Rong-Hua Li, Jeffrey Xu Yu, Lu Qin, Rui Mao, and Tan Jin. 2015. On random walk based graph sampling. In Proceedings of IEEE ICDE. 927–938.
[24]Miller McPherson, Lynn Smith-Lovin, and James M Cook. 2001. Birds of a feather: Homophily in social networks. Annual Review of Sociology 27, 1 (2001), 415–444.
[25]Mark EJ Newman. 2001. Clustering and preferential attachment in growing networks. Physical Review E 64, 2 (2001), 025102.
[26]Tien T Nguyen, Pik-Mai Hui, F Maxwell Harper, Loren Terveen, and Joseph A Konstan. 2014. Exploring the filter bubble: the effect of using recommender systems on content diversity. In Proceedings of WWW. 677–686.
[27]Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999. The PageRank citation ranking: Bringing order to the web. Technical Report. Stanford InfoLab.
[28]Javier Sanz-Cruzado and Pablo Castells. 2018. Enhancing structural diversity in social networks by recommending weak ties. In Proceedings of ACM RecSys. ACM, New York, NY, USA, 233–241.
[29]Kazutoshi Sasahara, Wen Chen, Hao Peng, Giovanni Luca Ciampaglia, Alessandro Flammini, and Filippo Menczer. 2021. Social influence and unfollowing accelerate the emergence of echo chambers. Journal of Computational Social Science 4, 1 (2021), 381–402.
[30]Han Zhang, Shangen Lu, Yixin Wang, and Mihaela Curmei. 2023. Delayed and Indirect Impacts of Link Recommendations. In Proceedings of ACM Conference on Fairness, Accountability, and Transparency. ACM, New York, NY, USA, 545–557.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime