The Effect of Recommendation Source and Justification on Professional Development Recommendations for High School Teachers
Lijie Guo, Christopher Flathmann, Reza Anaraky, Nathan McNeese, Bart Knijnenburg
Published in: HT ’22: Proceedings of the 33rd ACM Conference on Hypertext and Social Media · DOI: 10.1145/3511095.3531280
This Seed edition’s formatting was converted from the supplied ACM publisher HTML and checked against the supplied ACM PDF. Where a publisher figure endpoint refused retrieval, the embedded image is a facsimile of its supplied PDF source page.
Abstract
This paper describes a study conducted in the process of building a recommender system that provides personalized professional development pathways for high school teachers seeking to increase their disciplinary knowledge and/or their teaching skills. A controlled experiment (N = 190) was conducted to study the effects of the presented justification for the recommendations (teachers’ needs vs. their interests) and the presented source of the recommendations (a human expert vs. an AI algorithm) on users’ perceptions of and experience with the system. Our results show an interaction effect between these two system aspects: users who are told that the recommendations are based on their interests have a better experience when the recommendations are presented as originating from an AI algorithm, while users who are told that the recommendations are based on their needs have a better experience when the recommendations are presented as originating from a human expert.
CCS Concepts: • Human-centered computing → Human computer interaction (HCI); • Information systems → Users and interactive retrieval.
Keywords:recommender systems, professional development, algorithm aversion, justification
ACM Reference Format:
Lijie Guo, Christopher Flathmann, Reza Anaraky, Nathan McNeese, and Bart Knijnenburg. 2022. The Effect of Recommendation Source and Justification on Professional Development Recommendations for High School Teachers. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT '22), June 28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA 11 Pages. https://doi.org/10.1145/3511095.3531280
1 INTRODUCTION
The continued improvement of teaching professionals throughout their careers is an essential consideration for modern school systems. Past efforts have often targeted generic professional development methods that only aim to reach required standards, rather than try to fill the specific professional gaps of each individual teacher [14]. Targeting those specific gaps ultimately provides teachers with more effective professional development, which in turn increases the quality of education for the students they teach [10]. While personalized professional development has been implemented on a smaller scale, it has been found difficult to expand to wider audiences due to its reliance on experts collecting and evaluating user data to create personalized education plans [32]. Recommender systems can provide a means of expanding these expert decisions to wider audiences in an efficient manner by rapidly considering user data and recommending those professional development resources that are most beneficial and relevant to each individual teacher [38].
While the fields of entertainment and e-commerce were the first to adopt recommender systems in a commercial setting[3, 41, 42, 50], recommenders have more recently found their way to professional settings as well. Our project considers a system that recommends personalized professional development pathways to high school teachers seeking to increase their disciplinary knowledge and/or their teaching skills. The recommendations provided serve as promoted suggestions when signing up for professional development activities with the goal of guiding teachers towards development opportunities they will both enjoy and highly benefit from.
The first iteration of our recommender system, while simplistic compared to the state-of-the-art in this area, uses real-world teacher data to provide recommendations that benefit their professional development: teachers indicate their interests and needs by filling out a “needs assessment” questionnaire, which gets processed by a rule-based system that assigns weights to various professional development options (ranging from single-course microcredentials, to multi-course endorsements and comprehensive master programs) based on teachers’ answers to the needs assessment questions, subsequently listing the Top 3 options as recommendations. While these recommendations originate from our system, a lot of work by human experts has gone into the development of the “algorithm”—which is essentially a formalization of a vast body of expert knowledge about how teachers’ needs and interests would translate into professional development options that meet these needs and interests.
In addition to the content of the recommendations provided, our system must also consider the presentation of the recommendations: the characteristics of the interface that presents the recommendations to the teachers are critical for ensuring that they carefully consider the recommendations as valid and useful professional development options [7]. In this paper, we investigate two important design considerations regarding the way recommendations are presented to the end-users. Firstly, we acknowledge that teachers’ professional development decisions are driven by their interests (What courses would I like to take?) and their needs (What courses would be most beneficial for me to take?), and these may not always be perfectly aligned. The trade-off between these two possible considerations, especially when justifying the recommendations to the user, provides an important avenue for improving the perception of the recommendations provided. Hence, we posit the following research question:
RQ1: Do teachers have more favorable perceptions of professional development recommendations that are presented as items they would like, or as items that would be most beneficial to them?
Secondly, given that the recommendations originate from a system that embeds a vast amount of human expert knowledge, we ask ourselves whether it would be better to present the recommendations as originating from an AI algorithm or from a human expert. Past research has found conflicting results on this topic—some have carefully documented cases of “algorithm aversion”, where users tend to prefer to receive recommendations from a human rather than an algorithm[11, 25], while others have found situations where algorithmic suggestions are preferred to human suggestions[21, 31, 43]. Given that in our case, one could argue either way about the source of the recommendations, we posit the following research question:
RQ2: Do teachers have more favorable perceptions of professional development recommendations that are presented as originating from an AI algorithm or from a (human) expert?
Finally, recent research has suggested that the “algorithm aversion” phenomenon is task-dependent: it is stronger for subjective tasks and/or hedonic decisions than for objective tasks and/or utilitarian decisions[12]. In this light, one could argue that presenting the recommendations as something the user would like frames the decision as a subjective task and focuses them on the hedonic aspects of the decision, while presenting the recommendations as something that would benefit the user frames the decision as a more objective task and focuses them on the utilitarian aspects of the decision. This could result in an interaction effect between the justification and the source of the recommendations:
RQ3: Does the effect of recommendation source (AI vs. human) on teachers’ perceptions differ depending on the justification for the recommendations (interest vs. needs)?
Our work is the first to reconcile the “algorithm aversion” phenomenon (and its inverse) with research on explainable AI (xAI). In particular, whereas the existing research suggests that “algorithm aversion” depends on the recommendation domain, our study investigates the interaction between the recommendation source and the type of justification within a single domain. A significant effect in our study has considerable practical implications: for one, it would give researchers the opportunity to overcome algorithm aversion—or its inverse—by justifying the recommendations in a needs- or interest-oriented manner. Conversely, our results would provide guidance for recommender system developers to present the recommendations as stemming from either an AI system or a human, depending on how the recommendations are justified.
We answer the stated research questions in the context of our professional development recommender for high school teachers but argue that our core findings are likely applicable to a broader spectrum personalized professional development scenario and perhaps even to recommender systems in general.
2 RELATED WORK
2.1 Recommendations for Professional Development
One of the hallmarks of the modern education system is the continued learning of teachers through professional development [18]. While teachers’ professional development has historically been shown to be a critical component of high quality education, creating a more personalized approach to the professional development of teachers is envisioned to lead to even higher quality teaching, and, in turn, better student outcomes[14]. Specifically, personalized professional development has recently gained a lot of attention, as it can significantly impact the long-term growth of teachers and their abilities[32, 33]. As such, professional development programs that seek to improve education for both teachers and students have recently started using professional development resources that are manually tailored to each individual teacher to provide the most relevant instruction to improve their specific capabilities[32].
While the process of manually tailoring professional development advice to each individual teacher is practically feasible when applied to a handful of teachers, it becomes a much more daunting task when entire schools, districts, and even states are of concern. Rapa et al.[36] argue that a recommender system can be used to make this task manageable: unlike human experts, a carefully designed recommender system can efficiently consider a large amount of teacher data and tailor professional development resource suggestions to each individual user simultaneously—in fact, the more data it receives, the more accurate its recommendations tend to get[38]. The use of recommender systems can take what is traditionally a manual and time-consuming process, embed it into a recommendation algorithm, and thereby extend the results to a much larger population [1]. Specifically, Rapa et al.[36] argue that content-based and expert systems serve as a great starting point for this process, as they can codify the knowledge of human professional development experts into rules that can be used to parse user data into recommendations[37, 39]. Then, once a large portion of teachers run through the program, collaborative filtering approaches can be integrated alongside the rule-based system to create a hybrid system that also considers the experiences fellow teachers had with professional development[44, 48]. Thus, the recommender system's computational capabilities can continue to grow and improve alongside the professional development program it services.
Thus, an initial expert recommender system can be created to rapidly provide teachers with professional development resources that target their specific wants and needs[36]. Since the creation of these recommender systems can be done alongside domain experts, these recommendations can then be reviewed by the creators of both the professional development resources and the academic programs that offer them, and this approval would effectively act as a ground truth baseline[4]. Updates can then be made to the rules considered by the recommender system to better reflect the ground truth data. Ultimately, while the result of this process would be less personable than hand-reviewing each teacher's personal profile, the potential scale at which these recommender systems could operate far outweighs small decreases in individualization[36].
2.2 Justifying Recommendations
Most recommender systems are black boxes, giving users little insight into how the system has modeled and acted upon their preferences. Providing explanations in recommender systems can increase users’ perception of transparency and trust[46, 47]. Initial efforts to explain recommendations can be traced back to recommender systems for news articles, books, and movies[9, 22, 34]. In the two decades following these work, a wide variety of different types of explanations has been developed and tested[17, 19].
An important distinction must be made between explanations and justifications of recommendations: explanations describe the mechanism by which the recommender system arrived at the recommendations, while justifications provide a reason for the recommendation that is independent of the underlying algorithm[45]. In our study, we consider justifications, since it is neither realistic nor necessary for the end-user to understand the exact mechanism by which the recommendations are derived.
Another important distinction considers the goal of the justifications: they could be employed in an attempt to promote the recommendations (i.e., convincing users to adopt the recommendations), or they could be designed to increase users’ knowledge about the recommendations (i.e., allowing them to make more informed decisions about the recommendations)[8]. The idea that the goal of “good explanation should not be to ‘sell’ the user on a recommendation, but rather, to enable the user to make a more accurate judgment of the true quality of an item”[8] is more in line with the goal of recommendations for professional development. Hence, in our study, we specifically aim to help our target users (i.e., high school teachers) make a more accurate judgment of the quality of the recommended pathways.
Given that we aim to provide helpful justifications for our recommendations, the question remains as to what information we should use to justify the recommendations. As mentioned in the introduction, the recommendations our system provides are based on teachers’ professional development needs and their professional interests. We are arguably the first to study whether teachers prefer needs-based vs. interest-based justifications regarding the professional development recommendations.
2.3 Algorithm Aversion and the Source of the Recommendations
Algorithm aversion is the phenomenon that users prefer to receive advice or support from other humans rather than from AI-based systems. Research does not reach a consensus regarding the existence of algorithm aversion. On the one hand, a sizable body of work has shown or acknowledged the existence of algorithm aversion[2, 11, 16, 25, 30, 35]. This work shows that individuals are more likely to delegate strategic decisions to other humans rather than to an AI system [30]. A potential reason for this phenomenon may be that the emotional responses to the outcomes of delegated decisions are more intense when responsibility is delegated to another human being rather than to an AI-enabled system. For example, Promberger et al.[35] compared computer-generated recommendations against physician-generated recommendations in the context of physical health. They found that patients trust a physician more, and consequently, are more likely to follow the physician's advice. On the other hand, the same researchers found that patients feel less responsible when following a physician's recommendation rather than an AI agent's recommendation. Such inconsistency makes algorithm aversion an important topic for research.
Contrasting the work on algorithm aversion is a growing body of empirical evidence suggesting that users actually prefer algorithmic advice[15, 21, 31, 43]. For example, Dijkstra et al. found that individuals find expert systems more rational than human advisors [15]. Gunaratne et al. studied human decisions in the context of an online retirement saving system[21], they found that while both types of advice increase users’ saving performance, users are more likely to follow advice coming from an algorithmic source rather than crowd-source advice.
More recent studies suggest that the occurrence of algorithm aversion or algorithm seeking behaviors crucially depends on external factors. For example, Beger et al. revealed that users do not prefer human advice to algorithmic advice when they are unfamiliar with the human advisor[6]. Similarly, Castelo et al. explored the ’algorithm aversion’ phenomenon in 6 studies with different tasks[12] and found that the algorithm aversion phenomenon is task-dependent. The findings of their 6 studies were demonstrated in a conceptual model suggesting that the objectivity of the tasks decreases users’ discomfort with the algorithms and increases their perceptions of algorithm effectiveness and willingness to rely on algorithms—for objective tasks, users prefer the algorithm, while for subjective tasks, they prefer human advice. Castelo et al. note, though, that if the algorithm exhibits high affective human-likeness, this reduces the effect.
As a result of a literature review, Jussupow et al. [25] concluded that the existence of algorithm aversion may relate to other parameters such as performance, perceived capabilities, human involvement, human agents’ expertise, as well as social distance. Therefore, research needs to investigate these various different parameters and domains to better understand the algorithm aversion phenomenon. Indeed, one of the goals of our study is to understand how users’ perceptions of AI algorithm-based recommendations differentiate from those of human expert-suggested recommendations in an educational setting, with either interest-driven justification or needs-driven justification for the recommendations. As Logg et al. pointed out, ”algorithm appreciation” may appear in some domains where the algorithm has been historically used and popularly accepted by most people such as weather forecasts[31], which means the prior experience with the algorithmic or human advice could be a confounding factor in the phenomenon of algorithm aversion. Since our project considers such a system recommending personalized professional development pathways that specifically targets high school teachers, it avoids the possible additional effects that come from prior experiences. Arguably, this work reveals implications for the design of recommender systems for professional development and beyond.
3 STUDY DESIGN
We tested the effect of recommendation source (AI vs. human) and justification type (interest vs. needs) on teachers’ perceptions of the system in a scenario-based controlled experiment using a prototype of the real system. Testing users’ reactions to AI-based systems with prototypes is a common practice in HCI research[13]. While we plan to eventually test the effects of recommendation source and justification in our field trial where teachers receive real personalized professional development pathway recommendations, we decided to first run a more tightly controlled experiment to study these effects outside the inherently noisy environment of a field trial. In our controlled experiment, all participants receive the same recommendations after reading a scenario describing the abilities, needs and preferences that ostensibly served as input for these recommendations (more details in Section3.3). By using this scenario-based study setup, we can benefit from not having individual recommendation quality interfere with the effect of the presentation of the recommendations (i.e., recommendation source and justification). This setup also makes it possible to manipulate the justification type, as it allowed us to create a scenario where both interests and needs may conceivably underlie the recommendations. Conversely, in the real world some teachers’ recommendations follow only their interests or only their needs, making it impossible to claim otherwise (and thus difficult to manipulate the justification type).
Furthermore, the sample of our study is recruited via Prolific, and is thus different from the teachers in our real-world field trial. We took special care, though, that the participants in this study also identified as teachers, and that the recommendations were accompanied by a realistic scenario for which the presented recommendations would be appropriate. This ascertained that participants would be able to interpret the quality of the provided recommendations and (more importantly) the justifications, as their background as teachers would allow them to personally relate to the presented scenario. The remainder of this section outlines the setup of the controlled experiment. We note that our study procedures were approved by our Institutional Review Board (IRB).
3.1 Participants
We recruited participants for our experiment on Prolific, an online recruitment platform that has a set of detailed filters that can be used to target a particular sample of participants. Using these filters, we limited participation to adult teachers with a completed undergraduate degree or higher, so that the participants would closely match the users of the actual system under development. 207 participants took part in the study, yielding 190 usable data points after filtering out 17 participants who did not carefully read the presented information. Of the 190 participants, 120 identified as women, 66 as men, 3 as non-binary, and 1 participant preferred not to answer our gender question. The sample includes 13 participants between the ages of 18 and 24, 82 between 25 and 34, 55 between 35 and 44, 24 between 45 and 54, and 16 older than 54. Most participants completed the study in 6 minutes. They each received USD 1.20 for their participation.
3.2 Procedure
Participants were shown a welcome page with an introduction to our study and a consent form containing statements of possible risks, discomforts, and incentives. We then presented them with a scenario asking them to imagine being a teacher with certain professional development needs/interests (see Section 3.3), followed by a reading comprehension check question (“Which of the following is not true based on the scenario?”).
Next, participants were shown two screenshots of the proposed system: an information screen and a recommendation screen. The information screen (Figure 1) welcomed the participant to the system and explained that the recommendations were provided by either a human expert or an AI-based algorithm. The recommendation screen (Figure 2) displayed the recommended professional development pathways, including justifications for the recommendation process, as well as each of the individual recommendations, based on either the interests or the needs of the imagined teacher. This screen was followed by another reading comprehension check question1 (“Which of the following is not one of the recommendations?”).
Finally, participants were asked to answer a survey containing 38 questions (see Section 3.5) measuring their opinions about and user experience with the presented system. The user study procedure is shown in 3.
Figure 1:The information screen in the two source conditions: AI algorithm (left) and human expert (right).
Figure 2:The recommendation screen in the four conditions. Top left: recommendations based on interests presented by an AI algorithm; top right: recommendations based on needs presented by an AI algorithm; bottom left: recommendations based on interests presented by a human expert; bottom right: recommendations based on needs presented by a human expert.
Figure 3:The procedure of the online experiment.
3.3 Scenario
To provide enough context for participants to understand the recommendations, we created a scenario asking participants to imagine that they are a teacher with a carefully selected set of professional development needs and interests (Figure 4). To make the scenario match the source of the recommendations, it emphasized the teacher's interests for participants in the “interests” conditions, while emphasizing the teacher's needs for participants in the “needs” conditions. Note that this difference is merely one of presentation—the content of the two versions of the scenario remained the same. Also note that the participants all identified as teachers themselves, making the scenario (which was rooted in real-world teacher data) easily relatable.
Figure 4:The scenario presented to participants was manipulated alongside the justification manipulation: on the left, interests are presented as the primary source preferences and needs are presented as secondary; on the right, needs are presented as the primary source of preferences and interests are presented as secondary.
3.4 Experimental Manipulations
The experiment involved two between-subjects manipulations: recommendation source and justification. The recommendation source was presented as either an AI algorithm or a human expert. The source was printed in bold and accompanied by a robot or human avatar on both the information screen (Figure1) and the recommendation screen (Figure2). The recommendation source manipulation only manipulated whether the recommendations were provided by an AI algorithm or a human expert—all other information on the screens was kept as similar as possible, so as to be able to particularly test the effect of the source of the recommendations.
The justifications for the recommendations were presented on the recommendation screen (Figure2) as either the teacher's interests (“The [source] thinks you would like the following recommendations based on the information you provided.”) or their needs (“The [source] thinks the following recommendations would be most beneficial to you based on the information you provided.”). Furthermore, the “reason” listed for each recommendation also reflected the teacher's interests or their needs, depending on the experimental condition. Finally, as mentioned above, the scenario was framed in such a way that these reasons would match the reasons presented in the scenario.
We randomly assigned participants to one of the four experimental conditions in a 2 × 2 between-subjects design—a between-subjects manipulation was used to increase ecological validity and to prevent “demand characteristics” from influencing the study[28].
3.5 Dependent Variables
We used the following eight scales (adopted from related work) to measure participants’ perceptions of the system attributes and user experience with the system presented in the scenarios with eight subjective measurement constructs:
Understandability: participants’ self-reported understanding of the recommendation process, as derived from the justifications (adopted from [26, 27]).
Perceived recommendation quality: participants’ perception of the recommendation quality (adopted from [29]).
Perceived system effectiveness: participants’ perception of the effectiveness of the system (adopted from [29]).
Explainability: participants’ perception of how well the provided justifications explained the recommendations.
Fit with preference: the perceived fit of the recommendation with the participant's preferences (adopted from [20]).
Competence belief: participants’ perception of the ability, skills, and expertise of the system to perform effectively in its specific domain (a sub-scale of trust, adopted from [49]).
Benevolence belief: participants’ perception that the system cares about the consumer and acts in the consumer's interest (a sub-scale of trust, adopted from [49]).
Integrity belief: participants’ perception that the system adheres to a set of principles (e.g., honesty and keeping promises) that are generally accepted by consumers (a sub-scale of trust, adopted from [49]).
Each scale consists of multiple items, and a total of 38 items were administered in the questionnaire. Participants were asked to rate each item on a 5-point agreement scale (from strongly disagree to strongly agree).
4 RESULTS
Confirmatory Factor Analysis (CFA) was performed to validate the subjective scales that serve as dependent variables in our experiment. We subsequently fitted a Structural Equation Model (SEM) that demonstrates the causal relationships between the manipulations and the validated subjective constructs, as well as mediation effects.
4.1 Measurement Model (CFA)
Our CFA indicated 4 questionnaire items with either low loadings (< 0.70) or high modification indices (both of which indicate misfit). These items were removed from subsequent analyses. While all factors had an adequate convergent validity (AVE > 0.50)2, we found that several factors showed a lack of discriminant validity3. Particularly, we found that explainability was too highly correlated with perceived recommendation quality, perceived system effectiveness, competence belief, and benevolence belief; Perceived recommendation quality was too highly correlated with fit with preference; and competence belief was too highly correlated with perceived system effectiveness and benevolence belief. To avoid multicollinearity in our subsequent SEM model, we removed explainability, perceived recommendation quality, competence belief, and benevolence belief from subsequent analyses4.
We again performed a CFA with the remaining four factors (i.e., understandability, fit with preference, perceived system effectiveness, and integrity belief). In this analysis, one additional item was dropped from the model due to a low loading (< 0.70). The consistency coefficients (Cronbach's α) of the final four factors showed high to excellent scale reliabilities5 and the AVE of the four factors ranged from 0.745 to 0.885, indicating that the 4 constructs meet convergent validity requirements (see Table1). The final 4-factor model also meets the discriminant validity requirements (see Table 2).
Table 1:Items of the 4-factor model. Items without a factor loading were excluded from the analysis.
Considered aspects | Item | Factor loading |
|---|---|---|
Understandability AVE: 0.748 Cronbach's α: 0.91 | I understand how the system came up with the recommendations. | 0.871 |
The recommender explained the reasoning behind the recommendations. | 0.812 | |
I am unsure how the recommendations were generated. | -0.888 | |
The recommendation process is clear to me. | 0.878 | |
The recommendation process is not transparent. | -0.874 | |
Effectiveness AVE: 0.745 Cronbach's α: 0.92 | I would recommend the system to others. | 0.954 |
The system is useless. | -0.904 | |
The system makes me more aware of my choice options. | 0.768 | |
I make better choices with the system. | 0.884 | |
I can find better pathways without the help of the system. | -0.789 | |
I can find better pathways using the recommender system. | ||
The system showed useful pathways. | 0.904 | |
Fit with preference AVE: 0.885 Cronbach's α: 0.86 | The recommended pathways reflect what I want. | 0.938 |
The recommended pathways suit my needs. | 0.964 | |
The recommended pathways are exactly what I want. | 0.919 | |
Integrity AVE: 0.757 Cronbach's α: 0.83 | This system provides unbiased pathway recommendations. | 0.770 |
This system is honest. | 0.879 | |
I consider this system to be of integrity. | 0.952 |
Table 2:Factor-fit metrics. Off-diagonal values are correlations, diagonal values are the square roots of the average variance extracted $left(right. sqrt{A V E} left.right)$ per factor.
U | E | F | I | |
|---|---|---|---|---|
U nderstandability | 0.865 | 0.651 | 0.582 | 0.609 |
E ffectiveness | 0.651 | 0.863 | 0.780 | 0.779 |
F it with preference | 0.582 | 0.780 | 0.941 | 0.720 |
I ntegrity | 0.609 | 0.779 | 0.720 | 0.870 |
4.2 Structural Equation Model (SEM)
A Structural Equation Model (SEM) was fitted to the four constructs and the experimental manipulations (i.e., recommendation source and justification). An SEM enables one to specify the relationships between exogenous variables (the manipulations) and latent constructs (the CFA factors) as a structured model of regressions[28]. An important benefit of SEM is that fit statistics are provided for the model as a whole, as well as for the individual regression coefficients. We built our model following two principles:
Justification (RQ1), recommendation source (RQ2) and their interaction (RQ3) were hypothesized to influence understandability, fit with preference, integrity, and system effectiveness.
Each effect outlined in (1) is allowed to mediate the subsequent effects (e.g., understandability is allowed to mediate the effect of justification and/or recommendation source on fit with preference).
We first specified a saturated model with all hypothesized effects and mediations. We then iteratively trimmed non-significant effects. The resulting model is displayed in Figure 5. This model has a good overall fit with χ 2(158) = 246.918, p< 0.001, CFI = 0.990, TLI = 0.991, RMSEA = 0.055 with a 90% confidence interval of [0.041, 0.067]6.
The model shows that the recommendation source and justification manipulations have significant interaction effects on the dependent variables. Particularly, the manipulations have a significant interaction effect on the understandability of the system (p =.036) and a marginally significant effect on participants’ perceived system effectiveness (p =.098); understandability mediates the interaction effects on the perceived fit of the recommendations with the teacher's presented preferences and on participants’ perception of the integrity of the system.
Figure 5:The structural equation model for the data of the experiment.
Figure6 displays the total (mediated + direct) effects of the two manipulations on the dependent variables. The total interaction effect between source and justification is significant for all dependent variables—understandability (p =.036), fit with preference (p =.036), integrity (p =.040), and effectiveness (p =.011). In particular, Figure6 shows that when participants were told that the source of the recommendation is a human expert, there was no significant difference in understandability, fit with preference, integrity, or effectiveness between participants who were told that the recommendations were based on the teacher's interests vs. the teacher's needs. However, among participants who were told that the source of the recommendation is an AI algorithm, those who were told that the recommendations are based on the teacher's interests perceived a significantly higher level of understandability (a large, significant effect; Cohen's d = 0.78, p<.001), fit with preference (a medium-sized, significant effect; Cohen's d = 0.56, p<.001), integrity (a large, significant effect; Cohen's d = 0.74, p<.001), and effectiveness (a large, significant effect; Cohen's d = 1.12, p<.01) than participants who were told that the recommendations are based on the teacher's needs.
Figure 6:Total effects of recommendation source and justification on the perceived understandability (left) and the system effectiveness (right). The effect of the “Human” source with the ”Needs” based justification condition is set to zero, and the y-axis is scaled by the sample standard error.
5 DISCUSSION
5.1 Revisiting the Research Questions, and Comparison to Related Work
We set out to test the effects of justification (interests vs. needs, RQ1), recommendation source (human expert vs. AI algorithm, RQ2), and their interaction (RQ3) on teachers’ perceptions of and experience with a personalized professional development pathway recommender in a 2 × 2 between-subjects controlled experiment. In light of our research questions, the results show that an interest-based justification outperforms a needs-based justification (RQ1), but only for users who are told that the recommendations originate from an AI algorithm rather than a human expert (RQ3). Conversely, the effect of presenting a human expert vs. an AI algorithm as the source of the recommendations (RQ2) completely depends on the presented justification (RQ3): users who are told that the recommendations are based on their interests have a better experience when the recommendations are presented as originating from an AI algorithm, while users who are told that the recommendations are based on their needs have a better experience when the recommendations are presented as originating from a human expert.
Notably, the uncovered interaction effect runs counter to existing research, which shows that the “algorithm aversion” phenomenon is stronger for subjective tasks and/or hedonic decisions than for objective tasks and/or utilitarian decisions[12]—assuming that interest-based justifications align with a perception of the recommendations as subjective/hedonic while needs-based justifications align with a perception of the recommendations as objective/utilitarian, our results show that subjective/hedonic recommendations actually perform better when presented by an AI algorithm, while objective/utilitarian recommendations perform worse. Perhaps, then, there is no clear connection between interest vs. needs-based justifications and subjective vs. objective tasks. Research has shown that lay people perceive a task that can be approached by measuring and analyzing relevant quantitative variables as objective, while perceiving a task that can be approached using intuition or gut feelings as subjective[24]. From this perspective, both types of justifications can be considered objective, as each version portrays the system as taking a decidedly calculative approach to the recommendation process.
A possible alternative explanation for our findings is that while both interest and needs-based justifications are considered objective (and hence principally better suited for an AI system), users find a needs-based justifications condescending when coming from an AI system—if so, needs-based recommendations would indeed best be presented as originating from a human expert. Alternatively, one could argue that needs-based recommendations for teachers’ professional development (i.e., recommendations that may have a serious impact on their career) are more consequential than interest-based recommendations (i.e., recommendations that simply align with what they enjoy)—if so, teachers may be less willing to trust in algorithms for (high-risk) needs-based recommendations[35] than for (low-risk) interest-based recommendations[31].
5.2 Design Implications
Overall, users are most satisfied with interest-based recommendations presented by an AI algorithm—according to the results in Figure6, users find this system the most understandable, they find that the recommendations better fit their preferences, they find that the system has a higher level of integrity, and they find the system more effective. Arguably, users are most excited about interest-based recommendations, but they may believe that only an AI algorithm would be able to handle the complexity of translating their nuanced interests into a series of recommended professional development activities.
Given these results, we suggest that, when both are possible, the presentation of recommendations should emphasise their algorithmic nature, and the justification of recommendations should relate back to users’ interests over their needs. In our study, this presentation was implemented both visually and textually, with the system displaying a robot-like icon and the text explicitly stating that “the AI algorithm thinks you would like the following recommendations”, and the justification was implemented with each individual recommendation having a reference to the teacher's interests (see Figure2). Furthermore, while the design elements chosen were not designed to be overly distracting, the medium to large effect sizes of our study would indicate that more subtle designs implementations could still be effective.
Conversely, an AI algorithm based presentation of recommendations justified by users’ needs performed the worst. As mentioned above, one possible explanation for this could be that users find a needs-based explanation condescending when presented by an AI system; another possible explanation is that needs-based explanations are too consequential to trust to an AI system. Regardless, needs-based recommendations are best presented as originating from a human expert.
This finding has implications for situations where recommendations are exclusively based on needs—e.g., in situations where interests are not elicited, or where the recommender system is built to prioritize needs in cases where a system prioritizes targeting a user's deficiencies or needs over their interests. In these cases, it would be disingenuous to justify the recommendations by referring to the user's interests. Instead, when justifying recommendations with a user's needs, the system's presentation should downplay the algorithmic nature of the recommendation selection process. In our study, we did this by displaying a human-like icon on the recommendation page, and by explicitly mentioning that “the expert thinks the following recommendations would be most beneficial to you” (see Figure2). Note that one does not have to completely hide the involvement of a recommender system—on the introduction page of the human source condition, we did explicitly mention that “the expert has entered your preferences and information into the system and calculated the most relevant professional development options for you.” (see Figure1). Rather the presentation of these recommendations (i.e., the delivery of them to the user) should be perceived as coming from a human expert rather than an algorithmic system.
5.3 Limitations and Future Work
An obvious limitation of this work is that in an effort to control the quality of the recommendations and the justification between subjects, the manipulations were introduced in a scenario-based experiment rather than a real recommender system. To some extent, this limits participants’ deeper understanding of the personal relevance of the presented recommendations, and the justifications alike. To mitigate this limitation, we carefully outlined a scenario explaining the fictitious teacher's needs and interests. Moreover, we made sure to recruit participant among actual teachers, who are arguably more qualified to understand the presented scenario and recommendations than the general population.
In our future work, we will confirm these effects in the real personalized professional development pathway recommender and attempt to verify these potential reasons for the uncovered effects. The deployment of this study in the “live” system will also give us the opportunity to test the effects on users’ choice behavior: do the recommendation source and justification type have an effect on which and how many professional development items they agree to enroll in? And are there perhaps differences between conditions in terms of users’ attrition rates (e.g. dropping classes or abandoning them mid-semester—something that tends to happen frequently, as teachers have to balance their professional development commitments with the demands of their teaching job and their personal lives)? While our current study focused on opinions, the upcoming study with the real system will provide a unique opportunity to carefully study these behavioral effects as well. That said, experimental control will be more difficult in the “live” system, since the teachers will approach the system with different goals, constraints, and ambitions. Hence, the current, more carefully controlled study provides valuable insights into the attitudinal effects of recommendation source and justification type—the “live” system study will complement these results.
Another limitation of our study is that in order to carefully single out the effect of the recommendation source (i.e., to not ascribe differing capabilities to either source), the recommendations are presented as the outcome of a system supporting the recommendation calculation process, even in the “human expert” condition. This may have given our scenario a more calculative emphasis, regardless of the recommendation source or the justification type. To further emphasize the difference between human and AI recommendations, future work could present the recommendations in the “human expert” condition as manually curated rather than calculated. One must however be careful about the ethical ramifications of misrepresenting the true source of the recommendations.
6 CONCLUSION
In this paper, we presented a study to investigate the best way to present recommendations to teachers seeking to advance their professional development. In a carefully controlled, scenario-driven experiment, we tested the effect of the justification behind the recommendations (i.e., the teachers’ personal interests vs. their needs) and the source of the recommendations (i.e., a human expert vs. an AI algorithm). The results show that this recommender system benefits teachers most if they are told that the recommendations originate from an AI algorithm and are based on their interests. In our future work, we will confirm the uncovered interaction effect in the real recommender system and attempt to verify its underlying cause.
ACKNOWLEDGMENTS
This research was supported by the U.S. Department of Education (Award number S423A20008).
REFERENCES
Muhammad Tanvir Afzal and Hermann A Maurer. 2011. Expertise Recommender System for Scientific Community.J. Univers. Comput. Sci. 17, 11 (2011), 1529–1549.
Jorge A Alvarado-Valencia and Lope H Barrero. 2014. Reliance, trust and heuristics in judgmental forecasting. Computers in human behavior 36 (2014), 102–113.
Amos Azaria, Avinatan Hassidim, Sarit Kraus, Adi Eshkol, Ofer Weintraub, and Irit Netanely. 2013. Movie recommender system for profit maximization. In Proceedings of the 7th ACM conference on Recommender systems. 121–128.
Joeran Beel, Marcel Genzmehr, Stefan Langer, Andreas Nürnberger, and Bela Gipp. 2013. A comparative analysis of offline and online evaluations and discussion of research paper recommender system evaluation. In Proceedings of the International Workshop on Reproducibility and Replication in Recommender Systems Evaluation(RepSys ’13). Association for Computing Machinery, New York, NY, USA, 7–14. https://doi.org/10.1145/2532508.2532511
Peter M Bentler and Douglas G Bonett. 1980. Significance tests and goodness of fit in the analysis of covariance structures.Psychological bulletin 88, 3 (1980), 588.
Benedikt Berger, Martin Adam, Alexander Rühr, and Alexander Benlian. 2021. Watch Me Improve—Algorithm Aversion and Demonstrating the Ability to Learn. Business & Information Systems Engineering 63, 1 (2021), 55–68.
Émilie Bigras, Pierre-Majorique Léger, and Sylvain Sénécal. 2019. Recommendation agent adoption: how recommendation presentation influences employees’ perceptions, behaviors, and decision quality. Applied Sciences 9, 20 (2019), 4244.
Mustafa Bilgic and Raymond J Mooney. 2005. Explaining recommendations: Satisfaction vs. promotion. In Beyond personalization workshop, IUI, Vol.5. 153.
Daniel Billsus and Michael J Pazzani. 1999. A personal news agent that talks, learns and explains. In Proceedings of the third annual conference on Autonomous Agents. 268–275.
Hilda Borko and Carol Livingston. 1989. Cognition and improvisation: Differences in mathematics instruction by expert and novice teachers. American educational research journal 26, 4 (1989), 473–498.
Jason W Burton, Mari-Klara Stein, and Tina Blegind Jensen. 2020. A systematic review of algorithm aversion in augmented decision making. Journal of Behavioral Decision Making 33, 2 (2020), 220–239.
Noah Castelo, Maarten W. Bos, and Donald R. Lehmann. 2019. Task-Dependent Algorithm Aversion. Journal of Marketing Research 56, 5 (2019), 809–825. https://doi.org/10.1177/0022243719851788 arXiv:https://doi.org/10.1177/0022243719851788
Fabio Colella, Pedram Daee, Jussi Jokinen, Antti Oulasvirta, and Samuel Kaski. 2020. Human strategic steering improves performance of interactive optimization. In Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization. 293–297.
Linda Darling-Hammond, Maria E. Hyler, and Madelyn Gardner. 2017. Effective Teacher Professional Development. Learning Policy Institute. https://eric.ed.gov/?id=ED606743 ISSN: ISSN- Publication Title: Learning Policy Institute.
Jaap J Dijkstra, Wim BG Liebrand, and Ellen Timminga. 1998. Persuasiveness of expert systems. Behaviour & Information Technology 17, 3 (1998), 155–163.
Aaron C Elkins, Norah E Dunbar, Bradley Adame, and Jay F Nunamaker. 2013. Are users threatened by credibility assessment systems?Journal of Management Information Systems 29, 4 (2013), 249–262.
Gerhard Friedrich and Markus Zanker. 2011. A Taxonomy for Generating Explanations in Recommender Systems. AI Magazine 32, 3 (June 2011), 90–98. https://doi.org/10.1609/aimag.v32i3.2365
Michael Fullan. 2007. The New Meaning of Educational Change. Routledge. Google-Books-ID: dvc84eFzKkkC.
Fatih Gedikli, Dietmar Jannach, and Mouzhi Ge. 2014. How should I explain? A comparison of different explanation types for recommender systems. International Journal of Human-Computer Studies 72, 4 (April 2014), 367–382. https://doi.org/10.1016/j.ijhcs.2013.12.007
Ulrike Gretzel and Daniel R Fesenmaier. 2006. Persuasion in recommender systems. International Journal of Electronic Commerce 11, 2 (2006), 81–100.
Junius Gunaratne, Lior Zalmanson, and Oded Nov. 2018. The persuasive power of algorithmic and crowdsourced advice. Journal of Management Information Systems 35, 4 (2018), 1092–1120.
Jonathan L. Herlocker, Joseph A. Konstan, and John Riedl. 2000. Explaining collaborative filtering recommendations. In Proc. of the 2000 ACM conference on Computer supported cooperative work. ACM Press, Philadelphia, PA, 241–250. https://doi.org/10.1145/358916.358995
Li-tze Hu and Peter M Bentler. 1999. Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural equation modeling: a multidisciplinary journal 6, 1(1999), 1–55.
Yoel Inbar, Jeremy Cone, and Thomas Gilovich. 2010. People's intuitions about intuitive insight and intuitive choice.Journal of personality and social psychology 99, 2(2010), 232.
Ekaterina Jussupow, Izak Benbasat, and Armin Heinzl. 2020. Why are we averse towards Algorithms? A comprehensive literature Review on Algorithm aversion. (2020).
Bart P. Knijnenburg, Svetlin Bostandjiev, John O'Donovan, and Alfred Kobsa. 2012. Inspectability and control in social recommenders. In Proceedings of the sixth ACM conference on Recommender systems(RecSys ’12). ACM, New York, NY, USA, 43–50. https://doi.org/10.1145/2365952.2365966
Bart P. Knijnenburg, Nikhil Rao, and Alfred Kobsa. 2012. Experimental Materials Used in the Study on Inspectability and Control in Social Recommender Systems. Institute of Software Research, UC Irvine: Technical Report UCI-ISR-12-4.
Bart P Knijnenburg and Martijn C Willemsen. 2015. Evaluating recommender systems with user experiments. In Recommender systems handbook. Springer, 309–352.
Bart P Knijnenburg, Martijn C Willemsen, Zeno Gantner, Hakan Soncu, and Chris Newell. 2012. Explaining the user experience of recommender systems. User Modeling and User-Adapted Interaction 22, 4 (2012), 441–504.
Michael Leyer and Sabrina Schneider. 2019. Me, you or AI? How do we feel about delegation. (2019).
Jennifer M Logg, Julia A Minson, and Don A Moore. 2019. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes 151 (2019), 90–103.
Jeff C. Marshall and Daniel M. Alston. 2014. Effective, Sustained Inquiry-Based Instruction Promotes Higher Science Proficiency Among All Groups: A 5-Year Analysis. Journal of Science Teacher Education 25, 7 (Nov. 2014), 807–821. https://doi.org/10.1007/s10972-014-9401-4
Monica Martinez. 2019. Personalization turns learning into a journey. The Learning Professional 40, 4 (2019), 9–12. Publisher: National Staff Development Council.
Raymond J Mooney and Loriene Roy. 2000. Content-based book recommending using learning for text categorization. In Proceedings of the fifth ACM conference on Digital libraries. 195–204.
Marianne Promberger and Jonathan Baron. 2006. Do patients trust computers?Journal of Behavioral Decision Making 19, 5 (2006), 455–468.
Luke J. Rapa, Jeff C. Marshall, Stephanie M. Madison, Christopher Flathmann, Bart P. Knijnenburg, and Nathan J. McNeese. 2022. Clemson University's Teacher Learning Progression Program: Personalized Advanced Credentials for Teachers. In Handbook of Research on Credential Innovations for Inclusive Pathways to Professions. https://doi.org/10.4018/978-1-7998-3820-3.ch016 ISBN: 9781799838203 Pages: 313-334 Publisher: IGI Global.
Tajul Rosli Razak, Muhamad Arif Hashim, Noorfaizalfarid Mohd Noor, Iman Hazwam Abd Halim, and Nur Fatin Farihin Shamsul. 2014. Career path recommendation system for UiTM Perlis students using fuzzy logic. In 2014 5th International Conference on Intelligent and Advanced Systems (ICIAS). 1–5. https://doi.org/10.1109/ICIAS.2014.6869553
Paul Resnick and Hal R. Varian. 1997. Recommender systems. Commun. ACM 40, 3 (March 1997), 56–58. https://doi.org/10.1145/245108.245121
Francesco Ricci. 2002. Travel recommender systems. IEEE Intelligent Systems 17, 6 (2002), 55–57.
Ian Ruginski. 2019. Structural Equation Modeling in R Tutorial 6: Confirmatory Factor Analysis using lavaan in R. https://www.ianruginski.com/post/semhandout6/
J Ben Schafer, Joseph A Konstan, and John Riedl. 2001. E-commerce recommendation applications. Data mining and knowledge discovery 5, 1 (2001), 115–153.
Markus Schedl, Peter Knees, Brian McFee, Dmitry Bogdanov, and Marius Kaminskas. 2015. Music recommender systems. In Recommender systems handbook. Springer, 453–492.
Donghee Shin, Bouziane Zaid, and Mohammed Ibahrine. 2020. Algorithm Appreciation: Algorithmic Performance, Developmental Processes, and User Interactions. In 2020 International Conference on Communications, Computing, Cybersecurity, and Informatics (CCCI). IEEE, 1–5.
Xiaoyuan Su and Taghi M Khoshgoftaar. 2009. A survey of collaborative filtering techniques. Advances in artificial intelligence 2009 (2009).
P. Symeonidis, A Nanopoulos, and Y. Manolopoulos. 2008. Providing Justifications in Recommender Systems. IEEE Transactions on Systems, Man and Cybernetics, Part A: Systems and Humans 38, 6 (Nov. 2008), 1262–1272. https://doi.org/10.1109/TSMCA.2008.2003969
Nava Tintarev. 2007. Explanations of recommendations. In Proceedings of the 2007 ACM conference on Recommender systems. 203–206.
Nava Tintarev and Judith Masthoff. 2012. Evaluating the effectiveness of explanations for recommender systems. User Modeling and User-Adapted Interaction 22, 4-5 (Oct. 2012), 399–439. https://doi.org/10.1007/s11257-011-9117-5
Bogdan Walek and Vladimir Fojtik. 2020. A hybrid recommender system for recommending relevant movies using an expert system. Expert Systems with Applications 158 (2020), 113452.
Weiquan Wang and Izak Benbasat. 2007. Recommendation agents for electronic commerce: Effects of explanation facilities on trusting beliefs. Journal of Management Information Systems 23, 4 (2007), 217–246.
Kangning Wei, Jinghua Huang, and Shaohong Fu. 2007. A survey of e-commerce recommender systems. In 2007 international conference on service systems and service management. IEEE, 1–5.
FOOTNOTE
1The two reading comprehension check questions each had 4 options. Participants were allowed to fail each question twice before answering correctly. If they failed the question a third time, they would be redirected to the end of the survey, and we would discard their data.
2Average Variance Extracted (AVE) is an indicator of the convergent validity of the measurement scales, with the recommended lower bound threshold of 0.5.
3Discriminant validity is established when the $sqrt{A V E}$ of a factor is larger than its correlations with each of the other factors
4Note that the concept of explainability is still represented in the model by its close analog understandability, and the concept of perceived recommendation quality is represented by fit with preference. In addition, while we removed two trust factors (competence and benevolence), one trust factor (integrity) remains in the model.
5A commonly used rule of thumb is that an α of 0.7 indicates acceptable reliability, 0.8 or higher indicates good reliability, and 0.9 or higher indicates excellent reliability[40].
6Theoretically, a good model is not statistically different from the fully specified model (i.e., the p-value of the χ 2 should be > 0.05), but this statistic is commonly regarded as too sensitive[5]. As such, Hu and Bentler proposed cut-off values for the alternative fit indices to be: CFI > 0.96, TLI >0.95, and RMSEA < 0.05, with the upper bound of its 90% CI falling below 0.10 based on extensive simulations[23].
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org.
HT '22, June 28–July 01, 2022, Barcelona, Spain
© 2022 Association for Computing Machinery.
ACM ISBN 978-1-4503-9233-4/22/06…$15.00.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime