Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis
Converted the full HT'24 paper on Remote Entailment Analysis (a group-creativity measurement procedure comparing human dyads to ChatGPT) to faithful Markdown with all 15 figures cropped from the PDF and Table 1 transcribed, plus a stub.md since the paper carries no CC license.

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis

Authors: Daniel Baumartz, Maxim Konca, Alexander Mehler, Patrick Schrottenbacher, Dominik Braunheim Published in: HT '24: 35th ACM Conference on Hypertext and Social Media, Poznan, Poland, September 10-13, 2024 DOI: 10.1145/3648188.3675140 License: © Copyright held by the owner/author(s).

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis

Daniel Baumartz Goethe University Frankfurt Frankfurt, Germany baumartz@em.uni-frankfurt.de

Maxim Konca Goethe University Frankfurt Frankfurt, Germany konca@em.uni-frankfurt.de

Patrick Schrottenbacher Goethe University Frankfurt Frankfurt, Germany schrottenbacher@em.unifrankfurt.de

Abstract

We present a procedure for assessing group creativity that allows us to compare the contributions of human interlocutors and chatbots based on generative AI such as ChatGPT. We focus on everyday creativity in terms of dialogic communication and test four hypotheses about the difference between human and artificial communication. Our procedure is based on a test that requires interlocutors to cooperatively interpret a sequence of sentences for which we control for coherence gaps with reference to the notion of entailment. Using NLP methods, we automatically evaluate the spoken or written contributions of interlocutors (human or otherwise). The paper develops a routine for automatic transcription based on Whisper, for sampling texts based on their entailment relations, for analyzing dialogic contributions along their semantic embeddings, and for classifying interlocutors and interaction systems based on them. In this way, we highlight differences between human and artificial conversations under conditions that approximate free dialogic communication. We show that despite their obvious classificatory differences, it is difficult to see clear differences even in the domain of dialogic communication given the current instruments of NLP.

CCS Concepts

• Applied computing →Sociology; • Computing methodologies →Discourse, dialogue and pragmatics; Natural language generation; Speech recognition; Cluster analysis; Topic modeling; Supervised learning by classification; Classification and regression trees; Support vector machines; Neural networks.

Keywords

Creativity, Generative AI, Creative AI, NLP, Hermeneutics

Alexander Mehler Goethe University Frankfurt Frankfurt, Germany mehler@em.uni-frankfurt.de

Dominik Braunheim Johannes Gutenberg University Mainz Mainz, Germany dobraunh@uni-mainz.de

Interaction Systems by Means of Remote Entailment Analysis. In 35th ACM Conference on Hypertext and Social Media (HT ’24), September 10–13, 2024, Poznan, Poland. ACM, New York, NY, USA, 14 pages. https://doi.org/10.1145/ 3648188.3675140

1 INTRODUCTION

We present a procedure for the analysis of dialogic group creativity that allows us to compare the contributions of human interlocutors and AI chatbots such as ChatGPT. Using NLP methods, we automatically evaluate the dialogic contributions of interlocutors (human or artificial). In this way, we shed light on the (dis-)similarities of human and artificial conversational spaces across time scales. In order to underpin the dialogic character and group orientation of our creativity analysis, we follow objective hermeneutics [46]. That is, we design a test that is based on the sequential interpretation of sentences, for which we automatically guarantee a certain degree of coherence with sufficient semantic distance between them to simultaneously encourage the test subjects to make open, creative and text-guided interpretations. In this way, we develop a test scenario that addresses the creativity of groups, is based on extensive linguistic interactions between interlocutors, and allows fully automated comparisons with artificial chatbots. In contrast to other language-based creativity tests, which are based on the written output of test subjects in rather artificial laboratory situations, mostly at the lexical level, we are concerned with the openness of spoken language in the context of dialogic communication. In this sense, we are dealing with everyday (little-c) creativity [63], which is not person-related [6], but group-related. Approaches to measuring creativity in cognitive science involve association experiments with individuals. This includes lexical association tests such as the Remote Association Test [42], the Divergent Association Task (DAT) [47], the Alternative Uses Test (AUT) [22]. In the age of generative AI, large language models (LLMs) such as ChatGPT [48] have been shown to pass such creativity tests with ease, sometimes outperforming even humans [10, 12, 20, 24, 25]. We are therefore faced with the need to further develop our tools for measuring and testing creativity. In doing so, we follow a developmental path that focuses on the (1) interpretive, hermeneutic or (2) inferential abilities of humans, as studied by hermeneutics (1) [18] and cognitive science in the area of text comprehension (2) [36]. That is, we analyze creativity as a property of a system of

cooperating agents. This is what [59, p. 231] calls group creativity, an oversummative feature that characterizes the interactions of the group without being reducible to its members individually. [43, p. 436] speaks of interactive emergence to emphasize the social network as the locus of creative processes, i.e. the group, the team or the organization is the agent of creativity, not the individual [43, p. 438]. We develop a measurement model of group creativity using the example of dialog creativity [60]. This means that we focus on the smallest possible group, the dyad, in dialogical interaction. We extend our model by a systematic comparison with systems of quasi-interacting chatbots. The aim of this comparison is to identify the characteristics of human dialogical interaction systems that are indicative of their creativity. As one theoretical background to our approach, we refer to the theory of active inference [67]. This theory posits that interactants strive to generate aligned mental representations through cooperative communication. To this end, their communication creates a “hermeneutic niche” in support of establishing a “shared narrative” that allows their (verbal) behavior to be mutually disambiguated and made expectable. In Bakhtin’s sense (see [3, p. 1343]), their dialog follows a centripetal trajectory which – as we add – strengthens their social cohesion and supports the development of a convergent or aligned situation model [54, 71]. Based on these considerations, we predict three things:

H1 Interlocutors in cooperative dialogical communication will tend to form an understanding of relevant content as early as possible. That is, content that appears later in their communication will appear to be anticipated in relation to this prior understanding. Thus, the lower the semantic anticipation of later sentences by the interlocutors, the lower the degree of creativity of the dialog, and vice versa, the higher and the earlier the anticipation, the more creative the dialog. The moment of creativity in anticipation lies in the consensus of a distributed pre-understanding that the interlocutors must agree upon by virtue of their interactions, that is, something that emerges from interactive creation.

H2 Participants tend to align their mental states so that their communication contributions become increasingly similar as the length of the interaction increases. Thus, the lower the semantic alignment between the interlocutors, the more centrifugal the dialogical trajectory in Bakhtin’s terms, the lower the degree of creativity of the dialog, and vice versa.

H3 Although H1 and H2 assume that dialogic creativity manifests moments of convergence (in terms of anticipation (H1) or alignment (H2)), we hypothesize that it also manifests a kind of divergence: Interlocutors tend to balance between highly repetitive and highly divergent contributions in order to maintain a maximum of informational uncertainty between these poles, which preserves interpretative openness for as long as possible. Thus, the more repetitive the contributions of the interlocutors, the less creative the dialog. The creativity in this kind of divergence lies in the balancing of sufficiently divergent contributions, but not so divergent as to undermine the alignment of the interlocutors and the anticipation of what is to come.

According to our creativity test, a dialogic interaction system of 𝑛interlocutors is said to be creative if it develops a common (H2)

interpretation under the condition of a linguistic input controlled by the experimenter, where this input exhibits significant coherence gaps. This common interpretation should be developed in a way that anticipates the upcoming input (H1), while maintaining sufficient interpretive openness to allow for interpretive flexibility (H3). This type of creativity differs from the one addressed by traditional creativity tests in two fundamental ways: (1) It is about social systems to which (a kind of distributed) creativity is attributed. (2) The system to which creativity is attributed does not associate a word (e.g., DAT) or formulate a phrase or sentence (e.g., AUT) in the context of a set of other words, a picture or a short description. Rather, the system generates a dialog (𝑛= 2) or multilogue (𝑛> 2), and thus a whole discourse, which serves to create a shared mental model for the aligned interpretation of the given linguistic input. In order to test our hypotheses H1–H3, we design a test scenario in which several subjects interact in the task of sequentially interpreting a sequence of sentences: first a single sentence, then a second sentence in the context of the first, and so on. The sentences are chosen so that they come from the same source text, but do not immediately follow each other in this source. Rather, there are coherence gaps between them, which requires considerable interpretation in order to understand the sequence of sentences as a coherent whole. In other words, we create gaps in coherence in order to force the interlocutors to make creative interactive inferences to fill these gaps. This is what makes creativity observable in the first place. We use computational linguistic methods to check a textual entailment relationship between these sentences. Textual entailment refers to the relationship between an entailing text 𝑇 and an entailed text 𝐻, according to which people tend to believe 𝐻to be true when they read 𝑇[13, p. 3]. By referring to entailment, we control two parameters: on the one hand, the coherence gap that exists between the sentences and makes their mutual interpretation a challenging task. On the other hand, the possibility of such an interpretation, since the entailment relationship guarantees its very existence. Note that the lower the entailment, the higher the coherence gap. Based on this conception, we speak of a Remote Entailment Analysis (REA) for assessing dialogic group creativity – see fig. 1 for a visual representation of REA. Since our creativity test scenario deals with dialogic communication, for which there are no comparable reference data, such as those provided by the remote association task [42], and since we are conducting an open-ended test [51] not in the area of monologic (cf. [25] and AUT) but dialogic communication, we need a basis for comparison. For this purpose, we use ChatGPT, with the additional hypothesis that this medium is not creative, despite its frequently mentioned creativity in the literature (see section 2). We therefore measure the creativity of human dyads compared to artificial dyads and predict that the former will be more dialogically creative than the latter. This is our forth hypothesis H4. Our test scenario borrows from objective hermeneutics (OH) [45, 46]. OH is a method of text interpretation based on the sequentiality of sentences or text passages, and is therefore particularly suitable for substantiating our approach in terms of hermeneutics. OH concerns the interpretative reconstruction of text in a sequential manner [68]. This focus stems from the underlying premise of text being the central material instance for the validation of interpretations in the social sciences [45]. The paradigm of sequentiality

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis HT ’24, September 10–13, 2024, Poznan, Poland

essential to OH is echoed, e.g., in conversation analysis: turn taking, talk-in-action [58], and situated order [19] are crucial criteria for interpreting communication and are bound to sequence and time frames. OH is based on the following principles [see 68, pp. 21-38]:

(P1) Openness: Since context is central to understanding social action, the interpretation of texts as manifestations of action sequences requires a contextualization by the interpreter that cannot be presupposed.

(P2) Sequentiality: Interpreting along the sequentiality of texts promotes openness (P1), the hypothetical continuation of the sequence, and the formulation of alternative contextualizations, where the beginning is usually open to many alternative interpretations that become increasingly contextualized and condensed as the text progresses (P4, P5).

(P3) Closeness: The interpretation of every single token, sentence, and span of a text aims to avoid imposing subjective interpretations. Starting from close interpretations, one can distinguish between what is written and latent structures as candidate interpretations.

(P4) Detailedness: The interpretation should start at a high level of detail, leaving conclusions about the overall text for later steps. Each subsequence should be interpreted for a wide range of implications, semantic nuances, and contextualizations without committing to them too early.

(P5) Parsimony: Since objective hermeneutic interpretations tend to be very detailed, with a variety of competing contextualizations, they must be close to the text in question without adding anything that is not justified by it.

OH is sequential textual analysis that attempts to be open to alternative interpretations. The goal is not to interpret what a particular author meant, but to uncover the most likely interpretation of a text in a given sociocultural setting. OH is particularly suited to the type of creativity test we are developing: The sequential presentation of sentences in a text makes hypothetical continuations and the search for alternative contextualizations after each single sentence likely (P2, P3). For each sentence, the interlocutors have about five minutes to interpret it cooperatively with all the necessary details (P4), while orienting their interpretations to the given input sentence (P5). A higher level of creative interpretation is required due to the coherence gaps that we control for with our entailment criterion (P1). This setting supports the creation of interpretation spaces that make subsequent sentences semantically predictable (anticipation – H1) due to the cooperatively generated contextualizations (P2). However, to ensure sufficient interpretive openness to account for the coherence gaps experienced, a balance (P1, P5) between textboundedness and interpretive openness is motivated (H3). To this end, we organize the interpretation process dialogically from the outset, which promotes cooperation and thus alignment (H2). To implement and test this test of dialogic creativity, the paper is structured as follows: Section 2, discusses related work on creativity and ChatGPT. Section 3 introduces our procedure for measuring group-related dialogical creativity. Section 4 describes the experiment conducted to measure and compare creativity in the context of generative AI, while Section 5 presents an evaluation of our findings along with a discussion of our results. Section 7 concludes with an outlook on future work.

2 RELATED WORK

Since the paper is about measuring creativity, it is marginally related to efforts to study the effects of authoring hypertexts or hypermedia on creative thinking, a research topic with a long tradition in the hypertext community (see, e.g., [1, 2, 5, 7, 35, 39]). The antecedents of the present paper are more in the area of research on automatic measurement of creativity based on the linguistic output of test subjects. It is also related to studies of group creativity, again based on the linguistic output of the interactants. A third point of reference concerns efforts to assess the linguistic creativity of chatbots such as ChatGPT. We will review all three of these areas of related work. The relationship between language skills (e.g., L2 [17]) and creativity is a well-studied research topic. In addition to language skills, the creative potential of everyday language activities is emphasized [41]. Carter [8] (cited after [41]) sees collaborative dialogic communication in particular as a behavioral domain characterized by linguistic creativity. That is, instead of written manifestation, dialogs become an open, dynamic object of study of creative behavior. In line with this conception, [41] emphasizes the dialogical nature of linguistic creativity as something generated by a dialogical system of interlocutors. Creativity thus becomes a property of this dialogical system. [60] emphasizes the relationship between creativity and dialog. It is precisely this understanding that we follow here. With [61] we call this a kind of distributed creativity, since the interlocutors generate the dialogical output to which we attribute creativity by their very interaction – without being attributable to a single person. In line with [32], this notion of creativity is about features of language use, i.e. discursive phenomena, rather than about an internalized (mostly lexical) language system. As much as these approaches are elaborated from a linguistic point of view, they lack a measurement-theoretical approach to linguistic creativity, especially in the area of dialogic communication, especially with regard to automatic, computer-linguistic measurement. In contrast to these approaches to distributed dialogic creativity, which challenge measurement by this very distribution, approaches to creativity in the context of the Remote Association Test (RAT) (see [69] for a review), which examine subjects’ lexical competence as a factor influencing divergent thinking as an aspect of creative thinking, have well-developed measurement procedures. In particular, there are a number of approaches based on network science [9] that quantify the creativity of individuals as a function of graph patterns of lexical networks derived, for example, using RATs. One seminal approach is [33], who predict, for example, a lower level of connectedness in the networks of low performers on creativity tests (cf. [4, 34] for similar approaches regarding, e.g., clustering and network modularity in relation to creativity). These approaches are in line with [62], who consider networks of items (here: linguistic, lexical ones) in relation to a given construct (here: creativity). However, they focus on individuals from the perspective of a language subsystem (i.e., lexis) in terms of language skills, and thus fall short of the dynamics of dialogical communication, its discourse-related dynamics, and its distributed interactive nature. This paper aims to fill these two gaps: It aims at (1) distributed creativity using the example of dialogic communication, based on (2) a test that comes with a measurement procedure for quantifying the

Experimental Scenario: Remote Entailment Analysis (REA)

Coherence Gap 3

Coherence Gap 1 Coherence Gap 2

t0(= 0)

t1(= 5)

entails

Sentence (1)

Sentence (2)

entailment rate x

entails — entailment rate z

duration: 3 × t1

Interlocutor A

interaction system AB (here dialogic), i.e., the system to which distributed creativity is attributed

concept of creativity developed in the introduction. As a benchmark, the paper refers to ChatGPT. Consequently, we apply the same creativity test to pairs of human interlocutors on one side and artificial actors on the other side to test our hypotheses. This brings us into the context of creativity research using the example of LLMs and ChatGPT. [28] study prompt engineering – with prosaic, literary-looking text samples – with ChatGPT in supporting users’ hermeneutic perspectives on texts and their critical reflection: ChatGPT’s output is said to be hermeneutic if it is meaningful to the human reader. Each combination of human and artificial writers and readers is considered qualitatively, without any (quantitative) measurement. What is interesting for us is the assessment of the AI writer/reader variant, according to which the AI is able to write hermeneutically valuable texts starting from non-contextualized sentences. This is the property that we study, but in a dialogic variant of 15 minutes in length (see fig. 1), which we evaluate quantitatively and compare with human outputs under the same input conditions. One might ask whether it makes sense to include ChatGPT in a creativity test at all, since it produces output that is judged by human readers to be less creative than human-generated text. Contrary to such a view of algorithm aversion (cf. [44]), [11] show that ChatGPT generates a kind of poetry that is acceptable to (Chinese) readers and thus meets human creativity. Therefore, it may be a

t2(= 10)

entails

Sentence (3)

entailment rate y

Interlocutor B

difficult task to show that ChatGPT does not allow for the kind of dialogic creativity we are investigating here. ChatGPT as an assistant in the context of creative writing [15, 65] is also a well-researched topic. It shows, e.g., that it has problems adapting to the styles of competent writers because it prefers a certain style of writing [30], or that it is outperformed (though not significantly) by humans in narrative originality [15], emphasizing its role as an assistant rather than a substitute for human creativity [37]. In any case, these studies show the creative potential of ChatGPT in terms of generating texts to which readers can ascribe “hermeneuticity” [28], that is, as something that can be interpreted in a meaningful way; this should also challenge the comparison we are aiming at. Creative storytelling with ChatGPT is exemplified by [56], who demonstrates ChatGPT’s ability to collaborate with a human author. We consider the scenario where this collaboration relies on two instances of collaborating AI bots. Based on the Divergent Association Task (DAT), [12] explores the question of whether GPT (versions 3.5/4) outperforms humans, as statistical evidence would suggest. However, a closer analysis of the quantitative data suggests that a high-performing human is likely to produce more consistent association results. Nevertheless, GPT-4 is already a serious competitor here, even if not yet in a more creative way. In [12], the DAT does not consider temperature as

Figure 1: A Remote Entailment Analysis (REA) is performed by a dialogical interaction system AB consisting of two interlocutors to whom three sentences are presented one after the other (starting at 𝑡0). The task of AB is to interpret this sequence of sentences (presented at 𝑡0, 𝑡1 and 𝑡2) in a way that is consistent with the principles of objective hermeneutics (see text). The sentences are sampled from the same source text to ensure coherence, but in such a way that their entailment rates 𝑥, 𝑦and 𝑧are controlled to guarantee a certain coherence gap that makes this interpretation challenging. The gaps are calculated as functions of the respective rates or degrees of entailment (for the details see section 3). That is, Gap 1 is a function of entailment rate 𝑥, etc. 𝑥,𝑦,𝑧are chosen so that Gap 3 < Gap 1 + Gap 2 (note that these are discourse semantic gaps, not to be confused with temporal gaps 𝑡𝑖+1 −𝑡𝑖). The reason for this is to ensure that sentence 3 is located in a semantic sphere around sentence 1, which allows its semantic anticipation (see H1), and not in a sphere determined by sentence 2, which takes it far away from sentence

    The task of the interlocutors is then to develop a common interpretation sentence by sentence (P2). Each sentence presented

challenges the interpretation and its contextualization developed so far and may require adjustments, which in turn require a reopening (P1) of the interpretation space. Given this scenario, a REA attributes creativity to a system of two interlocutors.

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis HT ’24, September 10–13, 2024, Poznan, Poland

a parameter, although it is associated with creativity.1 It focuses on a more or less closed group of nouns, rather than considering the openness of dialogic communication. A better performance of GPT-4 in relation to humans is observed by [10], who experiment with temperature in the interval [0.1, 1.0] and observe for GPT-4 slightly decreasing DAT values with increasing temperature. [25] perform the AUT and are therefore closer to the kind of open-ended test performed here, albeit in terms of rather phrasal responses. They find that generative AI-based chatbots like ChatGPT perform the AUT in a way that they can be judged as creative (see [64] for related findings). Interestingly, they also experiment with an AI for evaluating AUT prompts [51], which cannot be applied here since we are considering dialogic communication and not just single responses at the level of phrases or sentences. In a review of such research, [20] argues that LLMs can match or even surpass humans, but are not creative because they only simulate outcomes, but do not emulate creative behavior, as we put it using the terminology of [52] (see [16, 37] for similar arguments). Analogously, [24] shows that GPT can master the Torrance Tests of Creative Thinking (TTCT) [66] in such a way that it challenges the expressiveness of current creativity tests, which do not seem to sufficiently reflect the nuances of human creativity (see [70] for related findings). Moreover, [24] show the need for a more precise distinction between human and artificial creativity, a task to which this paper is devoted, using dialogic group creativity as an example. All of these tests have in common that they deal with individual test takers and their responses, which are mostly lexical or consist of short responses (phrases or sentences) and, more rarely, shorter texts. This is very different from our approach to testing creativity, where there are no constraints on the communication of interlocutors who cooperate in organizing their dialog, which is constrained only by the sequential provision of the initial task, the three test sentences and the maximum length of about 15 minutes for the whole conversation (see fig. 1).

3 TESTING GROUP CREATIVITY: APPROACH

The procedure of our REA is illustrated in fig. 2 and generally described in this section, our implementation is then presented in section 4 in detail. The first step is the selection of 𝑛sentences in such a way that they are jointly similar and at the same time show a negative entailment, as described in fig. 1. We then perform REA in which the selected sentences are presented to the participants – both, groups of ≥2 human and/or chatbots – for interpretation. The instructions given to the human participants and the prompts used (“prompt engineering”2) need to be evaluated and, if needed, adjusted. The next step is the preprocessing of the experiment results, i.e. bringing the discussions into textual form to enable NLP-based analyze, e.g. by (automatically) transcribing audio or video recordings of the human participants and performing speaker diarization. Additionally, multimodal information could be introduced to the analysis. Validation of the experiment data follows, the textual data needs to

1For our creativity-related comparison, ChatGPT’s temperature parameter is of interest: values higher than 0.8 introduce randomness into word prediction, those lower than 0.2 make the prediction deterministic [50]. Temperature is relevant for the analysis of artificial creativity [14], underlining the need to take it into account, as done here. 2The prompt defines the problem to be solved creatively, as stated by [25, p. 2].

be evaluated to detect and minimize bias and errors introduced by the processing tools (transcription, diarization) and human influence, e.g. moderator bias, multiple participations and more. Ideally, the chatbot-generated text should be processed in the same way to minimize differences and align errors and bias by the processing tools, e.g. by utilizing speech synthesis to subsequent transcription as with the human answers. Finally the textual data can be analyzed in regards to the four hypotheses described above.

4 EXPERIMENTS

In our experiments we compare the results of tasking human participants (section 4.1) and ChatGPT (section 4.2) with discussing and interpreting a short text. We carried out 34 REAs as described in section 3 with pairs of human interlocutors. Each pair was provided with a short text sample, consisting of three sentences, to discuss and interpret. Each text sample was only used once, i.e. in one human discussion and one ChatGPT-based session. As mentioned, these sentences were chosen from the same source text to enable a creative interpretation: We generated groups of three sentences based on James Joyce’s Ulysses by randomly selecting sentences that are jointly similar and at the same time show a negative entailment. First, we selected subsequent 5000 sentences from the beginning of the book, then generated sentence embeddings [57] with the paraphrase-multilingual-MiniLM-L12-v2 model for each sentence and calculated the cosine similarity for each sentence pair. Subsequently, we calculated the entailment rates for all possible sentence combinations (50002 in total) using a DeBERTa-V3 [27] based model3 trained on [38]. To achieve both similarity and ambiguity within each sentence pair, we used a modified geometric mean of cosine similarities and entailment rates for ranking sentences. This modification was necessary because both cosine similarity and entailment rates are defined in the range -1 to 1. In the analysis, for each pair of juxtaposed sentences the computational model produced a triplet of probabilistic values corresponding to three different relations: "entailment", "neutral" and "contradiction". Each score lies within the interval [0, 1]. In order to derive a single entailment score from this tripartite output, several methodological approaches were considered. These include: (i) directly using the entailment score as a stand-alone measure, thereby restricting the range of the score to [0, 1]; (ii) adjusting the entailment score by subtracting the contradiction score, which re-calibrates the resulting values towards a median-centric distribution; and (iii) the preferred method adopted in this study, whereby the entailment score is used when it exceeds the contradiction score, otherwise the contradiction score is used with an inverted sign. This methodology ensures a coherent analytical framework for quantifying the extent of logical entailment between pairs of sentences, anchored in the probabilistic outputs generated by the model. This approach allowed us to effectively combine cosine similarities and entailment rates into a single ranking metric that preserved both similarity and ambiguity within each sentence pair. The first step in our approach was to calculate the absolute values of both the cosine similarity and entailment scores to ensure that all values

3https://huggingface.co/MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli- 2mil7, this is a multilingual model trained on hypothesis-premise pairs.

✓Validation

© Human Experiments

Ð Transcription

. Sentence Selection

✓Validation

á ChatGPT Experiments

Ó Prompt Engineering õ Queries

Figure 2: Complete overview of the experiment process.

Figure 3: Entailment rates of sampled sentences (x: entailment rate between sentence 1 and 2, y: entailment rate between sentence 2 and 3, z: entailment rate between sentence 1 and 3).

Figure 4: Cosine similarities of sampled sentences (x: entailment rate between sentence 1 and 2, y: entailment rate between sentence 2 and 3, z: entailment rate between sentence 1 and 3).

were non-negative. We then calculated the geometric mean of these absolute values. If either the cosine similarity or the entailment rate was originally negative, we multiplied the resulting geometric mean by -1. This method allowed for a nuanced assessment of the pairs of sentences and ensured that each pair of sentences was characterized by a balance between similarity and intended ambiguity. Finally, we randomly selected the first sentence of the group and paired it with the sentence that had the lowest ambiguity similarity score (ASS). The third sentence was chosen in the same way. We follow the principle of a hermeneutic text interpretation session, i.e. we consider the text sentence by sentence.

4.1 Human Participants

We performed 34 discussions with groups of two persons, in total 64 human interlocutors participated in our experiments. Of these, 5 persons participated in two experiments (each time with a partner taking the experiment for the first time). The participants were

✓Hypothesis Tests

✓Validation

Ü Speaker Diarization

ç Output Texts ✓Validation

✓Bias Validation

Figure 5: ASS values of sampled sentences (x: entailment rate between sentence 1 and 2, y: entailment rate between sentence 2 and 3, z: entailment rate between sentence 1 and 3).

Figure 6: ASS values of 1000 sentences sampled using the modified geometric mean method compared to sentences sampled from discrete uniform distribution.

divided into two groups: (1) undergraduate computer science students and (2) a mixture of postgraduate students, doctoral students and postdoctoral fellows. There were 16 experiments filled with students (1) and the remaining 18 with participants of segment (2). The experiments were accompanied by one of three moderators (responsible for 14, 12 and 8 experiments respectively), who showed the sentences and explained the task. In each session, the participants were presented with three sentences and asked to freely discuss and interpret them. As suggestion, the following two requests were given: “What could the text that begins with these sentences describe?” and “How does this sentence relate to the previous one?” The sentences were given one by one, previous sentences were visible at all time. For each sentence, the participants were allowed to discuss for up to five minutes; they were free to advance to the next sentence at any time. On average, a session lasted for 15 minutes

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis HT ’24, September 10–13, 2024, Poznan, Poland

Figure 7: Comparison of sentence embedding similarities of input and output sentences transcribed using different Whisper model sizes for experiment exp01. (R)soag: (relative) sum of absolute gradients, nozg: number of zero gradients, soH sum of entropies. Vertical lines mark the presentation of each new sentence.

Figure 8: Comparison of sentence embedding similarities of input and output sentences. From left to right: human transcripts, ChatGPT output, auto-correlation coefficients of the similarity scores for human output and for ChatGPT output.

and 29 seconds (with a standard deviation of 1:23, a minimum of 12:13 and a maximum of 17:56 minutes) and about 185 sentences were spoken. All experiments (including the target sentences, the instructions and the discussions) were conducted in German and are translated here for easier understanding. The discussions were recorded and subsequently transcribed using the Whisper [55] model. In most cases we only recorded the sentences in video without the participants being visible. We thus do not evaluate or analyze the interlocutors on a multimodal layer but rely on the text transcript of their discussion generated from the audio recording. We found that using one of the default Whisper models (e.g. small or large), the transcript quality differed throughout the experiments in a non-systematic unpredictable way. Because we were working with a much smaller dataset, where essentially every word of the discussion counts, we had to be particularly careful in evaluating and validating the transcript results. For this we developed an automatic transcript selection algorithm that uses features of the series of similarity values extracted from each transcript. The features that showed the best results were the sum of the entropies of the absolute similarity values (soH), the number of points with gradients equal to zero (nozg), the sum of the absolute gradients (soag), and the relative sum of the absolute gradients (rsoag). The features are aimed at detecting repetitiveness and corruption of text structure. In a few cases, however, manual correction was necessary, which was easy and quick due to the way the transcripts were visualised. To this end, we selected the most suitable Whisper model size for processing the transcript of an experiment. For further analysis of the participants we performed speaker diarization. For this we used the NVIDIA NeMo framework [26] as well as Whisper Timestamped [21, 40], an extension of OpenAI’s Whisper [55] allowing for more accurate, per word, timestamps. More specifically we used the NeMo framework to perform speaker diarization through the use of Automatic Speaker Recognition in combination with Voice Activity Detection. For this we used the models Frame-VAD Multilingual MarbleNet [31] and STT DE Conformer-CTC Large [23] respectively, both of which were created by NVIDIA. We then used Whisper Timestamped to derive a transcript of the conversation with each word being associated with a certain time range which we could then merge with the other dataset resulting in a transcript of not just the sentences but

also the associated speakers. Finally we manually added additional timestamps denoting a new sentence being discussed.

4.2 ChatGPT

For each of the 34 experiments conducted with human participants we simulate the REA using 6 variants of ChatGPT. Using the same short texts of three sentences each, we prompt ChatGPT to produce a discussion and interpretation similar to the human interlocutors. To stay comparable to the human experiments, all ChatGPT simulations were queried and further processed in German, we did however not explicitly instruct ChatGPT to use a specific language. We simulate two participants by using two ChatGPT agents taking turns, where the order of speaking is randomized and both are prompted from their own point of view, i.e. without shared context. We consider 4 different prompt variants, that are following the same structure consisting of a system prompt as first prompt, followed by a user prompt for each turn. The user prompt contains a summary of the discussion, the answers provided by ChatGPT – in which we encode which ChatGPT agent is speaking – and the currently active sentence. For each of the three input sentences of the text, each ChatGPT agent is queried for 5 answers, one experiment thus amounts to 30 requests to ChatGPT, in addition to 3 summary requests. We did not limit the length of the answer ChatGPT should provide, analogous to us not interfering directly during the human experiments. After 10 answers, we generate a summary of the discussion so far, again using ChatGPT. This is necessary for technical reasons, as the input and output size of ChatGPT is strictly limited to 128000 (GPT-4) and 16385 (GPT-3.5) subwords. We experiment with four different prompts consisting of two prompts that were additionally automatically optimized using PromptPerfect4. The first prompt (denoted as 𝑆𝑖𝑚𝑝𝑙𝑒and 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡) is a basic prompt that corresponds to the task that was given to the human participants. In this way we instruct ChatGPT by using a system prompt containing these information and a call for interpretation and a user prompt, which contains a summary of the discussion, the last answers provided by the chatbots and 1 to 3 sentences, depending on the progression. These prompts were used with two ChatGPT models, GPT-3.5 (gpt-3.5-turbo-0125)

4https://promptperfect.jina.ai/home

and GPT-4 (gpt-4-0125-preview). In addition, we used GPT-3.5 with a prompt (𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐and 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡) tasking Chat- GPT with performing the discussion and interpretation following the hermeneutics principles. In this prompt, we instruct ChatGPT to take on the role of an objective hermeneut discussing the sentences with a colleague. We experiment with employing student personas with these two hermeneutic prompts, i.e. a description of behavior, to both ChatGPT agents to further encourage a discourselike scenario on a level of a beginner to the hermeneutic process. The temperature parameter can be used to let ChatGPT choose more random or deterministic outputs and is often referred to as “creativity” [49]. We consider 13 values of temperature between 0 (deterministic), 1 (default value) and 2 (most random) in our experiments to present its effect in the context of our REA.

5 EVALUATION AND ANALYSIS

We developed a computational validation framework in order to ground our experiments and results. We consider different classification experiments to provide validation to the experiments performed and the data gathered and preprocessed, as well as to the analyses that follow and to test H4. The classifications operate on the principle of using machine learning to train a classification algorithm that is able to successfully separate data points towards a target class, i.e. we are able to attribute a sentence produced during an experiment with a specific parameter setting. We base the training of such a classifier on three different feature values and utilize two basic and proven algorithms, (1) Support Vector Machine (SVM) and (2) Random Forest Classification (RFC). We utilize the implementation provided by scikit-learn [53] using default parameters. Each discussion text in an experiment was preprocessed using spaCy [29] (with models decorenewssm and decorenewslg) and SentenceTransformers [57] (paraphrasemultilingual-MiniLM-L12-v2). We consider each sentence as a training sample and use the sentence embeddings, i.e. the sentence vector produced by spaCy, and the embedding encoded by SentenceTransformers, as feature vectors for this sample. As a baseline test we tested two additional features as separate classification experiments, where we use the number of tokens (i.e. the sentence length) and the number of characters in a sentence as feature values. We found that the sentence structures of the discussion of Chat- GPT and the human participants showed unalike behaviour based on length of the sentences. In this way, we show and validate the results of our experiments in a fully automatic process, on which the analyses concerning our hypotheses build on. Each pair of human participants was supported by a moderator during the experiment. The possibly different approaches to handling the instructions and presentation of the sentences might induce a bias into the discussion. As each tool generates slightly different sentence boundaries, this amounts to, on average, 5147 samples for the training set, and 1288 samples for the test set (20 %). The three moderators were used as target classes, to train 18 threeclass classifier to separate the sentences generated by the human participants towards the moderator who guided the experiment. We report the macro 𝐹1 score in fig. 9, with an average of 0.326 (with 10-fold cross validation the average is 0.324). The best result, based on sentence embeddings by SentenceTransformers, reaches

a score of 0.5 and the spaCy-based classifier an average of 0.415. These results show the vector-based classifiers to be better (with a p-value near 0) than a random classifier while not being able to distinguish between the moderators, and the word and character count based classifiers to be worse than random (p-value of 0). This shows we are not able to separate the moderators based on the discussion text and suggests little bias introduced by them in the experiments. As mentioned in the experiment description, five persons participated in two experiments, resulting in them being aware of the experiment structure and knowing about the selection of the three sentences to discuss. Using the same method we employed for examining the moderator bias, we validated the results regarding these participants. We trained binary classifiers with the target variable encoded as whether a participant in an experiment participated a second time. The results, depicted in fig. 10a, show that we can not distinguish sentences spoken in experiments where a person participated for the second time, with most classifier being significantly worse than a random classification at an average of 0.471 𝐹1 score overall – without much difference concerning the different feature generation methods and a score of 0.469 in a 10-fold cross validation. Considering this outcome, repeated participation seems to be possible without affecting the creativity of the interpretation output if presented with an unseen text. How much influence has the educational level on the REA? We separated the participants in two groups and performed 16 experiments with undergraduate students and 18 with the remaining participants. Following our validation framework, we evaluate the results in the same way. We show, see fig. 10b, that the results of the classifiers predicting the group membership indicate similar behavior in this regard. With a max 𝐹1 score of 0.658 (Sentence- Transformers embedding) and average of 0.532 (0.540 using 10-fold cross validation) the groups, and thus the educational level, are mostly indistinguishable. This indicates some sort a disconnection between the approach and discussion style and also the knowledge background and experience, unlike in a hermeneutic process. In each experiment, the interlocutors discussed a different set of three sentences. This leads to the expectation, that the experiments – the dialog generated by the pairs of participants – should be easily distinguishable. We show in fig. 10c, using the same technique as before, i.e. classifying the concrete experiment based on embeddings of all sentences from these experiments, that this is the case. We reach a maximum 𝐹1 score of 0.326 with an average below 0.08 (0.078 using 10-fold cross validation), with the classifiers trained on sentence embeddings being significantly better than random classification. Considering that we predict a large number of 34 target classes and the amount of training samples per class, i.e. per experiment, is quite low, the result suggests a certain distinguishability. In this way, the results show the different aspects the three feature variants capture: while the experiments are similar in structure, they cover different topics. Considering the results of the four analyses based on simple classification methods, we presented some validation of our proposed REA and its realization for the human participants. We now introduce the results of the experiments generated by ChatGPT.

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis HT ’24, September 10–13, 2024, Poznan, Poland

Figure 9: Macro 𝐹1 scores of separability of moderators in human experiments.

(a) Participants who took part twice. (b) Groups based on their knowledge level. (c) Pairs of human participants.

Figure 10: Classification results showing macro 𝐹1 scores and target classes.

The ChatGPT-based dialogs have been processed in the same way as the transcripts of the human participants, we computed sentence embeddings by means of SentenceTransformer and spaCy. Note that we did not have to transcribe the chatbot answers, preventing the introduction of conversion noise and bias in comparison to the dialogs generated by the human participants. We consider two ChatGPT models (GPT-3.5 and GPT-4) that were queried via the OpenAI API5, as well as up to four prompts as described in section 4.2. On average, the participants produced 185 sentences during their discussion, with a standard deviation of 62 sentences. The number of sentences, as well as the words in a sentence, by ChatGPT is much higher: ChatGPT always returned an answer and was never silent in contrast to the human participants – however generated invalid responses containing just random words – which resulted in 755 sentences on average (with a standard deviation of 1381) over all experiments, models, prompts and temperatures. Figure 11 shows an overview broken down by the parameters: The large average count is an artifact produced by the problematic temperature values, the average amount of sentence at the default temperature 1 is 339 (std 196), which is closer to but still much higher than that of the human participants. We now perform the same classification to separate the experiments as tested with human participants. In contrast to their dialogs, we reach higher scores with an average of 0.423 using Sentence- Transformer embeddings on all temperatures with GPT-4, with temperatures 0, 1 and 1.3 showing 𝐹1 scores of 0.957, 0.682 and 0.207, respectively. In all cases, the 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡performed worse.

5https://platform.openai.com/docs/api-reference

Figure 11: Overview of the amount of sentences produced by different ChatGPT agents.

This holds true for GPT-3, where in comparison the results are higher over the different prompts, with an average of 0.974, 0.920 and 0.740 for the three temperatures and 0.850 for all temperatures. Considering the number of classes (i.e. 34), one for each experiment, this shows an extreme separability of ChatGPT answers. It shows a stark difference in how the temperature manifests itself in the ChatGPT versions, and also the effect of optimizing the prompts. For each experiment, we queried ChatGPT 30 times to progress the interpretation dialog, based on 5 answers per 3 input sentence for 2 interlocutors. We evaluate the different temperatures supported by ChatGPT, which is somewhat advertised as the models creativity [49], i.e. from deterministic answers on one hand to more open, “creative” replies on the other. Figure 12 shows an overview of how many of these requests were successful for each experiment by using the ChatGPT versions 3.5 and 4 with 𝑆𝑖𝑚𝑝𝑙𝑒. The color indicates the amount of requests, up to a total of 30, and the size

Figure 12: Overview of the amount of answers and repetition produced by ChatGPT (top GPT-3.5, bottom GPT-4) using different temperature parameters for prompt 𝑆𝑖𝑚𝑝𝑙𝑒.

depicts the Jaccard similarity of the sets of five-character shingles of the pairs of neighboring sentences cumulated over the entire text, which can be used to asses the level of repetition (see fig. 13 for a more detailed example of the repetition structure of the texts produced in two experiments. The human output is shown on the left, followed by 4 prompt variants processed by GPT-3.5 and two prompt variants processed by GPT-4.). We noticed that the chatbots results from temperature 1.4 start to produce much less queries than expected, corresponding directly to the temperature rise. In some instances, the first query already fails. This is the result of the OpenAI API blocking the requests due to problematic or unsafe characters in the answer generated by ChatGPT. The higher temperature parameter prompts ChatGPT to produce more random output, which manifests itself in incoherent and incomprehensible words, sometimes in other languages and fragments as well as completely random characters. In the context of our REA, the temperature is

Figure 13: Jaccard similarities between shingle sets of consecutive sentences. Colors map dialogue phases: blue denotes the first, red the second, and green the third sentence.

ChatGPT Prompt Feature Mean F1

GPT-3.5 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐 character count 0.610 GPT-3.5 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐 sentence length 0.597 GPT-3.5 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐 sentence embedding 0.947

GPT-3.5 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡 character count 0.490 GPT-3.5 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡 sentence length 0.483 GPT-3.5 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡 sentence embedding 0.958

GPT-3.5 𝑆𝑖𝑚𝑝𝑙𝑒 character count 0.626 GPT-3.5 𝑆𝑖𝑚𝑝𝑙𝑒 sentence length 0.600 GPT-3.5 𝑆𝑖𝑚𝑝𝑙𝑒 sentence embedding 0.974

GPT-3.5 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡 character count 0.599 GPT-3.5 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡 sentence length 0.581 GPT-3.5 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡 sentence embedding 0.967

GPT-4 𝑆𝑖𝑚𝑝𝑙𝑒 character count 0.491 GPT-4 𝑆𝑖𝑚𝑝𝑙𝑒 sentence length 0.491 GPT-4 𝑆𝑖𝑚𝑝𝑙𝑒 sentence embedding 0.939

GPT-4 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡 character count 0.495 GPT-4 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡 sentence length 0.495 GPT-4 𝑆𝑖𝑚𝑝𝑙𝑒𝑜𝑝𝑡 sentence embedding 0.935

Table 1: Classifying humans vs. ChatGPT, averaged over experiments and temperatures.

thus not a suitable parameter to control or force a more creative output in dialogs of ChatGPT. A chatbot generating answers indistinguishable from humans might be considered as being creative. However, the answers of ChatGPT suggest a straightforward distinctiveness between sentences in dialogs generated by ChatGPT compared to human participants, as noted in H4. We train binary classifiers using SVM and RFC on the sentence embeddings generated by SentenceTransformers and spaCy for all ChatGPT models, prompts and temperatures tested, for all experiments. The classifiers are trained to separate human- and ChatGPT-generated sentences. This amounts to 36 results with an average 𝐹1 score of 0.953. Using the token or character count as features for training drops the score to 0.541 and 0.552, respectively. See table 1 for a breakdown based on the ChatGPT version, prompt and trainings feature used, this shows comparable results. Figure 14 shows a comparison of the 𝐹1 scores using the three training features from all experiments for the two ChatGPT models for each temperature, averaged over the different prompts and spaCy and SentenceTransformer tools. The sentence embedding features show a constant high score value with little deviation for all tools, independent of the tool that embedded the sentences, i.e. spaCy or SentenceTransformers. In contrast, the simple count of words or characters show solid performance only up to a temperature parameter of about 1.3. This shows, that we are able to

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis HT ’24, September 10–13, 2024, Poznan, Poland

Figure 14: Overview of the human vs ChatGPT classification results, for all experiments.

Figure 15: ChatGPT model, prompt and model-prompt classification using SentenceTransformer embeddings, showing the average macro 𝐹1 score over all experiments.

distinguish human and ChatGPT produced sentences quite easy, even with simple counting features, as the answering style of human participants and ChatGPT contrasts starkly. However, as we reach higher temperatures, the replies by ChatGPT become more and more erratic, resulting in shorter sentences making the training of the classifiers much harder, as also the human participants used shorter sentences. This holds true for all experiments, either combined or separate, via the results of a t-SNE dimension reduction and k-means clustering using scikit-learn to separate the two groups. One can easily make out two clusters in results generated by lower temperatures, and the amount of sentences, depicted by the increase in data points. Notably, there are two ChatGPT-clusters starting from temperatures about 1.3, indicating a separability between what might be two ChatGPT-based agents or between incomprehensible and proper sentences. The results of the k-means clustering further support the classification experiments. Again, the clusters are comparable between the different experiments. This ties into our hypothesis H4, we argue that if ChatGPT would be able to be creative on a level compared to humans we would not be able to differentiate between these two groups. Finally, we use our classification framework to evaluate the different ChatGPT versions and prompts. Figure 15 shows the results of using SentenceTransformer embeddings for each experiment individually, as well as for all experiments combined, across temperatures 0, 1, 1.3, and all 13 temperature values collectively. Shown is the average macro 𝐹1 score for the three classification targets where we try to detect the ChatGPT model (binary, GPT-3.5 vs GPT-4), the prompts used independently of the ChatGPT model (4 classes), as well as the combination (6 classes). In all cases, we report

scores that indicate a distinguishability of the replies given by the chatbots, with lower scores on higher temperatures, presumably due to the more random outputs. This again shows the need for validating and evaluating the ChatGPT prompts, as these are the basis to producing data that can be analyzed – probably even more important than the instructions given to the human participants. We applied two methods to test H1. Initially, for both methods, we measured the cosine similarity values of the sentence embeddings of the sentences presented to the participants and those generated by them. Then, for the first method, we calculated the average of these values over the duration of each phase, excluding similarity values related to sentences already presented to the participants. In phase 2, it was the interval between the red and green lines, averaging the values of the green curve. Our expectation is that human participants will have higher average scores than the AI. The one-sided t-test showed that the mean values for the human participants were indeed higher (statistically significant on 1% level) for 2 prompts out of 6 (ChatGPT 3.5 temperature 1; 𝑆𝑖𝑚𝑝𝑙𝑒and 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐). The second way to measure anticipation is to take the differences between the similarity of the seen sentences, i.e. sentence 1 in phase 1, sentences 1 and 2 in phase 2, and the values for the unseen sentences, i.e. sentences 2 and 3 in phase 1 and sentence 3 in phase 2. Here we would like the differences to be smaller for the human participants, as the AI agent would only focus on the current sentence, resulting in higher differences. The one-sided t-test affirmed the conjecture in 3 cases out of 6 (ChatGPT 3.5 temperature 1; 𝑆𝑖𝑚𝑝𝑙𝑒, 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐, and 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡). To test H2, we look at the similarities of the embeddings of the last sentence of one turn and the first sentence of the next turn of the other interlocutor separately for phase 1 (initiated by the 1st pivot sentence S1) and 3 (initiated by the 3rd pivot sentence S3), for all turn changes. The one-sided t-test showed that the mean values for the human participants were smaller (statistically significant on 1% level) for 1 prompt out of 6 (ChatGPT 3.5 temperature 1; 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡), if we take only last phase into account. When we only test for alignment in the first phase, all 6 tests are in support of H1. Finally, we propose three criteria to measure balancing (H3). In the endpoint-only variant we consider only the similarities of the last generated sentence in the last phase, none of the tests attest any statistically significant differences. For last-phase-all-values, where we consider all similarity values in the last phase, 5 out of 6 prompt variants show statistically significantly lower values for human participants, only 𝑆𝑖𝑚𝑝𝑙𝑒being an exception. The last method to measure balance is to test the differences in the slopes of the linear

fits of the curves produced by the human participants and AI agents. Here the t-test showed that the slopes for human participants were lower (statistically significant on 1% level) in two cases, 𝑆𝑖𝑚𝑝𝑙𝑒and 𝐻𝑒𝑟𝑚𝑒𝑛𝑒𝑢𝑡𝑖𝑐𝑜𝑝𝑡.

6 DISCUSSION

Our analysis has shown that – somehow in line with previous literature, but with a focus on dialogic communication – ChatGPT-based artificial dialogic communication cannot easily be classified as less creative than human dialogic communication. Our analysis results are quite mixed in this respect: Some of the anticipation experiments (H1) suggest that humans indeed verbalize interpretations that anticipate unknown future text to be interpreted. However, this result is not consistently unambiguous and therefore does not draw a strict line between humans and machines. Nevertheless, we find here an indication of a type of text comprehension that is alien to statistical language models per se: humans integrate what is to be understood into a comprehension horizon, they form mental models into which what is to come should be able to fit as coherently as possible so that the comprehension process converges. If this process is cooperative, we expect communicative alignment of the interlocutors (H2). With our measurement instruments we observe this alignment in significant difference to the machine in the first phase, but not in the third phase. Thus, either our measurement instruments are inadequate, or interlocutors develop a kind of more open horizon of understanding which, in the course of their communication, puts them in a state in which they differentiate from each other linguistically in a quasi-consensual way, since their alignment takes place at a deeper, semantic level which the linguistic surface – in contrast to the more repetitive language models in this respect – does not reflect. To study this, further investigations and further development of our NLP-based measurement procedure are necessary. When it comes to the degree of informational openness in the sense of a balance between repetitive and far-reaching, divergent interpretations, we again do not observe sufficient distinguishability between humans and machines. Here, too, either the language model we used (embeddings) is too imprecise to capture such semantic nuances and relationships, or the human pairs as a group are too divergent among themselves for H3 not to be falsified. What we do observe, however, is that human interaction systems are separated from their artificial competitor systems when the latter are based on GPT-3.5, or when the temperature parameter is changed beyond a critical level that actually destroys the artificial system’s ability to generate reasonable output. Here the picture is almost as predicted by our hypotheses H1-H3. Last but not least, our extensive classification experiments to distinguish between human and artificial pairs actually point to their differences, even though this classification is based on classical approaches using vector representations to embed human and artificial linguistic data. This can be seen as an indication of the fundamental distinctiveness of the two domains (human vs. machine dialogues) and as a challenge to further develop NLP tools in order to work out these differences in terms of a model that allows a hermeneutic and at the same time a social-cognitve interpretation, as addressed by hypotheses H1-H3.

From a technological point of view, we have demonstrated the need for validation at key points of data generation in our experiments. In the case of transcribing spoken text from audio recordings using the Whisper model, our evaluation shows that one cannot rely on high scores of model training results. Rather, we had to select different models for different dialogues in order to obtain optimal results, for which we developed a novel heuristic method. On the other hand, ChatGPT’s output, while readily available as text, is heavily influenced by prompt phrasing and parameters such as temperature, requiring extensive prompt engineering to control these dynamics. The same is true for speaker identification, which is an integral piece of information needed to determine whether an alignment process is taking place during the dialog between different interlocutors. So a fully automated process is possible, but currently not as accurate as desired. From this perspective, we advocate intensive research into NLP methods for the analysis of spoken data on the one hand, and the controlled generation of AI data for creativity research on the other. This will be in support for the kind of group-based creativity test we are pursuing.

7 CONCLUSION

In this paper, we have developed a shift (1) from traditional creativity tests that focus on individuals to a group-based test, (2) from tests that focus on constrained and reduced language output (mostly at the word or phrase level) to tests that address free dialogic communication. To this end, we developed a test scenario in terms of a remote entailment analysis that borrows from objective hermeneutics. It is based on the idea of a sequential interpretation of a sequence of sentences, which poses the task of filling coherence gaps by means of a cooperative interpretation, where these gaps are algorithmically controlled by relying on the concept of entailment. We measured entailment using large language models and addressed the task of automatically transcribing dialog data using Whisper, heuristically bypassing erroneous transcription output. We have explored several hypotheses about human dialogic creativity as opposed to artificial creativity, with results that are as promising as those that suggest a necessary deepening of our approach: although we find some differences in terms of anticipation (H1) and early alignment (H2), later alignment (H2) and balancing effects (H3) favor the machine. We find clear differences that distinguish human communication from that based on GPT-3.5, and also show that temperature can have a dramatic effect, letting the machine fall behind the human competitors. Taken together, these results show that we have opened up a research direction that allows for measuring creativity closer to the everyday creativity of dialogic communication. However, this also requires a further development of the NLP-based measurement instruments, for which we have provided a research framework with the our Remote Entailment Analysis (REA).

Acknowledgments

This work was supported by the research group CORE (FOR 5404, project number 462702138), sub-projects B05 and C08, funded by the German Research Foundation (DFG).

Measuring Group Creativity of Dialogic Interaction Systems by Means of Remote Entailment Analysis HT ’24, September 10–13, 2024, Poznan, Poland

References

[1] Oscar Ardaiz, Maria Luisa Sanz de Acedo, and Maria Teresa Sanz de Acedo.

    Wikideas and creativity connector: supporting group ideational creativity.

In Proceedings of the 4th International Symposium on Wikis (Porto, Portugal) (WikiSym ’08). Association for Computing Machinery, New York, NY, USA, Article 31, 2 pages. https://doi.org/10.1145/1822258.1822299

[2] Claus Atzenbeck, Dene Grigar, and Manolis Tzagarakis. 2023. Interdisciplinary Teaching Toward the Next Generation Hypertext Researchers. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT ’23). Association for Computing Machinery, New York, NY, USA, Article 43, 5 pages. https://doi.org/10.1145/3603163.3609055 — ACM HyperText copy

[3] Nic Beech, Robert MacIntosh, and Donald MacLean. 2010. Dialogues between academics and practitioners: The role of generative dialogic encounters. Organization Studies 31, 9-10 (2010), 1341–1367.

[4] Mathias Benedek, Yoed N. Kenett, Konstantin Umdasch, David Anaki, Miriam Faust, and Aljoscha C Neubauer. 2017. How semantic memory structure and intelligence contribute to creative thought: a network science approach. Thinking & Reasoning 23, 2 (2017), 158–183.

[5] Michael Mose Biskjaer, Jonas Frich, Lindsay MacDonald Vermeulen, Christian Remy, and Peter Dalsgaard. 2019. How Time Constraints in a Creativity Support Tool Affect the Creative Writing Experience. In Proceedings of the 31st European Conference on Cognitive Ergonomics (BELFAST, United Kingdom) (ECCE ’19). Association for Computing Machinery, New York, NY, USA, 100–107. https: //doi.org/10.1145/3335082.3335084

[6] Margaret A Boden. 1994. What Is Creativity? (1994), 75–117. [7] Jay David Bolter and Michael Joyce. 1987. Hypertext and creative writing. In Proceedings of the ACM Conference on Hypertext (Chapel Hill, North Carolina, USA) (HYPERTEXT ’87). Association for Computing Machinery, New York, NY, USA, 41–50. https://doi.org/10.1145/317426.317431

[8] Ronald Carter. 2004. Language and creativity: The art of common talk. Routledge. [9] Nichol Castro and Cynthia S. Q. Siew. 2020. Contributions of modern network science to the cognitive sciences: Revisiting research spirals of representation and process. Proceedings of the Royal Society A 476, 2238 (2020), 20190825.

[10] Honghua Chen and Nai Ding. 2023. Probing the “Creativity” of Large Language Models: Can models produce divergent semantic association?. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 12881–12888. https://doi.org/10.18653/v1/2023.findings-emnlp.858

[11] Yueying Chu and Peng Liu. 2023. Public aversion against ChatGPT in creative fields? The Innovation 4, 4 (2023), 100449.

[12] David Cropley. 2023. Is artificial intelligence more creative than humans?: Chat- GPT and the divergent association task. Learning Letters 2 (2023), 13–13.

[13] Ido Dagan, Dan Roth, Mark Sammons, and Fabio Massimo Zanzotto. 2013. Recognizing textual entailment: Models and applications. Morgan & Claypool Publishers, San Rafael.

[14] Joshua Davis, Liesbet Van Bulck, Brigitte N Durieux, and Charlotta Lindvall. 2024. The Temperature Feature of ChatGPT: Modifying Creativity for Clinical Research. JMIR Hum Factors 11 (8 Mar 2024), e53559. https://doi.org/10.2196/53559

[15] María-Isabel de Vicente-Yagüe-Jara, Olivia López-Martínez, Verónica Navarro- Navarro, and Francisco Cuéllar-Santiago. 2023. Writing, Creativity, and Artificial Intelligence: ChatGPT in the University Context. Comunicar: Media Education Research Journal 31, 77 (2023), 45–54.

[16] Giorgio Franceschelli and Mirco Musolesi. 2023. On the creativity of large language models. arXiv preprint arXiv:2304.00008 (2023).

[17] Guillaume Fürst and François Grin. 2018. Multilingualism and creativity: A multivariate approach. Journal of Multilingual and Multicultural Development 39, 4 (2018), 341–355.

[18] Hans-Georg Gadamer. 1990. Wahrheit und Methode: Grundzüge einer philosophischen Hermeneutik. Mohr Siebeck, Tübingen.

[19] Harold Garfinkel. 1967. Studies in Ethnomethodology. Prentice Hall, Malden, MA. [20] Ken Gilhooly. 2024. AI vs humans in the AUT: Simulations to LLMs. Journal of Creativity 34, 1 (2024), 100071. https://doi.org/10.1016/j.yjoc.2023.100071

[21] Toni Giorgino. 2009. Computing and Visualizing Dynamic Time Warping Alignments in R: The dtw Package. Journal of Statistical Software 31, 7 (2009). https://doi.org/10.18637/jss.v031.i07

[22] J Pv Guilford. 1964. Some new looks at the nature of creative processes. In Contributions to mathematical psychology, N. Fredrickson and H. Gilliksen (Eds.). Holt, Rinehart & Winston, New York.

[23] Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang.

    Conformer: Convolution-augmented Transformer for Speech Recognition.

arXiv:2005.08100 [eess.AS]

[24] Erik E. Guzik, Christian Byrge, and Christian Gilde. 2023. The originality of machines: AI takes the Torrance Test. Journal of Creativity 33, 3 (2023), 100065. https://doi.org/10.1016/j.yjoc.2023.100065

[25] Jennifer Haase and Paul H.P. Hanel. 2023. Artificial muses: Generative artificial intelligence chatbots have risen to human-level creativity. Journal of Creativity

33, 3 (2023), 100066. https://doi.org/10.1016/j.yjoc.2023.100066

[26] Eric Harper, Somshubra Majumdar, Oleksii Kuchaiev, Li Jason, Yang Zhang, Evelina Bakhturina, Vahid Noroozi, Sandeep Subramanian, Koluguri Nithin, Huang Jocelyn, Fei Jia, Jagadeesh Balam, Xuesong Yang, Micha Livne, Yi Dong, Sean Naren, and Boris Ginsburg. [n. d.]. NeMo: a toolkit for Conversational AI and Large Language Models. https://github.com/NVIDIA/NeMo

[27] Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023. DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=sE7-XhLxHA

[28] Leah Henrickson and Albert Meroño-Peñuela. 2023. Prompting meaning: a hermeneutic approach to optimising prompt engineering with ChatGPT. AI & SOCIETY (2023), 1–16.

[29] Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd.

    spaCy: Industrial-strength Natural Language Processing in Python. (2020).

https://doi.org/10.5281/zenodo.1212303

[30] Daphne Ippolito, Ann Yuan, Andy Coenen, and Sehmon Burnam. 2022. Creative Writing with an AI-Powered Writing Assistant: Perspectives from Professional Writers. arXiv:2211.05030 [cs.HC]

[31] Fei Jia, Somshubra Majumdar, and Boris Ginsburg. 2021. MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 6818–6822. https://doi.org/10.1109/ICASSP39728. 2021.9414470

[32] Rodney H Jones. 2010. Creativity and discourse. World Englishes 29, 4 (2010), 467–480.

[33] Yoed N. Kenett, David Anaki, and Miriam Faust. 2014. Investigating the structure of semantic networks in low and high creative persons. Frontiers in human neuroscience 8 (2014), 407.

[34] Yoed N. Kenett, Roger E Beaty, Paul J Silvia, David Anaki, and Miriam Faust.

    Structure and flexibility: Investigating the relation between the structure

of the mental lexicon, fluid intelligence, and creative achievement. Psychology of Aesthetics, Creativity, and the Arts 10, 4 (2016), 377.

[35] Andruid Kerne and Eunyee Koh. 2007. Creativity support: the mixed-initiative composition space. In Proceedings of the 7th ACM/IEEE-CS Joint Conference on Digital Libraries (Vancouver, BC, Canada) (JCDL ’07). Association for Computing Machinery, New York, NY, USA, 509. https://doi.org/10.1145/1255175.1255309

[36] Walter Kintsch. 1998. Comprehension. A Paradigm for Cognition. Cambridge University Press, Cambridge.

[37] Keith Kirkpatrick. 2023. Can AI demonstrate creativity? Commun. ACM 66, 2 (2023), 21–23.

[38] Moritz Laurer, Wouter van Atteveldt, Andreu Salleras Casas, and Kasper Welbers.

    Less Annotating, More Classifying – Addressing the Data Scarcity Issue

of Supervised Machine Learning with Deep Transfer Learning and BERT - NLI. Preprint (June 2022). https://osf.io/74b8k Publisher: Open Science Framework.

[39] Min Liu. 1998. The effect of hypermedia authoring on elementary school students’ creative thinking. Journal of Educational Computing Research 19, 1 (1998), 27–51.

[40] Jérôme Louradour. 2023. whisper-timestamped. https://github.com/linto-ai/ whisper-timestamped.

[41] Janet Maybin and Joan Swann. 2007. Everyday creativity in language: Textuality, contextuality, and critique. Applied linguistics 28, 4 (2007), 497–517.

[42] Sarnoff Mednick. 1962. The associative basis of the creative process. Psychological review 69, 3 (1962), 220.

[43] Reijo Miettinen. 2013. Creative encounters and collaborative agency in science, technology and innovation. In Handbook of research on creativity. Edward Elgar Publishing, 435–449.

[44] Carey K Morewedge. 2022. Preference for human, not algorithm aversion. Trends in cognitive sciences 26, 10 (2022), 824–826.

[45] Ulrich Oevermann. 1986. Kontroversen über sinnverstehende Soziologie. Einige wiederkehrende Probleme und Mißverständnisse in der Rezeption der “objektiven Hermeneutik”. In Handlung und Sinnstruktur: Bedeutung und Anwendung der objektiven Hermeneutik, S. Aufenanger and M. Lenssen (Eds.). Kindt München, 19–83.

[46] Ulrich Oevermann, With Tilman Allert, Elisabeth Konau, and Jürgen Krambeck.

    21. Structures of Meaning and Objective Hermeneutics. In Modern German

Sociology, Volker Meja, Dieter Misgeld, and Nico Stehr (Eds.). Columbia University Press, New York Chichester, West Sussex, 436–448. https://doi.org/doi:10.7312/ meja92024-024

[47] Jay A. Olson, Johnny Nahas, Denis Chmoulevitch, Simon J. Cropper, and Margaret E. Webb. 2021. Naming unrelated words predicts creativity. Proceedings of the National Academy of Sciences 118, 25 (2021), e2022340118. https://doi.org/10.1073/pnas.2022340118 arXiv:https://www.pnas.org/doi/pdf/10.1073/pnas.2022340118

[48] OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] [49] OpenAI. 2024. Best practices for prompt engineering with the OpenAI API. https://help.openai.com/en/articles/6654000-best-practices-for-promptengineering-with-the-openai-api. [Accessed April 12, 2024].

[50] OpenAI. 2024. OpenAI API. https://platform.openai.com/docs/apireference/completions. Accessed: April 14, 2024.

[51] Peter Organisciak, Selcuk Acar, Denis Dumas, and Kelly Berthiaume. 2023. Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity 49 (2023),

    https://doi.org/10.1016/j.tsc.2023.101356

[52] Howard H. Pattee. 1989. Simulations, Realizations, and Theories of Life. In Artificial Life. SFI Studies in the Sciences of Complexity, Christopher G. Langton (Ed.). Addison-Wesley, Redwood, 63–77.

[53] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12 (2011), 2825–2830.

[54] Martin J. Pickering and Simon Garrod. 2004. Toward a mechanistic psychology of dialogue. Behavioral and Brain Sciences 27 (2004), 169–226.

[55] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023. Robust Speech Recognition via Large-Scale Weak Supervision. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 28492–28518. https://proceedings.mlr.press/v202/radford23a.html

[56] Ralph Raiola. 2023. ChatGPT, Can You Tell Me a Story? An Exercise in Challenging the True Creativity of Generative AI. Commun. ACM 66, 5 (2023), 104–ff.

[57] Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. https://arxiv.org/abs/1908.10084

[58] Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. 1974. A Simplest Systematics for the Organization of Turn Taking for Conversation. Language 50 (1974), 696–735.

[59] Keith Sawyer. 2012. Explaining Creativity: The Science of Human Innovation. Oxford University Press, Oxford.

[60] Keith Sawyer. 2016. Creativity and dialogue: The improvisational nature of conversational interaction. In The Routledge handbook of language and creativity.

Routledge, 78–91.

[61] R Keith Sawyer and Stacy DeZutter. 2009. Distributed creativity: How collective creations emerge from collaboration. Psychology of aesthetics, creativity, and the arts 3, 2 (2009), 81.

[62] Verena D. Schmittmann, Angélique O.J. Cramer, Lourens J. Waldorp, Sacha Epskamp, Rogier A. Kievit, and Denny Borsboom. 2013. Deconstructing the construct: A network perspective on psychological phenomena. New Ideas in Psychology 31, 1 (2013), 43–53. https://doi.org/10.1016/j.newideapsych.2011.02.007

[63] Dean Keith Simonton. 2013. What is a creative idea? Little-c versus Big-C creativity. In Handbook of research on creativity. Edward Elgar Publishing, 69–83.

[64] Douglas Summers-Stay, Clare R Voss, and Stephanie M Lukin. 2023. Brainstorm, then select: a generative language model improves its creativity score. In The AAAI-23 Workshop on Creative AI Across Modalities.

[65] Viriya Taecharungroj. 2023. “What can ChatGPT do?” Analyzing early reactions to the innovative AI chatbot on Twitter. Big Data and Cognitive Computing 7, 1 (2023), 35.

[66] Ellis Paul Torrance. 2017. Torrance tests of creative thinking. [67] Jared Vasil, Paul B Badcock, Axel Constant, Karl Friston, and Maxwell JD Ramstead. 2020. A world unto itself: human communication as active inference. Frontiers in psychology 11 (2020), 480375.

[68] Andreas Wernet. 2006. Einführung in die Interpretationstechnik der objektiven Hermeneutik. Springer.

[69] Ching-Lin Wu, Shih-Yuan Huang, Pei-Zhen Chen, and Hsueh-Chih Chen. 2020. A Systematic Review of Creativity-Related Studies Applying the Remote Associates Test From 2000 to 2019. Frontiers in Psychology 11 (2020).

[70] Yunpu Zhao, Rui Zhang, Wenyi Li, Di Huang, Jiaming Guo, Shaohui Peng, Yifan Hao, Yuanbo Wen, Xing Hu, Zidong Du, Qi Guo, Ling Li, and Yunji Chen. 2024. Assessing and Understanding Creativity in Large Language Models. arXiv:2401.12491 [cs.CL]

[71] Rolf A. Zwaan. 2016. Situation models, mental simulations, and abstract concepts in discourse comprehension. Psychonomic bulletin & review 23, 4 (2016), 1028– 1034.

Figures

3648188.3675140 figure-01.png
3648188.3675140 figure-02.png
3648188.3675140 figure-03.png
3648188.3675140 figure-04.png
3648188.3675140 figure-05.png
3648188.3675140 figure-06.png
3648188.3675140 figure-07.png
3648188.3675140 figure-08.png
3648188.3675140 figure-09.png
3648188.3675140 figure-10.png
3648188.3675140 figure-11.png
3648188.3675140 figure-12.png
3648188.3675140 figure-13.png
3648188.3675140 figure-14.png
3648188.3675140 figure-15.png
3648188.3675140 figure-16.png
3648188.3675140 figure-17.png
3648188.3675140 figure-18.png
3648188.3675140 figure-19.png
3648188.3675140 figure-20.png
3648188.3675140 figure-21.png
3648188.3675140 figure-22.png
3648188.3675140 figure-23.png
3648188.3675140 figure-24.png
3648188.3675140 figure-25.png
3648188.3675140 figure-26.png
3648188.3675140 figure-27.png
3648188.3675140 figure-28.png
3648188.3675140 figure-29.png
3648188.3675140 figure-30.png
3648188.3675140 figure-31.png
3648188.3675140 figure-32.png
3648188.3675140 figure-33.png
3648188.3675140 figure-34.png
3648188.3675140 figure-35.png
3648188.3675140 figure-36.png
3648188.3675140 figure-37.png

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime