HyperCausal: Visualizing Causal Inference in 3D Hypertext
Author: Kevin Bönisch, Goethe-University, Germany, k.boenisch@outlook.com
Author: Manuel Stoeckel, Goethe-University, Germany, manuel.stoeckel@em.uni-frankfurt.de
Author: Alexander Mehler, Goethe-University, Germany, mehler@em.uni-frankurt.de
DOI: https://doi.org/10.1145/3648188.3677049 HT '24: 35th ACM Conference on Hypertext and Social Media, Poznan, Poland, September 2024
Abstract
We present HyperCausal, a 3D hypertext visualization framework for exploring causal inference in generative Large Language Models (LLMs). HyperCausal maps the generative processes of LLMs into spatial hypertexts, where tokens are represented as nodes connected by probability-weighted edges. The edges are weighted by the prediction scores of next tokens, depending on the underlying language model. HyperCausal facilitates navigation through the causal space of the underlying LLM, allowing users to explore predicted word sequences and their branching. Through comparative analysis of LLM parameters such as token probabilities and search algorithms, HyperCausal provides insight into model behavior and performance. Implemented using the Hugging Face transformers library and Three.js, HyperCausal ensures cross-platform accessibility to advance research in natural language processing using concepts from hypertext research. We demonstrate several use cases of HyperCausal and highlight the potential for detecting hallucinations generated by LLMs using this framework. The connection with hypertext research arises from the fact that HyperCausal relies on user interaction to unfold graphs with hierarchically appearing branching alternatives in 3D space. This approach refers to spatial hypertexts and early concepts of hierarchical hypertext structures. A third connection concerns hypertext fiction, since the branching alternatives mediated by HyperCausal manifest non-linearly organized reading threads along artificially generated texts that the user decides to follow optionally depending on the reading context.
CCS Concepts: • Information systems → Information systems applications; • Human-centered computing → Visualization systems and tools;
Keywords: 3D hypertext, large language models, visualization
ACM Reference Format: Kevin Bönisch, Manuel Stoeckel, and Alexander Mehler. 2024. HyperCausal: Visualizing Causal Inference in 3D Hypertext. In 35th ACM Conference on Hypertext and Social Media (HT '24), September 10--13, 2024, Poznan, Poland. ACM, New York, NY, USA 7 Pages. https://doi.org/10.1145/3648188.3677049
(b) Selecting the node And from a random branch of the network. The popup on the left provides a description of the inferred sequence alongside their respective probabilities. Figure 1: Overview of a HyperCausal network visualizing GPT-2 [30] with k = 3 branches, $max_tokens=7$ and input=”Once upon a time there was a boy“. The nodes are depicted in a golden color, while the edges are rendered in gray. The opacity of each element represents the probability of the corresponding token inferred by the LLM, ranging from 0 to 1. Thus, increased brightness of both edges and nodes signifies a higher likelihood of the token being present.
1 MOTIVATION
Starting with the seminal work on semantic spaces by [29], the graphical visualization of semantic structures based on feature spaces has a long tradition in computational linguistics. Early work mostly focused on lexical spaces to graphically visualize semantic associations or meaning relations. This includes tree-like models based, e.g., on dependency trees [31], similarity trees [20] or cohesion trees [23], all of which represent lexical associations hierarchically starting from seed words. It was long before the advent of network theory [27] that the visualization of general graph structures became common. This tradition is reflected in early hypertext models for making network structures accessible in overview representations. There is an abundance of graph models for the visualization of hierarchical [2, 15] or network-like hypertext structures [3, 26]. Most notably, spatial hypertexts [7, 22] can be considered as systems whose structure is induced or superimposed by spatial relations intended to facilitate orientation in hypertext.
In this paper we present HyperCausal, a system for visualizing semantic relations between linguistic units based on generative AI. That is, instead of visualizing genuinely hypertextual structures using methods from graph theory or computational linguistics, we take the opposite approach: we visualize semantic structures from high-dimensional feature spaces using concepts from spatial hypertexts to make them traversable for users. In this way, we make the options that are hidden from the user transparent by presenting alternative predictions based on the same word (sequence) according to the underlying AI, and show which subspaces these alternatives make accessible. This is important because parameters, such as temperature [28] , and the algorithms used for generating texts, such as beam search [12] , directly influence the selection from these alternatives without the user being aware of the underlying space of options and their consequences. In this sense, HyperCausal can be seen as a spatial hypertext that uses 3D spatial structures to make the prediction alternatives generated by AI models both traversable and interactively tangible. Since the underlying AI model can be varied, HyperCausal also serves as a tool for comparing different AI models. This comparison can refer to syntagmatic relations (How is the given sequence continued?) and paradigmatic relations (Which words are available as alternatives at the same position?). The paper is organized as follows: Section 2 reviews related work, Section 3 introduces HyperCausal, Section 4 describes future work, and Section 5 gives a conclusion.
2 RELATED WORK
In recent years, numerous visual analytics techniques have been developed to enhance the interpretability and transparency of machine learning models, particularly in the context of natural language processing (NLP) and deep learning in general. These techniques can be broadly categorized based on their application stage: prior to, during, or after model training, as surveyed by [37]. The importance of visualization in model interpretability is emphasized by [38], who highlight that visualizations provide intuitive, interactive, and explicit ways to diagnose bias, identify errors, and understand model mechanisms. For instance, [19] developed an interactive beam search visualization for sequence-to-sequence models, which includes a graph representation of attention scores between the input and the generated output sequence. Similarly, [21] visualized attention scores in natural language inference tasks, presenting them alongside dependency parse trees and offering zoomable attention graphs and matrices for a detailed view of machine comprehension processes. [17] introduced ActiVis, a model-oriented visualization tool that displays a model's computation graph and neuron activations in relation to downstream tasks. This enables engineers to analyze activation patterns and evaluate model performance more effectively. In the domain of sequence-to-sequence models, [33] created interactive visualizations for beam search, including a translation view that allows users to manipulate beam search paths and a neighborhood view that projects generated sequences into a 2D space for comparative analysis. [24] presented Text2voronoi, a system for mapping texts onto Voronoi diagrams based on word embeddings of their lexical constituents. Text2voronoi is used as visual input for a classifier to predict constructs based on subjects’ linguistic output, analogous to medical imaging techniques. Beyond specific NLP-related visualizations, [6] introduced Activation Atlas, a tool for understanding the representations learned by deep neural networks, especially convolutional neural networks (CNNs). This tool facilitates exploration of the high-dimensional feature spaces learned by CNNs by visualizing the image masks of the averaged activations on a 2D plane. Focusing on transformer-based models, [34] visualized attention scores and key-query-value vectors within BERT [10] and GPT-2 [30] models, enabling the exploration of attention mechanisms across different layers. [9] extended this work by visualizing attention flows within and across layers and among attention heads in transformer models, allowing users to trace and compare attention scores to understand classification decisions made by the models. [36] further advanced this by proposing AttentionViz, which includes a so-called global perspective that, unlike previous visualizations, enables the analysis of global patterns across multiple input sequences while focusing on the query-key vector interactions. AttentionViz also facilitates the inspection of both language and vision transformers. [18] use 3D visualizations to make embedding models for words comparable, e.g. by coloring the visualized lexical elements according to their part of speech. The embeddings can be varied according to the corresponding layer within the underlying neural network. Finally, [5] expanded the visualization capabilities beyond 2D by introducing a 3D representation of GPT, effectively illustrating each step of the transformer network. This 3D visualization is complemented by a 2D UI panel that guides users through each step with clear explanations and navigates seamlessly through each layer in the 3D transformer. Although this visualization emphasizes the educational aspect of transformers, it nevertheless reveals the inner workings of these models when provided with an input, potentially aiding even experts in their decoding process. While these advancements underscore the importance of visualization in comprehending NLP models, 3D visualizations remain exceedingly rare compared to their 2D counterparts.
The main difference between 3D and 2D visualizations is not just the addition of another dimension. Rather, it is the spatial representation of the information encoded in NLP models. According to this view, the interaction with HyperCausal stimulates spatial learning or thinking, allowing the user to map the underlying abstract (essentially numerical) information from high-dimensional feature spaces in a way that goes beyond point clouds, clusters, Voronoi diagrams, and the like. There is a second key difference from the latter visualization methods that brings us into the realm of hypertext: this is the interactive element of HyperCausal, according to which the user is not presented with a pre-generated view of the underlying information space. Rather, it is the user's own interactions that determine what they actually see. That is, the user actively chooses the perspective or direction under which the underlying information structure is unfolded as a graph to provide further branching options. This is different from interactive point clouds, where the user can zoom in or out or select a subset of points, because in those cases the user is initially shown a full view of all information units, which he then manipulates to get a partial view. In HyperCausal this is different: here the user starts with an initial sequence (in extreme cases with a single word) and then decides which of its alternative extensions to follow. Finally, instead of representing associations between words that occur independently of syntactic relations, HyperCausal explores inference scores to generate word sequences that form candidate syntactic or textual structures of the underlying language. In this way, the artificially generated structures that are displayed in 3D with HyperCausal appear as a non-linearly organized text somewhat reminiscent of hypertext fiction (see e.g. [32]) and hierarchically ordered NoteCards [15]. And while in most of the approaches listed above the information structure to be displayed is predetermined, making the user a selector of information units and their networking rather than a creator, in HyperCausal he can choose among parameters to change the sequencing and branching in the information structure unfolded by his very interaction.
3 HyperCausal
In this section, we describe HyperCausal, an open-source1 3D hypertext visualization framework for inference based on causal LLMs. Using the concepts of 3D spatial hypertext, it fills the gap in 3D visualization identified in Section 2.
A more zoomed-out cross-section excerpt of a GPT-2 HyperCausal network with k=6 and maxtokens=5, illustrating its structure and appearance in the 3D space. Figure 2: A more zoomed-out cross-section excerpt of a GPT-2 HyperCausal network with k = 6 and $max_tokens=5$, illustrating its structure and appearance in the 3D space.
3.1 LLM Inference as a Hypertext System
HyperCausal projects causal inference of Large Language Models (LLMs) into a 3D hypertext space, wherein each generated token during the LLM's generative process is represented as a node, with the associated probability of token generation depicted as an edge. Given that Causal LLMs are generative pre-trained transformers designed to autoregressively predict the subsequent token based exclusively on preceding context, a generated sequence can be conceptualized as a directed branch comprising tokens (nodes) and probabilities (edges), which recursively grows through its predecessors. This structure then encapsulates a closed, ideally semantically coherent sequence of discrete linguistic units. Additionally, besides visualizing the primary sequence generated by the LLM, HyperCausal also considers the top k potential alternative sequences at each node, incorporating these as individual branches. This results in a 3D directional network that expands in width, height, and depth within the 3D space as outlined in Figure 1 and Figure 2. Algorithm 1 outlines HyperCausal’s core, breadth-first procedure of building such a network. The visualization enables comparative analyses across various autoregressive language models and their parameters. This includes evaluating the token's probabilities, comparing different temperature settings, logit biases, or search algorithms and what impact they have on the LLM's causal space. As a result, we enable users to visualize branching in predicted word sequences with different numbers of search beams, top-K [11] and top-p (nucleus) sampling [16] settings. Further, we force the choice of certain tokens either by selection in the UI or by encouraging their generation using a sequence bias that directly modifies the model output logits. Additionally, it unveils an otherwise concealed underlying space to users, presenting alternative branches for exploration. Finally, as HyperCausal is implemented around the Hugging Face transformers library2 and visualized through Three.js3, it can be applied to existing codebases and used on any browser and device.
Algorithm 1. HyperCausal’s simple breadth-first graph generation
procedure PlantTree(inTxt, start)
r ← NewRoot(inTxt, start)
q ← [r]
while q do
branch ← q.dequeue() ▷ use stack for depth-first
Grow(branch) ▷ Visualize in UI
nextBranches ← FetchNext(branch)
for b in nextBranches do
if b.tok ≠ '<EOS>' then
q.enqueue(b)
end if
end for
end while
end procedureAn example of HyperCausal's potential for hallucination detection, demonstrated by visualizing its network for GPT-3 [4] with the input ”Question: What is the capital of Italy? Answer:“ using k=2, maxtokens=5 and greedy search as a decoding strategy. The thin green branch from the root of the tree to the top right indicates the primary inference the model would output, which shows a hallucination, as evidenced by zooming in on the generated sequence via the red box. The correct answer appears in the tertiary sequence with a much lower probability of 0.0078%, highlighted through the zoomed-in green box, which would not have been outputted. In this specific case, adjusting the prompt would be an adequate solution, as the very first generated token led to sequence branches where none provided the correct information. Furthermore, it may be advisable to consider changing the decoding strategy to something more contextually robust, such as Beam search. This recommendation stems from the observation that the more probable branches, indicated by the higher density of brighter edges in the bottom area of the network, were obscured by the low probability of the first token of the sequence branch. Figure 3: An example of HyperCausal’s potential for hallucination detection, demonstrated by visualizing its network for GPT-3 [4] with the input ”Question: What is the capital of Italy? Answer:“ using k = 2, $max_tokens=5$ and greedy search as a decoding strategy. The thin green branch from the root of the tree to the top right indicates the primary inference the model would output, which shows a hallucination, as evidenced by zooming in on the generated sequence via the red box. The correct answer appears in the tertiary sequence with a much lower probability of $0.0078%$, highlighted through the zoomed-in green box, which would not have been outputted. In this specific case, adjusting the prompt would be an adequate solution, as the very first generated token led to sequence branches where none provided the correct information. Furthermore, it may be advisable to consider changing the decoding strategy to something more contextually robust, such as Beam search. This recommendation stems from the observation that the more probable branches, indicated by the higher density of brighter edges in the bottom area of the network, were obscured by the low probability of the first token of the sequence branch.
3.2 Use Cases
This section outlines use cases for HyperCausal and provides guidance on how to use the framework in these contexts.
3.2.1 Explainability. Besides the apparent necessities for interpretable or explainable AI [1, 13, 14], developers of applications that leverage generative language models will also run into ubiquitous practical issues, like unexpected behaviour of models or deviation from an expected pattern. To this end, we tailored HyperCausal’s capabilities to leverage the intuitive nature of traversing a 3D space, thereby facilitating the easy identification of unexpected deviations, also showcased by Figure 3. Additionally, this also enables a form of information confidence verification, as firstly, the network would highlight weakly chosen probabilities through less opaque edges and nodes. Secondly, it facilitates further examination by detailing the probabilities, showcasing the sequence and its surrounding network. This process identifies potentially weak informational links, as in: uncertain yet possibly critical tokens (non-stopwords for example) that exhibit low probability.
3.2.2 Prompt Engineering. When prototyping during prompt engineering, developers could explore why a certain prompt yields unexpected results by traversing the output tree, identify critical nodes and experiment with different settings from the UI to achieve the desired outcome. Prompt engineering based on black-box LLMs is increasingly becoming an alternative to classical NLP-related modeling using explicit models. Therefore, methods are needed that allow both, making transparent the LLM-based co-generation of prompts and the analysis of the prompted results. HyperCausal is a method that supports these two approaches to prompt engineering.
3.2.3 Parameter Tuning. As the trees are embedded in 3D space, users can explore the effect of multiple parameters at the same time and compare the probabilities of single nodes or their joint probability – and thus quickly compare branches of different lengths. Furthermore, the Continue this branch button within HyperCausal’s user interface allows users to specifically continue and analyze a sequence branch even after the network has been generated, making it possible to inspect and compare different branches. This is especially useful when comparing different decoding strategies, such as greedy- and beam search, or different sampling techniques.
3.2.4 Hallucination Detection. By providing both prompt and answer, users can visualize how the predictions of the language model diverge from the given text and analyze the impact of generation hyperparameters. This need is underscored by the propensity of LLMs to generate hallucinations [35], which can lead to catastrophic misunderstandings of information [8] in critical fields such as law, medicine, education, or science. To this end, HyperCausal can help identify such hallucinations by visualizing each token in the sequence along with its associated low probabilities and alternative branches. This is demonstrated by Figure 3, where an incorrect information sequence detected by HyperCausal’s network traversal could be overlooked by the LLM. The correct inference appears only in the tertiary sequence branch, as highlighted by the green box. In cases like this, HyperCausal makes transparent word sequences that are alternatively predictable by a language model given an input string (e.g. a prompt). This is done in such a way as to reveal hallucinatory sequences among the underlying alternatives that would otherwise be overlooked (that is, by taking a single prediction alternative).
3.2.5 Educational context. HyperCausal’s 3D visualization of causal LLM inference, complemented by its intuitive navigation features, can provide significant educational value by translating abstract and often challenging concepts and results into a more comprehensible 3D space. The interactive elements further enhance understanding and make the content accessible to beginners. The reason for this evaluation is again the space of alternatives that HyperCausal makes interactively accessible: instead of leaving the user alone with a single possible answer to a prompt, he or she can interactively play with alternative answers that contextualize the preferred one and thus provide a discursive ground on which the user can think about the relevance of the preferred answer or other reasons for its justification. Consider the example of Critical Online Reasoning (COR) [25], where students seek information from the Web to begin reasoning in order to complete a particular task. Obviously, in such educational application scenarios, generative AI-based chatbots become a serious alternative to classic web search, but with the caveat that students would actually have to subject the chatbot-generated answer to one or more validity tests. With HyperCausal, we create a first such testing possibility within the environment of the underlying AI software, allowing learners to contextualize given answers in relation to alternative answers that the system could have given with certain deviating probabilities. By comparing these answers, we ultimately encourage students to activate their critical thinking and at least question any prompted answer. Such a HyperCausal-based test cannot replace content-related validity tests, but it can prepare for them by encouraging a critical attitude in terms of COR.
4 FUTURE WORK
In future work, we intend to enhance HyperCausal’s user interface features to enable direct parameter tuning within the interface itself. This improvement aims to streamline the process, eliminating the need for users to utilize the command-line interface (CLI) to pass parameters, as is currently required. Furthermore, we aim to place increased emphasis on the use cases outlined in Section 3.2, particularly regarding hallucination detection. We find that HyperCausal’s 3D visualization facilitates the identification of weakly predicted sequences and tokens that may indicate potential hallucinations. To enhance this capability, we intend to implement a heuristic that automatically detects such instances based on various factors, including the tokens themselves, their probabilities, and crucially, the surrounding HyperCausal network and its alternative branches. Moreover, we aim to incorporate these findings into the 3D space for enhanced visualization and analysis. Finally, as this is the first version of HyperCausal, we aim to conduct a systematic evaluation of it, encompassing both expert users and novices, to assess the ease of interface facilitation.
5 CONCLUSION
In this paper, we introduced HyperCausal, a 3D hypertext visualization system that represents causal inference relations between linguistic items based on LLMs in a 3D hypertext space. By doing so, our framework offers a visualization of the semantic relations between linguistic units and provides insights into the inner workings of generative LLMs. Through its intuitive interface and interactive features, HyperCausal enables users to explore alternative sequences and their branching, compare different language models and their parameters, and analyze the impact of generation hyperparameters. Additionally, it facilitates developers in comprehending and debugging the behavior of generative language models, aids in prompt engineering, and enables detecting potential hallucinations produced by the models. Finally, HyperCausal can serve as a useful tool for improving the interpretability and transparency of machine learning models, especially in the context of natural language processing, but also in application contexts such as higher education. As our framework is built upon the widely-used transformers library, it can be easily integrated into existing codebases and accessed across various devices, making it accessible to a broad audience of researchers and practitioners.
HyperCausal is currently a tool for visualizing sequences of word predictions based on autoregressive language models. As such, its visualization logic draws on several concepts from hypertext research: first, the visualization space is 3D, using spatial relations to distribute word sequences and their branching in a partially hierarchical manner. Second, HyperCausal relies heavily on interaction with the user, who decides how to unfold the prediction alternatives, using hyperparameters that allow context-sensitive adaptation of the underlying information structure. Third, by unfolding different branching alternatives, the user explores a non-linear text, albeit artificially generated, reminiscent of text structures studied in hypertext fiction. This is all the more so since the user can generate the initial text sequence, which is then expanded by HyperCausal interacting with the user. This means that the final text is more or less generated by user input, machine output, and user interaction behavior. In this way, HyperCausal ultimately shows how to integrate information logic elaborated in hypertext research with visualization techniques used to make understandable and comprehensible information generated by current large language models.
ACKNOWLEDGMENTS
This work was supported by the project CORE (Critical Online Reasoning in Higher Education; FOR 5404, project number 462702138), subprojects B05 and C08, and the BIOfid project (grant number ME 2746/5-1), both funded by the German Research Foundation (DFG).
REFERENCES
Saranya A. and Subhashini R.2023. A systematic review of Explainable Artificial Intelligence models and applications: Recent developments and future trends. Decision Analytics Journal 7 (2023), 100230. https://doi.org/10.1016/j.dajour.2023.100230 - Espen Aarseth. 1995. Cybertext: perspectives on ergodic literature. University of Bergen. - Mark Bernstein. 2002. Storyspace 1. In Proceedings of the Thirteenth ACM Conference on Hypertext and Hypermedia (College Park, Maryland, USA) (HYPERTEXT ’02). Association for Computing Machinery, New York, NY, USA, 172–181. https://doi.org/10.1145/513338.513383 - Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. arxiv:2005.14165 [cs.CL] - Brendan Bycroft. 2023. LLM Visualization. https://github.com/bbycroft/llm-viz. 3D Visualization of a GPT-style LLM. - Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah. 2019. Activation Atlas. Distill (2019). https://doi.org/10.23915/distill.00015 https://distill.pub/2019/activation-atlas. - C. Chen and M. Czerwinski. 1998. From Latent Semantics to Spatial Hypertext: An Integrated Approach. In Proceedings of 9th ACM Conference on Hypertext and Hypermedia, K. Grønbæk, E. Mylonas, and F. M. Shipman (Eds.). ACM, New York, 77–86. - Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E. Ho. 2024. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. arxiv:2401.01301 [cs.CL] - Joseph F. DeRose, Jiayao Wang, and Matthew Berger. 2021. Attention Flows: Analyzing and Comparing Attention Mechanisms in Language Models. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2021), 1160–1170. https://doi.org/10.1109/TVCG.2020.3028976 arxiv:2009.07053 [cs.HC] - Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186. https://doi.org/10.18653/v1/N19-1423 - Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical Neural Story Generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Iryna Gurevych and Yusuke Miyao (Eds.). Association for Computational Linguistics, Melbourne, Australia, 889–898. https://doi.org/10.18653/v1/P18-1082 - Markus Freitag and Yaser Al-Onaizan. 2017. Beam Search Strategies for Neural Machine Translation. In Proceedings of the First Workshop on Neural Machine Translation. Association for Computational Linguistics. https://doi.org/10.18653/v1/w17-3207 - Julie Gerlings, Arisa Shollo, and Ioanna Constantiou. 2021. Reviewing the Need for Explainable Artificial Intelligence (xAI). arxiv:2012.01007 [cs.HC] - Prashant Gohel, Priyanka Singh, and Manoranjan Mohanty. 2021. Explainable AI: current status and future directions. arxiv:2107.07045 [cs.LG] - Frank G. Halasz. 1988. Reflections on NoteCards: Seven Issues for the Next Generation of Hypermedia Systems. Commun. ACM 31, 7 (1988), 836–852. - Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net. https://openreview.net/forum?id=rygGQyrFvH - Minsuk Kahng, Pierre Y. Andrews, Aditya Kalro, and Duen Horng Chau. 2018. ActiVis: Visual Exploration of Industry-Scale Deep Neural Network Models. IEEE Transactions on Visualization and Computer Graphics 24, 1 (2018), 88–97. https://doi.org/10.1109/TVCG.2017.2744718 arxiv:1704.01942 [cs.HC] - Rebecca Kehlbeck, Rita Sevastjanova, Thilo Spinner, Tobias Stähle, and Mennatallah El-Assady. 2021. Demystifying the Embedding Space of Language Models. https://bert-vs-gpt2.dbvis.de/. - Jaesong Lee, Joong-Hwi Shin, and Jun-Seok Kim. 2017. Interactive Visualization and Manipulation of Attention-based Neural Machine Translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Lucia Specia, Matt Post, and Michael Paul (Eds.). Association for Computational Linguistics, Copenhagen, Denmark, 121–126. https://doi.org/10.18653/v1/D17-2021 - Dekang Lin. 1998. Automatic Retrieval and Clustering of Similar Words. In Proceedings of the COLING-ACL ’98. 768–774. - Shusen Liu, Tao Li, Zhimin Li, Vivek Srikumar, Valerio Pascucci, and Peer-Timo Bremer. 2018. Visual Interrogation of Attention-Based Models for Natural Language Inference and Machine Comprehension. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Eduardo Blanco and Wei Lu (Eds.). Association for Computational Linguistics, Brussels, Belgium, 36–41. https://doi.org/10.18653/v1/D18-2007 - Catherine C. Marshall and Frank M. Shipman III. 1993. Searching for the Missing Link: Discovering Implicit Structure in Spatial Hypertext. In Proceedings of the Fifth ACM Conference on Hypertext. ACM, 217–230. - Alexander Mehler. 2002. Hierarchical Orderings of Textual Units. In Proceedings of the COLING-ACL ’02. Morgan Kaufmann, 646–652. - Alexander Mehler, Tolga Uslu, and Wahed Hemati. 2016. Text2voronoi: An Image-driven Approach to Differential Diagnosis. In Proc. of the 5th Workshop on Vision and Language, hosted by ACL 2016 (VL@ACL 2016). - Marie-Theres Nagel, Svenja Schäfer, Olga Zlatkin-Troitschanskaia, Christian Schemer, Marcus Maurer, Dimitri Molerov, Susanne Schmidt, and Sebastian Brückner. 2020. How do university students’ web search behavior, website characteristics, and the interaction of both influence students’ critical online reasoning?. In Frontiers in Education, Vol. 5. Frontiers Media SA, 565062. - T. H. Nelson. 1965. Complex information processing: a file structure for the complex, the changing and the indeterminate. In Proceedings of the 1965 20th National Conference (Cleveland, Ohio, USA) (ACM ’65). Association for Computing Machinery, New York, NY, USA, 84–100. https://doi.org/10.1145/800197.806036 - Mark E. J. Newman. 2010. Networks: An Introduction. Oxford University Press, Oxford. - OpenAI. 2023. GPT-4 Technical Report. arxiv:2303.08774 [cs.CL] - Charles Egerton Osgood, George J. Suci, and Percy H. Tannenbaum. 1957. The measurement of meaning. University of Illinois Press, Urbana, IL. - Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9. - Burghard Rieger. 1984. Semantic Relevance and Aspect Dependency in a Given Subject Domain. In Proceedings of the COLING-ACL ’84. 298–301. - Simon Rowberry. 2011. Vladimir Nabokov's pale fire: the lost ’father of all hypertext demos’?. In Proceedings of the 22nd ACM Conference on Hypertext and Hypermedia (Eindhoven, The Netherlands) (HT ’11). Association for Computing Machinery, New York, NY, USA, 319–324. https://doi.org/10.1145/1995966.1996008 - Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M. Rush. 2019. Seq2seq-Vis: A Visual Debugging Tool for Sequence-to-Sequence Models. IEEE Transactions on Visualization and Computer Graphics 25, 1 (2019), 353–363. https://doi.org/10.1109/TVCG.2018.2865044 arxiv:1804.09299 [cs.CL] - Jesse Vig. 2019. A Multiscale Visualization of Attention in the Transformer Model. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Marta R. Costa-jussà and Enrique Alfonseca (Eds.). Association for Computational Linguistics, Florence, Italy, 37–42. https://doi.org/10.18653/v1/P19-3007 - Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024. Hallucination is Inevitable: An Innate Limitation of Large Language Models. arxiv:2401.11817 [cs.CL] - Catherine Yeh, Yida Chen, Aoyu Wu, Cynthia Chen, Fernanda Viégas, and Martin Wattenberg. 2023. AttentionViz: A Global View of Transformer Attention. arxiv:2305.03210 [cs.HC] - Jun Yuan, Changjian Chen, Weikai Yang, Mengchen Liu, Jiazhi Xia, and Shixia Liu. 2021. A survey of visual analytics techniques for machine learning. Comput. Vis. Media 7, 1 (2021), 3–36. https://doi.org/10.1007/S41095-020-0191-7 - Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for Large Language Models: A Survey. ACM Trans. Intell. Syst. Technol. 15, 2, Article 20 (feb 2024), 38 pages. https://doi.org/10.1145/3639372
FOOTNOTE
1 https://github.com/TheItCrOw/PrismAI, accessed July 9th, 2024
2 https://huggingface.co/docs/transformers/index, accessed May 16th, 2024
3 https://threejs.org/ accessed May 16th, 2024
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org.
HT '24, September 10–13, 2024, Poznan, Poland
© 2024 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 979-8-4007-0595-3/24/09. DOI: https://doi.org/10.1145/3648188.3677049
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime