WriteAssist: A Personalized Generative AI System for Autonomous Authoring of Scholarly Literature Reviews
Authors
Kallol Naha — Department of Computer Science, University of Idaho, Moscow, ID, USA naha7197@vandals.uidaho.edu; ORCID: 0000-0002-1815-234X
Sajratul Y Rubaiat — Department of Computer Science, University of Idaho, Moscow, ID, USA ruba3062@vandals.uidaho.edu; ORCID: 0009-0001-0367-644X
Syed N. Sakib — Department of Computer Science, University of Idaho, Moscow, ID, USA saki3064@vandals.uidaho.edu; ORCID: 0009-0000-2653-2767
Hasan M Jamil — Department of Computer Science, University of Idaho, Moscow, ID, USA jamil@uidaho.edu; ORCID: 0000-0002-3124-3780; corresponding author
Published in Proceedings of the 36th ACM Conference on Hypertext and Social Media (HT '25) · DOI: 10.1145/3720553.3746679 · Pages: 33–37 · License: CC BY-NC-ND 4.0
Abstract
In an era of information overload, research writing, particularly literature review composition, has become increasingly burdensome due to the sheer volume of scholarly publications released each year. This paper introduces WriteAssist, a novel standalone authoring system that helps researchers efficiently generate literature review sections. Given the title and abstract of a work-in-progress manuscript, WriteAssist automatically retrieves relevant and recent peer-reviewed articles, highlighting portions that offer supporting or contrasting perspectives. A key innovation lies in its personalized recommendation engine, which tailors results based on the user’s prior publications and research profile, enabling context-aware synthesis. We position WriteAssist within the landscape of intelligent writing assistants, academic search platforms, and personalized recommender systems, and we detail its architecture – integrating natural language processing and user modeling to streamline academic writing. The system represents a significant step toward alleviating cognitive overload in scholarly composition and offers a blueprint for smarter, adaptive tools in academic research support.
Introduction
Writing a comprehensive literature review has become an increasingly burdensome task in the face of exponential growth in scholarly publications [[2], [8]]. Each year, millions of new articles are added to digital libraries, making it progressively more difficult for researchers to identify, organize, and synthesize relevant literature [[8]]. This explosion of information not only increases the risk of overlooking critical work but also contributes to what many scholars describe as digital clutter – an overwhelming accumulation of PDFs, notes, and search results that hinders focused writing.
Compounding the challenge is the fragmented nature of current academic workflows. Researchers must toggle between academic search engines (e.g., Google Scholar, DBLP, ACM Digital Library (DL)), reference managers (e.g., Zotero, EndNote), and writing platforms such as LaTeX, leading to frequent context switching and reduced efficiency [[12]]. Despite their utility, these tools operate in isolation and fail to assist with the cognitive task of integrating retrieved sources into coherent scholarly narratives. As a result, the process of literature synthesis remains largely manual and time-intensive, especially for graduate students and early-career researchers [[9]]. Emerging AI tools offer some assistance by suggesting relevant works or summarizing findings, yet they are often decoupled from the writing interface and lack personalization to the author’s research context.
In this paper, we introduce WriteAssist, a standalone intelligent authoring system designed to streamline and personalize the literature review process. WriteAssist accepts as input the title and abstract of a draft manuscript and performs four integrated functions: it automatically retrieves relevant and recent peer-reviewed literature, highlights passages that support or challenge the author’s claims, incorporates the user’s research profile to personalize recommendations, and presents all results within a unified writing environment. Unlike existing systems that focus solely on retrieval [[12]], WriteAssist closes the loop between search and composition by embedding natural language processing and user modeling directly into the authoring workflow. By reducing context switching and integrating relevant insights seamlessly into the writing process, WriteAssist represents a novel step toward intelligent, context-aware academic writing support.
Related Works
Recent advances in AI have sparked the development of intelligent writing assistants capable of supporting higher-level composition tasks beyond grammar and syntax. While tools like Grammarly and Google’s Smart Compose offer general writing support, academic writing requires deeper assistance with argumentation, content integration, and literature synthesis. Academic phrase banks aid non-native speakers [[1]], and systems like FigurA11y assist with tasks like drafting figure captions [[10]], but they don’t help integrate scholarly content. Citation recommenders suggest references based on draft text [[6]], and LLMs like GPT-4 assist with idea generation and rephrasing [[12]], yet these tools remain separate from the writing process and lack personalized, context-aware support.
Academic search tools such as Google Scholar, Semantic Scholar, and domain-specific databases (e.g., ACM DL, IEEE Xplore) are widely used to locate prior work [[12]], but they require manual query formulation and offer limited support for literature synthesis.
Visualization tools (Connected Papers, Action Science Explorer) reveal citation networks [[3], [5], [7]], and managers (Zotero, Mendeley, EndNote) organize references [[9]], but neither guides selection or integration of citations within the authoring flow.
More recent research has developed citation recommendation systems using semantic similarity to match relevant papers with draft content [[4], [6], [11]], but these prototypes often lack usability and integration with writing tools. WriteAssist advances this work by highlighting excerpts from recommended papers and categorizing them as supportive or contrasting, similar to “smart citation” approaches such as Scite. This helps writers construct stronger arguments and maintain scholarly rigor.
Personalized recommendation systems help scholars find more relevant and interesting papers by using their past work [[12]]. WriteAssist uses the author’s publication and citation history to improve search results, avoid repeated content, and suggest new ideas. Unlike general tools, it gives suggestions during writing and lets users control how results are chosen [[9], [12]]. This makes WriteAssist a useful tool for easier and smarter academic literature reviews.
System Architecture of WriteAssist
The proposed architecture of WriteAssist is shown in Figure [1] representing its main functional components.
Figure 1: WriteAssist system architecture.
WriteAssist is designed as a standalone application or writing interface plugin that unifies literature retrieval, content analysis, personalization, and in-text integration into a cohesive, intelligent authoring workflow. As illustrated in Figure [1], the architecture comprises seven core modules: input processing, retrieval, knowledge graph generator, document analysis, personalization, HyperText generation and assembler. Together, these components form an end-to-end pipeline that allows researchers to move seamlessly from idea to literature-supported writing.
Input Processing
The system begins with Input Processing, where the user submits the title and abstract of a manuscript in progress, denoted as T, via the User Dashboard. This input is processed by an LLM/NLP Analyzer function (Λe). This module is responsible for extracting a set of key research dimensions DT and their corresponding values VT from T. These extracted dimensions and values, representing the core thematic and methodological aspects of the input manuscript, form the foundation for subsequent retrieval and knowledge graph construction phases. The representations aim to capture both lexical cues and semantic nuances of the research topic.
Retrieval Module
Following input processing, the Retrieval Module aims to gather a set of relevant existing publications. For the current implementation of WriteAssist, the primary source corpus P is an indexed version of the DBLP bibliographic database. The module employs a semantic search strategy, denoted as 𝒮, which takes the input manuscript T and the corpus P to return a focused set of N most contextually relevant papers, P_r = 𝒮(T, P). This retrieval internally uses a two-stage approach, involving an initial keyword-based candidate selection followed by semantic re-ranking using embeddings. Filters are applied to ensure retrieved results meet criteria such as being peer-reviewed or falling within a specific publication date range, with considerations for foundational works based on user context or if querying a private, locally maintained article database. The set Pr serves as the initial context for the Knowledge Graph Generator.
Knowledge Graph Generator
Algorithm 1 provides the core logic driving WriteAssist’s Knowledge Graph Generator. The tool builds one small, directed knowledge graph KGd for each dimension d ∈ DT extracted from the draft title and abstract T. For each d ∈ DT, the system checks whether d already appears among the dimensions of papers in Pr. If it does, those papers plus the draft itself form the node set Pd. If not, d is treated as novel: its label or value is embedded via Eembed, and a k-nearest-neighbor search finds matching existing dimensions or values. Finally, each nonempty Pd moves on to graph construction.
In the 𝓔 function (Algorithm 2), we are making a heuristics relationship graph. It first computes embeddings ej for each paper j’s value vj, d corresponding to dimension d. These embedding vectors are then processed by a clustering algorithm Acluster (e.g., DBSCAN) to assign cluster labels {ck}.
\begin{aligned}
e_j &= E_{\text{embed}}(v_{j,d}) \\\n\{c_k\} &= A_{\text{cluster}}(\{e_j\})
\end{aligned}Next, the routine establishes temporal relationships by considering every ordered pair of papers (Pi, Pj) within the node set Pd such that Pi was published before Pj (denoted ti < tj). It computes a confidence weight wij using a function 𝓔. This function assigns the highest weight if both papers share a non-noise cluster and exhibit high lexical overlap; lower weights for lower overlap or noise cluster (N), while intermediate weights can be assigned if an LLM Λr confirms relatedness under other conditions. Pairs with wij > 0 define potential directed edges Pi → Pj with associated weights. This set of weighted potential edges forms an initial directed graph over the nodes Pd. To extract the core structure, a Maximum Spanning Tree (MST) algorithm is applied to this initial graph. The MST identifies the set of edges Ed with the highest possible total weight that connects the nodes within a component without forming cycles. The final result is a compact, interpretable knowledge graph
KG_d = (V_d, E_d)where Vd = Pd. This graph captures thematic clusters (reflected in edge weights) and the temporal evolution of ideas along dimension d. Algorithm 1 governs this per-dimension process, calling Algorithm 2 for the connection steps (embedding, clustering, MST).
Document Analysis
The Document Analysis Module then examines the abstracts and, where available, full texts of candidate articles to extract relevant content. It identifies excerpts that support or contrast the target work through a combination of claim detection, keyword overlap, thematic mapping, and stance analysis and then assembles a new structured document. For example, if a target abstract claims improved accuracy in gesture recognition for AR/VR applications, the module surfaces excerpts from related papers that either report similar performance gains (supporting) or null findings (contrasting). Argumentative relation mining techniques are used to classify each excerpt, with contextual metadata (e.g., section of origin, citation status) retained for interpretability.
Personalization Module
In parallel, the Personalization Module tailors recommendations using a dynamic user profile. This profile includes prior publications, frequently cited works, topical interests, and preferred venues or authors. The module adjusts ranking scores to highlight sources likely unfamiliar yet relevant, balancing personalization with diversity. For instance, a paper outside the user’s primary field but relevant to the abstract’s methods may be up-weighted to broaden perspective. Highlighted excerpts are also re-prioritized based on alignment with the user’s past focus or writing patterns, enhancing the likelihood of meaningful integration.
HyperText Generation
The generated document is then passed to the HyperText Generation module, which interlinks the citations within the text to corresponding entries in a subset of the DBLP knowledge graph. This overlay highlights not only the direct citation relationships but also secondary, lower-level connections, along with the specific portions of text that informed each citation. Users can interactively explore the citation network: hovering over a citation in the text reveals its position within the knowledge graph, while hovering over a node in the graph displays the associated cited passage. This bidirectional interaction allows for intuitive navigation between the document and its underlying scholarly context.
Figure 2: The WriteAssist interface shows the submitted title and abstract (top-left), AI-generated Related Work (bottom-left), and hyperlinked citations (top-right). Hovering over a citation reveals its source article with matching color-coded highlights (bottom-right). Upon acceptance, WriteAssist exports a.tex and.bib file, updating the user’s citation library. This integration supports real-time evidence tracking and enhances scholarly transparency during literature review writing.
To support efficiency and responsiveness, WriteAssist caches intermediate results and pre-computes embedding indices and claim structures for a subset of indexed literature. Real-time NLP operations are minimized where possible using hybrid retrieval and analysis strategies. By integrating these architectural elements, WriteAssist transforms the traditionally disjointed and labor-intensive process of literature review into a context-aware, intelligent authoring experience – bridging scholarly search, content synthesis, and academic writing in a single unified workflow.
Assembler
Finally, the Assembler and User Interface module delivers curated results to the user within the Authoring Tool. As shown in Figure [1], each recommended paper is presented with essential metadata and a highlighted passage that is either supportive or contrasting relative to the user’s work. These excerpts are actionable – users can cite directly, save notes, or view expanded context. The interface allows dynamic updating as the manuscript evolves: When the user modifies the abstract or completes a citation, the system refreshes recommendations to maintain relevance. Buttons for citation insertion, note management, and feedback ensure interactive engagement and iterative refinement of suggestions.
Implementation
WriteAssist is a compact yet capable research authoring tool that merges modern web technologies with large-scale AI functionality. Optimized for scalability, it runs efficiently on cloud infrastructure with optional GPU support for LLM inference.
The front end is a React single-page app styled with Material UI (MUI) for a streamlined user experience, and it uses Plotly.js, a highly customizable JavaScript library that enables interactive visualizations like citation networks and article overlays. On the back-end, Django, a python-based framework manages logic and API endpoints, while MySQL stores structured data. RESTful APIs with async support ensure responsive, efficient, and non-blocking communication between client and server.
For citation analysis, claim extraction, and fairness-aware recommendations, WriteAssist integrates various large language models – including DeepSeek-R1:7B, Gemma3, Mistral-small, and LLaMA3.2 – selected dynamically based on task needs. Secure user authentication provides personalized access and safeguards data across the platform.
Conclusion and Future Directions
WriteAssist introduces a novel paradigm for academic writing by unifying literature retrieval, citation contextualization, and user-specific recommendation into a single, intelligent interface. Unlike conventional tools that separate search from composition, WriteAssist embeds dynamically relevant, context-aware suggestions directly into the writing workflow as “interconnected and color-coded hyperlinks.” This integration not only reduces cognitive effort, but promotes deeper, comparative engagement with scholarly discourse – enhancing both writing quality and pedagogical value.
Grounded in personalization, WriteAssist adapts to users’ research profiles, surfacing foundational works for novices and frontier research for experts. By highlighting both supporting and contrasting evidence, it fosters critical argumentation and offers potential integration into editorial pipelines for citation completeness and review assistance. Its design draws from advances in intelligent writing systems, recommender models, and HCI, with an emphasis on augmenting rather than automating authorship.
Nonetheless, challenges remain. NLP techniques for claim detection and stance classification are imperfect and over-personalization may bias discovery. Ethical safeguards are needed to prevent misuse, such as uncritical copying of system-generated text. Technical hurdles, including limited access to full-text papers and vague user queries, require solutions like hybrid retrieval models, smart caching, and clearer user intent modeling.
Looking forward, future work will focus on building a deployable prototype and conducting user studies to assess trust, efficiency, and academic impact. Expanding personalization beyond publication history to include reading patterns, writing stages, and collaboration dynamics will further improve adaptivity. Enhancing NLP precision and adapting WriteAssist to domain-specific norms; especially in medicine, law, or the humanities; will be essential for broader adoption. As publication volumes grow, tools like WriteAssist have the potential not just to accelerate scholarly writing but to deepen its rigor, relevance, and inclusivity.
Acknowledgments
This research was supported in part by the National Institutes of Health IDeA grant P20GM103408, the National Science Foundation CSSI grant OAC 2410668, the US Department of Energy grant DE-0011014, and the USDA ISAID grant 2020-69012-31871.
References
[1]Micaela Aguiar and Silvia Araujo. 2024. Designing A Phrase Bank For Academic Learning And Teaching: A European Portuguese Case Study. Studia Universitatis Babeș-Bolyai Philologia (Dec. 2024), 169–192. doi:https://doi.org/10.24193/subbphilo.2024.4.08
[2]Lutz Bornmann and Rudiger Mutz. 2015. Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references. Journal of the Association for Information Science and Technology 66, 11 (April 2015), 2215–2222. doi:https://doi.org/10.1002/asi.23329
[3]Jaegul Choo, Hannah Kim, Edward Clarkson, Zhicheng Liu, Changhyun Lee, Fuxin Li, Hanseung Lee, Ramakrishnan Kannan, Charles D. Stolper, John Stasko, and Haesun Park. 2018. VisIRR: A Visual Analytics System for Information Retrieval and Recommendation for Large-Scale Document Data. ACM Transactions on Knowledge Discovery from Data 12, 1 (Jan. 2018), 1–20. doi:https://doi.org/10.1145/3070616
[4]Thi N. Dinh, Phu Pham, Giang L. Nguyen, and Bay Vo. 2024. Enhancing local citation recommendation with recurrent highway networks and SciBERT-based embedding. Expert Systems with Applications 243 (June 2024), 122911. doi:https://doi.org/10.1016/j.eswa.2023.122911
[5]Cody Dunne, Ben Shneiderman, Robert Gove, Judith Klavans, and Bonnie J. Dorr. 2012. Rapid understanding of scientific paper collections: Integrating statistics, text analytics, and visualization. J. Assoc. Inf. Sci. Technol. 63, 12 (2012), 2351–2369. doi:https://doi.org/10.1002/ASI.22652
[6]Travis Ebesu and Yi Fang. 2017. Neural Citation Network for Context-Aware Citation Recommendation. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, Shinjuku, Tokyo, Japan, August 7-11, 2017, Noriko Kando, Tetsuya Sakai, Hideo Joho, Hang Li, Arjen P. de Vries, and Ryen W. White (Eds.). ACM, 1093–1096. doi:https://doi.org/10.1145/3077136.3080730
[7]Robert Gove, Cody Dunne, Ben Shneiderman, Judith Klavans, and Bonnie J. Dorr. 2011. Evaluating visual and statistical exploration of scientific literature networks. In 2011 IEEE Symposium on Visual Languages and Human-Centric Computing, VL/HCC 2011, Pittsburgh, PA, USA, September 18-22, 2011, Gennaro Costagliola, Amy J. Ko, Allen Cypher, Jeffrey Nichols, Christopher Scaffidi, Caitlin Kelleher, and Brad A. Myers (Eds.). IEEE, 217–224. doi:https://doi.org/10.1109/VLHCC.2011.6070403
[8]Anna Małgorzata Kamińska. 2025. Tailoring Scientific Knowledge: How Generative AI Personalizes Academic Reading Experiences. Publications 13, 2 (April 2025), 18. doi:https://doi.org/10.3390/publications13020018
[9]Xiaoqiu Le, Chenyu Mao, Yuanbiao He, Changlei Fu, and Liyuan Xu. 2017. Dpaper: An Authoring Tool for Extractable Digital Papers. J. Data Inf. Sci. 1, 1 (2017), 86–97. doi:https://doi.org/10.20309/JDIS.201607
[10]Nikhil Singh, Lucy Lu Wang, and Jonathan Bragg. 2024. FigurA11y: AI Assistance for Writing Scientific Alt Text. In Proceedings of the 29th International Conference on Intelligent User Interfaces(IUI ’24). ACM, 886–906. doi:https://doi.org/10.1145/3640543.3645212
[11]Deepa Tilwani, Yash Saxena, Ali Mohammadi, Edward Raff, Amit P. Sheth, Srinivasan Parthasarathy, and Manas Gaur. 2024. REASONS: A benchmark for REtrieval and Automated citationS Of scieNtific Sentences using Public and Proprietary LLMs. CoRR abs/2405.02228 (2024). doi:https://doi.org/10.48550/ARXIV.2405.02228 arXiv:https://arXiv.org/abs/2405.02228
[12]Zitong Zhang, Braja Gopal Patra, Ashraf Yaseen, Jie Zhu, Rachit Sabharwal, Kirk Roberts, Tru Cao, and Hulin Wu. 2023. Scholarly recommendation systems: a literature survey. Knowledge and Information Systems 65, 11 (June 2023), 4433–4478. doi:https://doi.org/10.1007/s10115-023-01901-x
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime