Text2SceneVR: Generating Hypertexts with VAnnotatoR as a Pre-processing Step for Text2Scene Systems
The automatic generation of digital scenes from texts is a central task of computer science.

Authors: Giuseppe Abrami, Alexander Henlein, Attila Kett, Alexander Mehler

This Seed edition’s formatting was converted from the ACM version of record under the supplied ACM publication authorization.

Abstract

The automatic generation of digital scenes from texts is a central task of computer science. This task requires a kind of text comprehension, the automation of which is tied to the availability of sufficiently large, diverse and deeply annotated data, which is freely available. This paper introduces Text2SceneVR, a system that addresses this bottleneck problem by allowing its users to create a sort of spatial hypertexts in Virtual Reality (VR). We describe Text2SceneVR’s data model, its user interface and a number of problems related to the implicitness of natural language in the manifestation of spatial relations that Text2SceneVR aims to address while trying to remain language independent. Finally, we present a user study with which we evaluated Text2SceneVR.

CCS Concepts

Human-centered computing → Visualization; Applied computing → Annotation; Hypertext / hypermedia creation.

Keywords

Virtual Reality; Spatial Hypertext; VAnnotatoR; Text2Scene; 3D Annotations

ACM Reference Format: Giuseppe Abrami, Alexander Henlein, Attila Kett, and Alexander Mehler. 2020. Text2SceneVR: Generating Hypertexts with VAnnotatoR as a Pre-processing Step for Text2Scene Systems. In Proceedings of the 31st ACM Conference on Hypertext and Social Media (HT ’20), July 13–15, 2020, Virtual Event, USA. ACM, New York, NY, USA, 10 pages. https://doi.org/10.1145/3372923.3404791

1 Introduction

Human information processing is strongly spatial, not only in perception but also in language [45]. It is not a problem for us, for example, to describe scenes or mentally reconstruct scenes from linguistic descriptions. Take the following text (Sample 1): “During my last conference, I stayed in a beautiful hotel room with a

red sofa, dark blue curtains and a breathtaking view of the old town, which was offered to me through my window beside the desk.” Imagining the scene described by this sentence is no problem for us, while computers still have fundamental difficulties in visualizing even elementary aspects of it. In particular, relations that are not directly mentioned, but are part of general knowledge, for example, are particularly difficult for computers to process (e.g. the fact that the curtains being mentioned are probably attached to the window). Although texts can be processed with various linguistic tools (e.g. [5, 26, 35, 49, 65]), entities be recognized and mapped to 3D objects (e.g. [15, 20, 48]), we are still very far from understanding such texts on a similar level as humans. In this paper we introduce Text2SceneVR, an open hypermedia system for generating a special type of spatial hypertext, which aims to generate data for the training of Text2Scene systems [20]. The spatial data available so far usually have no textual description (e.g. SUNCG [74]), and if they do, they are rather sparse and only connect complete rooms with generic statements (e.g. Stanford Text2Scene [19]). There is no assignment between component objects and the corresponding text sections, which would provide additional information for training. In this way, the basic problem of such text understanding systems, namely the lack of sufficiently large, deeply annotated and openly accessible data, is addressed. This data bottleneck problem currently prevents the effective development of systems that automatically map texts to computerbased scene representations. By transferring this annotation problem to virtual reality and thus associating it with 3D, spatial hypertext, we benefit from the extended annotation possibilities offered by such systems (e.g. [76]). Our approach is to specifically address the problem of the implicitness of natural language: for this purpose, we allow users to extend input texts with sentences and text segments which, from their point of view, are related to the input (by being entailed by it) but not explicitly mentioned. Annotation with Text2SceneVR then means to connect segments of the input text with virtual objects or their spatial relations and to do the very same with entailed descriptions. The resulting hypertexts can than be used to train systems that are ideally able to do this themselves. To generate such systems, we distinguish the following relations: (1) Object recognition: The first task concerns the identification of described objects. This requires entity recognition methods that recognize successive descriptions of the same objects, their attributes and relations, in texts.

(2) Referential meaning relations: In order to recognize objects correctly one has to understand the meanings of their linguistic descriptions. This relates to explicit as well as implicit descriptions. Explicit descriptions are usually manifested by definite noun phrases (e.g. the red sofa). Implicit descriptions concern under-specified, possibly contradictory, vague or otherwise informationally uncertain descriptions. When referring, for example, to a hotel room it can be implicitly assumed that a bed is likely contained in it. But this does not need to be mentioned in its description. In any event, it is expected that the result of an automatic processing of scene descriptions itself is not under-specified or too low in content. With Text2SceneVR we introduce a tool for generating annotation data for training systems that automatically interpret under-specified scene descriptions. (3) Part-whole-relations: A related challenge concerns implicit descriptions of part-whole-relations of objects. Depending on the level of detail of the scene description, it may be necessary to additionally refer to components of objects and the materials of which they consist. In the above text, the visualization of the bed may include, for example, references to mattresses, sheets, pillows, quilts, covers, etc. (4) Topological relations: Beyond part-whole-relations we have to consider that objects are topologically arranged - and again, this may not be explicitly expressed (as in “The printer is placed besides the PC” – left or right?). This also refers to spatial distance, perspectivation, scaling and contextual relations (such as the left-right distinction). The following sentence illustrates how scale descriptions can be underspecified: The big ant is on the little elephant. [40] We can assume that the elephant is significantly larger than the ant. That is, the attribute is scaled relative to the referenced entity. But also the pronoun “on” is ambiguous: in the sentence “The ant is on the airplane”, on means “inside” and not “on top”. Based on these preliminaries, we distinguish relations (rows in Table 1) and their arguments (columns in Table 1) in order to specify 16 sub-tasks of Text2Scene systems.

Table 1. Matrix of arguments and relations.

Relation

Relation explicitness

Explicit argument

Implicit argument

Referential meaning

explicit

OR₁

OR₂

Referential meaning

implicit

OR₃

OR₄

Part-whole

explicit

OH₁

OH₂

Part-whole

implicit

OH₃

OH₄

Topological

explicit

OM₁

OM₂

Topological

implicit

OM₃

OM₄

Given the example “I leafed through my newspaper in my garden in the shade of a tree while leaning on its root”, Table 2 lists sentences that are either explicitly or implicitly entailed by this sample and thereby exemplify the cases distinguished by Table 1. Since these cases are usually mixed, it is very difficult to correctly identify object relations expressed in texts and to convert them into scenic representations. This task of automatically generating scenes based on text descriptions is approached under the name Tex2Scene. Only humans are currently capable of solving this task. In order to train Text2Scene systems

appropriately, we need both: sufficiently deep and accurate annotations of texts like Sample 1, but also annotations of the same quality of sentences and texts entailed by such samples in order to get a better grip on the problem of under-specified space and object descriptions. Text2SceneVR is dedicated to this task. Different disciplines have varying views on how to define a scene. In linguistics, the “narrated space” is often divided hierarchically. These range from the lowest level of the “spatial framework” in which the current action takes place to the highest level of the “narrative universe” [22, chap. 2.3], which describes “the world (in the spatio-temporal sense of the term) presented as actual by the text, plus all the counterfactual worlds constructed by characters as beliefs, wishes, fears, speculations, hypothetical thinking, dreams, and fantasies [68, chap. 2.1e]”. For psychology, a scene is much more object-bound. Thus, in experiments, a scene is often understood as a set of all objects included by the corresponding scene (e.g. [30]), or of all objects that are perceived (e.g. [82]) in the scene. Existing scene synthesis systems interpret objects in a similar way to psychologists and describe scenes as a series of objects in space arranged in a certain way [84]. The creation of scenes from text are usually realized in three steps: preprocessing/parsing, optimization/inference and generation [15, 84]. The formalisation is as follows [15]:

𝑃(𝑠|𝑢) = 𝑃(𝑡|𝑢)𝑃(𝑡′|𝑡)𝑃(𝑠|𝑡′) (1)

where u is the original utterance used to generate the scene s, t is the original scene template and t’ is the optimized one. P(t|u) is therefore the parsing phase in which the template is generated from the input, P(t’|t) the interference phase in which the template is optimized and P(s|t’) the generation phase in which the final scene is generated from the template. The recognition of objects and their direct relationships can be processed directly in the parsing phase through extensive NLP pre-processing. Usually the problems arise from the implicit relationships, which can be resolved in the Inference phase. However, interpretations of meaning representations of objects, part-whole relations and contiguity relations mostly depend on the respective context and the availability of

Table 2. Sentences exemplifying referential meaning, topological, and part-whole relations as distinguished by Table 1. All these sentences are entailed by Sample (1). Mentions of objects and relations are identified by o and r, with implicit mentions in bold.

Type

Example

OR₁

“I (o) have (r) a garden (o).”

OR₂

“I lean (r) with my back (o) against the root (o).”

OR₃

“I (o) read (r) the newspaper (o).”

OR₄

“The weather (o) is (r) sunny (o).”

OH₁

“The root (o) is partof (r) the tree (o).”

OH₂

“The tree (o) has (r) branches (o).”

OH₃

“The newspaper (o) consistsof (r) pages (o).”

OH₄

“Leaves (o) hang (r) from the branches (o) of the tree.”

OM₁

“I (o) am in (r) my garden (o).”

OM₂

“The grass (o) beneath me is in (r) the shadow (o) of the tree.”

OM₃

“The newspaper (o) is infrontof (r) me (o).”

OM₄

“The sun (o) is behind (r) the tree (o).”

general knowledge. Many systems try to solve the corresponding bottleneck problem by using knowledge bases like WordNet [57] or ConceptNet [75] (e.g. WordsEye [20] or SceneMaker [32, 33]); alternatively they use data-driven methods (e.g. SceneSeer [15] or Ma et al. [48]). It becomes clear that Tex2Scene is a highly challenging task that requires certain conditions for its implementation: A high degree of machine learning is necessary, which in turn requires the availability of appropriate, deeply annotated training data that contain a variety of under-specified representations of spatial relationships and explicate these as far as possible. Our approach is to generate these training data in the form of spatial hypertexts. Since there is currently no patent solution for this scenario, we extended VAnnotatoR [3, 76] – an open hypermedia system for the visualisation and annotation of graph structures for the representation of natural language texts. More specifically, we added the functionality of visualizing and annotating spatial hypertexts in Virtual Reality (VR). Moreover, the relations described in Table 1 are processed succesively and beside our focus on the generation of training data for machine learning, the interaction of users with objects can be included in later stages. This paper describes an extension of VAnnotatoR as a system for generating spatial hypertexts, the underlying annotation model, its exemplification and evaluation, the so-called Text2- SceneVR and is structured as follows: Section 2 provides an overview of related work. Section 3 describes the annotation environment, the architecture, and the dataset used for evaluation, while the annotation model is described in Section 4. Afterwards, Section 5 presents the evaluation of VAnnotatoR and Section 6 outlines future work. The paper is summarized in Section 7.

2 Related Work

There are some projects that focus on spatial hypertexts but the aspect of three-dimensionality is not considered in our knowledge. For this reason this overview is basically functional and we refer to comprehensive overview articles. Therefore a well overview of these mostly older projects and the observation of the absence of a common vocabulary regarding spatial hypertexts, see [11]. The understanding of spatial hypertexts used to consist more in the visualization of these graph structures. Accordingly, the origins were rather browser-based procedures with the goal of visualizing the underlying networks [52, 79]. An example for this is the tool VIKI [51]. The system supports the reader by its visual representation through the spatial usability of relative nodes as well as the writer by means of a visual language. In the following years, the system was continuously further-developed [52, 53]. The Visual Knowledge Builder (VKB) [72] considered itself a “second generation” of spatial hypertexts. The focus of this system was on long-term cooperation and linking through the introduction of processing histories. Since then, many other applications have been developed and enhanced according to the Spatial Hypertext principle, e.g. Wikis [73], visualization of relevant content [66], use of a document store for data storage [67] or interpretation of spatial ambiguity [25]. Related work also concerns VR-based systems. Since it is impossible to list all relevant VR projects, we highlight selected projects

which focus on dynamic virtual environments. 3D visualizations of software code that enable immersive “flights” by users, where classes are represented as virtual solar systems or as cities is described by [61]. Likewise Kett et al. [42] introduce resources2city, a system for visualizing and interacting with file systems represented as cities, while [83] examine the effects of virtualized architectural structures on users. In addition, [60] introduce Vremiere, a video editing tool designed to break the boundaries of 2D applications by processing and visualizing in 3D environments. There is also a range of earlier projects that address information management and retrieval using 3D environments such as [10, 12, 13]. All these tools have in common that they allow for rich object-related annotations in VR making use of spatial metaphors for information modeling. Our task will be to add the generation of training data for the automatic recognition of spatial structures to this area. Applications in VR are also available in different fields of application such as medicine [43]), psychiatry [9] and learning [59, 69]. The first uses VR to develop immersive therapies for patients with post-traumatic stress disorders. The second investigates the potential of VR in forensic psychiatry. Thirdly, [59] describe the use of 3D environments in teaching autistic children. Although there are many projects of this kind, there is no system that allows to generate training data especially for Text2Scene systems. Text2Scene- VR is being developed to fill this gap. A second field of related work concerns semiotic analyses of VR, which are rather rare (see [8] for a review of this literature). We concentrate on the few articles that focus on an operative, at least classificatory concept of semiotics. [50] provide a semiotic analysis from the point of view of pragmatics and especially rhetoric. [8] extend this approach by considering syntactic, semantic and pragmatic aspects of classifying VR systems. In doing so, they focus primarily on visual, iconic signs. [7] use this classification in a user study of eight VR systems. In contrast to these approaches, we start with an analysis of linguistic signs to enter the field of indexical (hyperlinks) and iconic signs (3D simulations). Furthermore, there are numerous works focusing on the recognition of objects in texts and their transformation into three-dimensional representations. The first successful system was Words- Eye [20]. This has been further developed until today and is one of the linguistically most flexible Text2Scene systems [34] because of the resulting resources like VigNet [21] and SpatialNet [80]. WordsEye is largely based on manually annotated rules for processing input texts. The StanfordText2Scene [17–19] project is based on the Stanford- NLP pipeline [49] and therefore includes a wide range of pre-processing tools. Since the placement of objects is based on statistically learned spatial knowledge and the system enables interaction with the user, StanfordText2Scene learns from user behavior, improving both the selection and placement of objects. In more recent works, the focus continues to be on the realistic representation of rooms and groups of objects, rather than on language analysis (e.g. [48]). In general, most Text2Scene systems lack sufficient linguistic pre-processing or the ability to post-correct generated scenes or rooms manually [34]. There are efforts to learn from human corrections (e.g. [18]) or map certain linguistic expressions to spatial references, such as VigNet [21] (an extension of FrameNet [6]) or SpatialNet [80], but

Figure 1: The software landscape into which VAnnotatoR is embedded: on the left side, several web applications are displayed. This includes TextImager (providing NLP pipelines for automatic text processing), TextAnnotator (enabling the multiuser-based annotation of texts), Wikidition [56] (providing a wiki-based user interface) and VAnnotatoR. In the middle of the diagram, the infrastructure of TextImager is presented. It offers NLP pipelines used by VAnnotatoR via TextAnnotator. TextImager’s backend (see TextImager Service Repository) processes a number of input formats also available via ResourceManager (top right). The DUCC [14] component (right side) serves for the horizontal and vertical distribution of processes to enable the processing of large amounts of text data. Finally, Calamari is a database based on Blazegraph [77] which enables the management of ontological knowledge which is currently extracted from Wikidata, Wikipedia and other resources or areas in which the present architecture has been applied. The colors of the elements allow a clearer differentiation.

these only refer to individual linguistic phenomena and the data is not publicly available. In the meantime, efficient systems have been developed that map linguistic descriptions to images [78, 85] or generate descriptions for images [81]. But even these try to avoid the problems mentioned above by using ever larger neural end-to-end systems, which require even larger large data sets, such as COCO [47] or Conceptual Captations [71]. However, these datasets are not yet available for 3D. More specifically, in the present context, end-to-end learning means that the entire model is differentiable so that it can therefore be trained via gradient descent. Since these models often consist of millions of parameters that are trained via training data, correspondingly large amounts of data are required [27]. Note that end-to-end learning has established itself as a state-of-the-art method in many NLP areas such as coreference resolution [46] or speech recognition [31]).

3 From VAnnotatoR to Generating and Annotating Virtual Rooms

For the generation of spatial hypertext it is necessary to learn topological as well as part-whole relations from texts. For this, the implementation of a system for the generation spatial hypertexts has already been the object of previous work [55], the so-called V- AnnotatoR [76]. VAnnotatoR allows for creating, visualizing and interacting with multimedia data (texts, images, segments of texts and images, geo-coordinates, video and audio files, URLs (by

means of virtual browsers) and 3D models of objects and especially of (virtual reconstructions of) buildings. For this purpose, 3D glasses (HTC Vive[^1] and Oculus Rift[^2]) are used as VR devices, while Google’s ARCore[^3] is used as a platform for Augmented Reality (AR) devices [55]. In addition to visualization and interaction with objects, annotation, i.e. the explicit relation of signs and objects, is an essential feature of VAnnotatoR. This functionality depends on the type of the object: texts and images can be segmented with VAnnotatoR, for example, links are processed with its virtual browser and video files are processed using a virtual viewer. Beyond that, VAnnotatoR includes a variety of methods for the interaction with 3D content: • Highlighting: links between objects can be highlighted by the user to get an overview or to create reminder marks. • Looking ahead: remote objects linked to an object can be visualized with a preview function, especially if they are out of sight in virtual space. This preview serves as a preparation step for what we call teleportation. • Teleportation: in order to bridge the spatial distances between objects, portals can be created which visualize a preview of the target and, when used (selected or entered), provide a virtual transportation to the remote object [55].

1https://www.vive.com/de/product/#viveseries 2https://www.oculus.com/rift/ 3https://developers.google.com/ar

The automatic generation of scenes in VR based on text descriptions makes it necessary to train suitable machine-learning models, which in turn are bound to the sufficient availability of training data. To this end, we implemented and tested an annotation model (see Section 4) which is based on the core technology of VAnnotatoR. The annotation model is exemplified by virtual rooms. Thus, the present paper describes the extension of VAnnotatoR for generating and annotating virtual rooms as a means to generate annotation data for training Text2Scene systems.

3.1 VAnnotatoR’s Core Functionality

VAnnotatoR[^4] was developed as a virtual research platform for the visualization, annotation and processing of multimedia content. VAnnotatoR processes a wide range of content objects: this includes (segments of) texts, images, videos and audio streams as well as 3D representation of buildings or places [55]. Figure 1 shows the software landscape in which VAnnotatoR is embedded (c.f. [3, 42, 44, 76]). By integrating TextAnnotator [2], a platform-independent annotation tool, VAnnotatoR allows for annotating texts on various levels of text structuring [41]. Amongst other things, this includes anaphoric relations, propositional structures, argument structures and rhetorical text structures [2]. TextAnnotator operates on texts using the UIMA [23, 29] format. In this way, external NLP tools can be easily integrated and, conversely, the output of TextAnnotator can be exchanged interoperably [37]. In fact, any UIMA document that is serialized and interchangeable via XML can be processed with TextAnnotator in this way. Thanks to the additional integration of TextImager [35], VAnnotatoR dispenses with the need to manually annotate documents virtually in raw format on all levels. That is, a wide spectrum of language levels is automatically pre-processed and annotated using the NLP pipeline (including tools for tokenization, named entity recognition, relation extraction, semantic role labeling, etc.) of TextImager. In addition, by means of DUCC [14][^5], TextImager allows for processing large amounts of text data in a horizontally and vertically distributed server landscape. TextImager generates UIMA documents, which are managed by the so-called ResourceManager and the UIMA Database Interface (UIMA-DI) [1]. UIMA-DI is a database solution for the document-based approach of UIMA and enables the real-time use of UIMA documents for annotation processes. Annotations of UIMA documents are defined by means of annotation schemes, that is, so-called UIMA Type System Descriptors. With the help of ResourceManager [28] UIMA documents can be given user and group related access rights. In addition to these documents, VAnnotatoR can process a number of other resources, as shown in Figure 2. The communication between VAnnotatoR and TextAnnotator takes place via a web socket. This 1-to-1 connection of both tools enables direct interaction between different users without time-consuming requests for changes [4]. TextAnnotator allows the simultaneous annotation of the same text by several users [4]. To this end, views are generated so that texts can be annotated by different users in a collaborative manner or logically and contextually separated from each other. And since the views are provided

4For videos introducing into VAnnotatoR see https://tinyurl.com/w4jctvv 5Distributed UIMA Cluster Computing

Figure 2: VAnnotatoR gives access to various external resources, whereby its users can choose between different sources: resources can be selected from the local computer, the Stolperwege server [54], the ResourceManager [28], or from ShapeNetSem [70].

with access rights, a very flexible use is guaranteed. In addition, the real-time evaluation of different views of the same documents in terms of the inter-annotator agreement allows their selection for training machine learning systems from a quality perspective [4].

Figure 3: Text pre-processed by TextImager is made accessible to VAnnotatoR via TextAnnotator and displayed in an annotation box. Only one sentence is displayed at a time; however, users can switch between the sentences.

3.2 VAnnotatoR’s Text2SceneVR

Up to this point already existing features of VAnnotatoR were described. The following enhancements of VAnnotatoR is about creating spatial structures based on their textual descriptions (see Table 1) to arrive at annotation data for training Text2Scene systems, each annotation task begins with an input text as illustrated

Figure 4: Virtual rooms are created by first defining their dimensions.

in Figure 3: the text is pre-processed by TextImager, loaded via TextAnnotator and visualized in an annotation box of VAnnotatoR, in which words are separated on the token level. Within the box, tokens can be merged to map multi-word expressions and to relate them to spatial objects created by the user (see Figure 5). After a connection to TextAnnotator has been established, a spatial hypertext can be generated from the input text with VAnnotatoR’s so-called Text2SceneVR. For this purpose, references to the 3D objects created by the user and their contents must be generated from the text and its segments. In addition, textual relations manifested in the text must be mapped to corresponding spatial relations; in other words: spatial configurations must be created that correspond to these textual relations (see Section 1). Finally, the user may generate additional sentences that are entailed by corresponding sentences of the input text from his point of view and process them according to the same procedure. In this way, the relational spectrum described in Section 1 is mapped in such a way that the respective input text and its user-dependent text extensions are interwoven with the user-generated object space and its spatial arrangement. This is what we call a spatial hypertext in VR. To generate such hypertexts, the following operations are available for users: (1) Creating rooms: the first step for creating a spatial hypertext in VR is to create a room. For this purpose there is a menu item in the annotation box that allows to draw the room’s dimensions on the floor as a grid (Figure 4). The dimensions between the corners of the room, which need not be square, are then visualized. The grid spacing can be freely configured, so that a flexible design of rooms is possible. After the outlines of a new room have been defined, it can be configured in detail (Figure 5). (2) Creating windows, doors and use textures: the rooms can be equipped with doors and windows (Figure 6) and also textured (Figure 7). The rooms can be placed and arranged as desired in the virtual environment. It is possible to arrange them next to each other, to connect them and to form room ensembles (Figure 7). (3) Object placement: further functions include the selection and configuration of room contents and their spatial arrangement. As shown in Figure 8, objects as provided by Shape- NetSem [70] can be placed anywhere in the virtual environment. Besides positioning, objects can be scaled, rotated and clustered into organizational groups.

Figure 5: After defining the corners of a room, its walls are created, whose height depends on the user settings.
Figure 6: Freely configurable doors and windows are positioned on the walls of a room. If two rooms are next to each other, it is possible to create a passage between them.
Figure 7: Virtual rooms can be provided with textures that reflect information contained in the underlying text. All textures are taken from 3dtextures.me.
Figure 8: A look at the annotation from our evaluation. The room described in the input text was created, textured, objects were placed and positioned and the multi-word unit “old high-legged chair” was linked to the chair-object (annotation). The representation of the chair does not fully correspond to the object mentioned in the text; this is because a corresponding object does not exist in ShapeNetSem. The blue line visualizes the pointing gesture [44] used to create the annotation; the green line visualizes the annotation of the chair in the room by a segment of the input text.
Figure 9: VAnnotatoR uses a control window for visualizing and annotating spatial structures. It allows for accessing and modifying the input text (document tab), the properties of rooms (room tab) and objects (object tab). With the help of the object tab, doors, windows and other objects are created, scaled or rotated.

An important step in generating spatial hypertexts is the specification of objects, which are explicitly or implicitly mentioned in the texts, and their placement as contents of the previously generated rooms (see object tab). To this end, a wide range of 3D objects and textures are available. Objects are taken from ShapeNet- Sem [70], a sub-project of ShapeNet [16], which includes more than

12 000 3D objects from 270 categories. Each of these objects is annotated with semantic features such as scaling, orientation, estimated weight and volume. Textures are taken from 3dtextures.me[^6]. About 700 textures from 45 main categories and 200 subcategories are available. VAnnotatoR thus has a large number of degrees of freedom for creating rooms and their contents.[^7]

Text2SceneVR is currently being used by students as part of a practical course at the Goethe University Frankfurt. Until now, twelve paragraphs have been annotated and virtual rooms have been created with Text2SceneVR, which forms a corpus of annotated romms. In addition, two students separately used Kafka’s “The Metamorphosis” as a basis for modeling and annotating the corresponding apartment with an average of 45 objects (without walls, windows and doors).

4 Text2SceneVR’s Annotation Model

In order to map spatial objects to linguistic expressions, a data model is required that is flexible, extensible and interoperable by using known formats. For this purpose, we developed a data model, which is largely based on UIMA type system descriptors. The data model is shown in Figure 10. Though we implemented this model by means of ShapeNetSem, it can be extended to include related object models as generated, for example, with VoxML [62].

Figure 10: Text2SceneVR’s data model for the annotation of 3D scenes.

Our data model requires that annotations are selected in the input text and anchored with a so-called RoomObject. Our model does not require that an object specified in this way is a concrete object that can be mapped to ShapeNetSem; rather, it can also be an abstract object that is composed of several sub-objects. In any event, the entire scene described by the input text itself is considered a RoomObject (e.g. a kitchen scene). The walls of a room are saved as a sorted list of nodes and assigned to the corresponding room object as attributes, as shown in Figure 4. This approach enables not only the hierarchical structuring of the scene representation, but also the linking of text segments with groups of objects. This regards, for example, the modeling of quantifiers (e.g. all glasses) or of expressions that denote

6https://3dtextures.me/ 7For the demonstration of the use of VAnnotatoR for the creation of spatial hypertexts see our YouTube videos ( https://tinyurl.com/y87wtveq).

Garden

Figure 11: Example scene according to the example in Table 2. Arrows mark the child/part of relationship and the objects linked to the text are blue. The other objects are derived and the spatial arrangement is based on a 3D representation.

groups of objects (e.g. seating group). Finally, room objects that are not mentioned directly in the input text are classified as parts of object groups (partly with reference to decoration purposes) and assigned to the overall text. A representation of the examples from Table 2 is shown in Figure 11. Implicit referential meaning is resolved by the selection of objects, implicit part-whole relations by child links and implicit topological relations by the spatial placement of objects.

5 Evaluation

We conducted a user study to evaluate Text2SceneVR. Starting from a text sample, the task of the test persons was to create a room, select and place objects within this room and to annotate the objects by assigning them to corresponding text segments (see Figure 8 for a snapshot of this evaluation task). In addition, a UMUX test [24] (Usability Metric for User Experience) was carried out following the annotation task. The UMUX test included the following questions, which had to be answered on a scale of 1-7, where 1 means that one strongly disagrees and 7 that one strongly agrees: (1) Using VAnnotatoR to annotate spatial structures is a frustrating experience. (2) The functions in VAnnotatoR, to annotate spatial structures, meets my requirements. (3) VAnnotatoR is easy to use. (4) The creation of spatial structures in 3D environments are easy to perform with VAnnotatoR. (5) The annotations of text to spatial structures in VAnnotatoR are easy to perform. Each participant created an individually designed room with partly different objects placed in it. This is due to the fact that each test person imagines the room somehow differently, while the text sample is somehow under-specified with regard to all spatial details. This relates, for example, to the choice of objects, their positioning and design. For this reason, a comparison between the individual annotations is only possible to a limited degree. Therefore, an analysis was performed to compare how much time the annotators needed per object to recognize them in the text, select them from the database and afterwards placing them in the room. Because individual participants have created the rooms

in various degrees of detail, only the concrete object creation and placement process is used for comparison. The results are shown in Figure 12, where the participants needed on average 3.2 minutes to create and place an object. Unfortunately two participants could not be evaluated because of problems with motion sickness. With state-of-the-art VR hardware and alternative movement options, this problem could be solved in the future. The participants spent most of their time looking for suitable objects in the database. Since the database did not contain a suitable 3D object for all object descriptions, sometimes long searches were the result and the next suitable object was selected. This shows that the selection process for objects must be optimized. For example, proposals can be generated based on selected tokens in a text and their textual contexts. But as the UMUX test results of Figure 13 show, the participants were mostly satisfied with the usability of VAnnotatoR. The greatest frustration was caused by getting used to the controls and unfamiliarity with VR. But the participants got used to them after annotating 1-2 objects. This also explains the slightly worse results for Question 3. On the other hand, Question 4 and 5, and thus the focus of our tool, were rated best, which speaks for its handling.

Figure 12: The results of the time measurement with 15 participants of which two participants could not be analysed. On average it took each of the 13 participants 3.2 minutes to place each object. Time includes: recognizing objects in the text, searching for suitable objects from the database and finally placing them in the room.
Figure 13: The average results of the UMUX test with 15 participants.

6 Future Work

There is a growing need in computational linguistics to extract spatial and temporal relations from texts [64]. This led to linguistic schemes such as ISOSpace [39, 63], which serve to model spatial relations of the referents of linguistic expressions. We aim to map the spatial model of VAnnotatoR directly to ISOSpace. This will facilitate the recognition and learning of links between expressions and spatial relations. Further, by using SemAF-ISO (Semantic Annotation Framework) [36, chap. 4.2], we plan to map our model to ISOTimeML [38] in order to annotate temporal structures and to connect them with spatial annotations. Furthermore, our model currently only allows the annotation of entire objects or groups of objects by linking them to text segments. To overcome this bottleneck, we plan to integrate PartNet [58]. This will make it possible to annotate components of objects. A transformation of ShapeNetSem into VoxML [62] is desirable as soon as its development has progressed sufficiently and thus far more objects are available than up to now. VoxML is a modeling language that represents semantic knowledge about 3D objects and links it to representations of actions. A final task will be to speed up the process of selecting 3D objects by recommending candidate objects as soon as a token or multi-word expression is selected in the input text.

7 Summary

We introduced Text2SceneVR, a VAnnotatoR-based tool for generating spatial hypertexts that can be used as training data for Text2Scene systems. It uses TextAnnotator to link texts with spatial objects. The resulting hypertexts can be used to train Text2- Scene systems that automatically generate virtual scenes from textual descriptions. In this way we expect to make a significant contribution to solving the bottleneck problem of Text2Scene systems regarding the lack of training data. Based on the analysis of textual manifestations of spatial content, we distinguished three types of object relations and gave examples of when these are explicitly or implicitly expressed linguistically. In this way we distinguished 12 tasks for Text2Scene systems, which we address with Text2Scene- VR. By using existing databases such as ShapeNetSem and 3dtextures, Text2SceneVR already achieves a high degree of freedom in modeling scene descriptions. In an evaluation, we successfully tested the usability of Text2SceneVR. Future work will deal with the integration of ISOSpace, ISOTimeML and PartNet to increase the expressiveness of Text2SceneVR by far. Text2SceneVR will be published in GitHub (https://github.com/ texttechnologylab/VAnnotatoR) under the AGPL license.

Acknowledgements

The support by the Förderfonds Lehre of Goethe University Frankfurt, by the Stiftung Polytechnische Gesellschaft (SPTG) and by the Federal Ministry of Education and Research (BMBF) via the project CEDIFOR (https://www.cedifor.de/) are gratefully acknowledged.

References

[1] Giuseppe Abrami and Alexander Mehler. 2018. A UIMA Database Interface for Managing NLP-related Text Annotations. In Proc. of LREC (LREC 2018). Miyazaki, Japan.

[2] Giuseppe Abrami, Alexander Mehler, Andy Lücking, Elias Rieb, and Philipp Helfrich. 2019. TextAnnotator: A flexible framework for semantic annotations. In Proc. of ISA-15 (Gothenburg, Sweden) (ISA-15).

[3] Giuseppe Abrami, Alexander Mehler, and Christian Spiekermann. 2019. Graphbased Format for Modeling Multimodal Annotations in Virtual Reality by Means of VAnnotatoR. In Proc. of HCI 2019 (Orlando, Florida, USA) (HCII 2019), Constantine Stephanidis and Margherita Antona (Eds.). Springer International Publishing, Cham, 351–358.

[4] Giuseppe Abrami, Manuel Stoeckel, and Alexander Mehler. 2020. TextAnnotator: A UIMA based tool for simultaneous and collaborative annotation of texts. In Proc. of LREC 2020 (Marseille, France) (LREC 2020).

[5] Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual String Embeddings for Sequence Labeling. In Proc. of COLING 2018. 1638–1649.

[6] Collin F Baker, Charles J Fillmore, and John B Lowe. 1998. The berkeley framenet project. In Proc. of COLING 98. ACL, 86–90.

[7] B R Barricelli, A De Bonis, S Di Gaetano, and S Valtolina. 2018. Semiotic Framework for Virtual Reality Usability and UX Evaluation. In Proc. of GHItaly18.

[8] B R Barricelli, D Gadia, A Rizzi, and D L R Marini. 2016. Semiotics of virtual reality as a communication process. Behav Inform Technol 35, 11 (2016), 879– 896.

[9] M Benbouriche, K Nolet, D Trottier, and P Renaud. 2014. Virtual Reality Applications in Forensic Psychiatry. In Proc. of VRIC ’14 (Laval). ACM, New York, 7:1–7:4.

[10] S. Benford, Ch Greenhalgh, and D. Lloyd. 1997. Crowded Collaborative Virtual Environments. In Proc. of CHI 1997 (Atlanta, Georgia, USA). ACM, New York, 59–66.

[11] Mark Bernstein. 2011. Can We Talk about Spatial Hypertext. In Proc. of HT 11 (Eindhoven, The Netherlands) (HT ’11). ACM, New York, NY, USA, 103–112.

[12] S K Card, G G Robertson, and J D Mackinlay. 1991. The Information Visualizer, an Information Workspace. In Proc. of CHI 1991 (New Orleans, USA). ACM, New York, 181–186.

[13] S K Card, G G Robertson, and W York. 1996. The WebBook and the Web Forager: An Information Workspace for the World-Wide Web. In Proc. of CHI 1996.

[14] James R Challenger, Jaroslaw Cwiklik, Louis R Degenaro, Edward A Epstein, and Burn L Lewis. 2016. Distributed UIMA cluster computing (DUCC) facility. US Patent 9,396,031.

[15] Angel X. Chang, Mihail Eric, Manolis Savva, and Christopher D Manning. 2017. SceneSeer: 3D scene design with natural language. arXiv preprint arXiv:1703.00050 (2017).

[16] Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. 2015. ShapeNet: An Information-Rich 3D Model Repository. Technical Report arXiv:1512.03012 [cs.GR]. Stanford University — Princeton University — Toyota Technological Institute at Chicago.

[17] Angel X. Chang, Will Monroe, Manolis Savva, Christopher Potts, and Christopher D. Manning. 2015. Text to 3D Scene Generation with Rich Lexical Grounding. In Proc. of IJCNLP 15. ACL, Beijing, China, 53–62.

[18] Angel X. Chang, Manolis Savva, and Christopher D. Manning. 2014. Interactive Learning of Spatial Knowledge for Text to 3D Scene Generation. In Proc. of ILLVI.

[19] Angel X. Chang, Manolis Savva, and Christopher D Manning. 2014. Learning Spatial Knowledge for Text to 3D Scene Generation. In Proc. of EMNLP 14.

[20] Bob Coyne and Richard Sproat. 2001. WordsEye: an automatic text-to-scene conversion system. In Proc. of SIGGRAPH 01. 487–496.

[21] Robert Eric Coyne, Daniel Bauer, and Owen C Rambow. 2011. Vignet: Grounding language in graphics using frame semantics. (2011).

[22] Katrin Dennerlein. 2009. Narratologie des Raumes. Vol. 22. Walter de Gruyter.

[23] David Ferrucci, Adam Lally, Karin Verspoor, and Eric Nyberg. 2009. Unstructured Information Management Architecture (UIMA) Version 1.0. OASIS Standard. https://docs.oasis-open.org/uima/v1.0/uima-v1.0.html

[24] Kraig Finstad. 2010. The usability metric for user experience. Interacting with Computers 22, 5 (2010), 323–327.

[25] Luis Francisco-Revilla and Frank Shipman. 2005. Parsing and Interpreting Ambiguous Structures in Spatial Hypermedia. In Proc. of HT 05 (Salzburg, Austria) (HT ’05). ACM, New York, NY, USA, 107–116.

[26] Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke S. Zettlemoyer. 2017. AllenNLP: A Deep Semantic Natural Language Processing Platform. arXiv:arXiv:1803.07640

[27] Tobias Glasmachers. 2017. Limits of end-to-end learning. arXiv preprint arXiv:1704.08305 (2017).

[28] Rüdiger Gleim, Alexander Mehler, and Alexandra Ernst. 2012. SOA implementation of the eHumanities Desktop. In Proc. of the Workshop on Service-oriented Architectures (SOAs) for the Humanities: Solutions and Impacts, Digital Humanities 2012, Hamburg, Germany.

[29] T. Götz and O. Suhre. 2004. Design and implementation of the UIMA Common Analysis System. IBM Systems Journal 43, 3 (2004), 476–489.

[30] Michelle R Greene. 2013. Statistics of high-level scene context. Frontiers in psychology 4 (2013), 777.

[31] Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al. 2014. Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567 (2014).

[32] Eva Hanser, Paul Mc Kevitt, Tom Lunney, and Joan Condell. 2009. SceneMaker: automatic visualisation of screenplays. In Proc. of AAAI 09. Springer, 265–272.

[33] Eva Hanser, Paul Mc Kevitt, Tom Lunney, Joan Condell, and Minhua Ma. 2010. SceneMaker: multimodal visualisation of natural language film scripts. In Proc. of KES 2010. Springer, 430–439.

[34] Kaveh Hassani and Won-Sook Lee. 2016. Visualizing natural language descriptions: A survey. ACM Computing Surveys (CSUR) 49, 1 (2016), 1–34.

[35] Wahed Hemati, Tolga Uslu, and Alexander Mehler. 2016. TextImager: a Distributed UIMA-based System for NLP. In Proc. of COLING 2016 System Demonstrations (Osaka, Japan). Federated Conference on Computer Science and Information Systems.

[36] Nancy Ide and James Pustejovsky. 2017. Handbook of linguistic annotation. Springer.

[37] Nancy Ide and Keith Suderman. 2009. Bridging the Gaps: Interoperability for GrAF, GATE, and UIMA. In Proc. of LAW III. ACL, Suntec, Singapore, 27–34.

[38] ISO. 2012. Language resource management — Semantic annotation framework (SemAF) — Part 1: Time and events (SemAF-Time, ISO-TimeML). Standard ISO/IEC TR 24617-1:2012. International Organization for Standardization, Geneva, CH. https://www.iso.org/standard/37331.html

[39] ISO. 2014. Language resource management — Semantic annotation framework (SemAF) — Part 7: Spatial information (ISOspace). Standard ISO/IEC TR 24617- 7:2014. International Organization for Standardization, Geneva, CH. https:// www.iso.org/standard/60779.html

[40] Hans Kamp. 1975. Two Theories about Adjectives. In Formal Semantics of Natural Language, Edward L. Keenan (Ed.). Cambridge University Press, 123–155.

[41] Attila Kett. 2020. text2City: Räumliche Visualisierung textueller Strukturen. Bachelor Thesis. , 61 pages. Goethe University of Frankfurt.

[42] Attila Kett, Giuseppe Abrami, Alexander Mehler, and Christian Spiekermann. 2018. Resources2City Explorer: A System for Generating Interactive Walkable Virtual Cities out of File Systems. In Proc. of UIST 2018 (Berlin, Germany).

[43] B M Kuehn. 2018. Virtual and augmented reality put a twist on medical education. JAMA 319, 8 (2018), 756–758.

[44] Vincent Kühn, Giuseppe Abrami, and Alexander Mehler. 2020. WikiNectVR: A Gesture-based Approach for Interacting in Virtual Reality Based on WikiNect and Gestural Writing. In Proc. of HCII 2020 (Copenhagen, Denmark) (HCII 2020).

[45] George Lakoff. 1987. Women, Fire, and Dangerous Things: What Categories Reveal about the Mind. University of Chicago Press, Chicago.

[46] Kenton Lee, Luheng He, Mike Lewis, and Luke Zettlemoyer. 2017. End-to-end neural coreference resolution. arXiv preprint arXiv:1707.07045 (2017).

[47] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision. Springer, 740–755.

[48] Rui Ma, Akshay Gadi Patil, Matthew Fisher, Manyi Li, Sören Pirk, Binh-Son Hua, Sai-Kit Yeung, Xin Tong, Leonidas Guibas, and Hao Zhang. 2018. Languagedriven synthesis of 3D scenes from scene databases. In SIGGRAPH Asia 2018 Technical Papers. ACM, 212.

[49] Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP Natural Language Processing Toolkit. In Proc. of ACL System Demonstrations. 55–60.

[50] D Marini, R Folgieri, D Gadia, and A Rizzi. 2012. Virtual reality as a communication process. Virtual Reality 16, 3 (2012), 233–241.

[51] Catherine C. Marshall, Frank M. Shipman, and James H. Coombs. 1994. VIKI: Spatial Hypertext Supporting Emergent Structure. In Proc. of ECHT 94 (Edinburgh, Scotland) (ECHT ’94). ACM, New York, NY, USA, 13–23.

[52] Catherine C. Marshall and Frank M. Shipman III. 1995. Spatial hypertext: designing for change. Commun. ACM 38, 8 (1995), 88–97.

[53] Catherine C. Marshall and Frank M. Shipman III. 1997. Spatial hypertext and the practice of information triage. In Proc. of HT 97. 124–133.

[54] Alexander Mehler, Giuseppe Abrami, Steffen Bruendel, Lisa Felder, Thomas Ostertag, and Christian Spiekermann. 2017. Stolperwege: An App for a Digital Public History of the Holocaust. In Proc. of HT 17 (Prague, Czech Republic) (HT ’17). ACM, New York, NY, USA, 319–320. https://doi.org/10.1145/3078714.3078748

[55] Alexander Mehler, Giuseppe Abrami, Christian Spiekermann, and Matthias Jostock. 2018. VAnnotatoR: A Framework for Generating Multimodal Hypertexts. In Proc. HT 2018 (Baltimore, Maryland). ACM, New York, NY, USA.

[56] Alexander Mehler, Benno Wagner, and Rüdiger Gleim. 2016. Wikidition: Towards A Multi-layer Network Model of Intertextuality. In Proc. of DH 2016 (Kraków) (DH 2016). http://dh2016.adho.org/abstracts/250

[57] George A Miller. 1995. WordNet: a lexical database for English. Commun. ACM 38, 11 (1995), 39–41.

[58] Kaichun Mo, Shilin Zhu, Angel X. Chang, Li Yi, Subarna Tripathi, Leonidas J Guibas, and Hao Su. 2019. Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding. In Proc. of CVPR 2019. 909– 918.

[59] C A Naranjo, J S Ortiz, V M Álvarez, J S Sánchez, V M Tamayo, F A Acosta, L E Proaño, and V H Andaluz. 2017. Teaching Process for Children with Autism in Virtual Reality Environments. In Proc. of ICETC 17 (Barcelona, Spain). ACM, New York, 41–45.

[60] C Nguyen, S DiVerdi, A Hertzmann, and F Liu. 2017. Vremiere: In-Headset Virtual Reality Video Editing. In Proc. of CHI 17 (Denver). ACM, New York, 5428– 5438.

[61] R Oberhauser and C Lecon. 2017. Virtual Reality Flythrough of Program Code Structures. In Proc. of VRIC 17. ACM, New York, 10:1–10:4.

[62] James Pustejovsky and Nikhil Krishnaswamy. 2016. VoxML: A visualization modeling language. arXiv preprint arXiv:1610.01508 (2016).

[63] James Pustejovsky, Jessica L Moszkowicz, and Marc Verhagen. 2011. ISO-Space: The annotation of spatial information in language. In Proc. of SIGSEM, Vol. 6. 1–9.

[64] James Pustejovsky, Jessica L Moszkowicz, and Marc Verhagen. 2011. Using ISO- Space for annotating spatial information. In Proc. of the International Conference on Spatial Information Theory.

[65] Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. arXiv:2003.07082 [cs.CL]

[66] Daniel Roßner, Claus Atzenbeck, and Tom Gross. 2019. Visualization of the Relevance: Using Physics Simulations for Encoding Context. In Proc. of HT 19 (Hof, Germany) (HT ’19). ACM, New York, NY, USA, 67–76.

[67] Jessica Rubart. 2019. On Managing Spatial Hypermedia with Document Stores. In Proc. of HUMAN 19 (Hof, Germany) (HUMAN ’19). ACM, New York, NY, USA, 13–18.

[68] Marie-Laure Ryan. 2012. Space. Hühn, Peter et al. (eds.): the living handbook of narratology (2012). http://www.lhn.uni-hamburg.de/article/space view date:12 Feb 2019.

[69] A Z Sampaio, D Rosario, A Gomes, and J Santos. 2013. Virtual reality applied on civil engineering education: Construction activity supported on interactive models. Int. Journal of Engineering Education 29, 6 (2013), 1331–1347.

[70] Manolis Savva, Angel X. Chang, and Pat Hanrahan. 2015. Semantically-Enriched 3D Models for Common-sense Knowledge. CVPR 2015 Workshop on Functionality, Physics, Intentionality and Causality (2015).

[71] Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In Proceedings of ACL.

[72] Frank M. Shipman III., Haowei Hsieh, Preetam Maloor, and J. Michael Moore. 2001. The visual knowledge builder: a second generation spatial hypertext. In Proc. of HT 01. 113–122.

[73] Carlos Solís and Nour Ali. 2008. ShyWiki-A Spatial Hypertext Wiki. In Proc. WikiSym 08 (Porto, Portugal) (WikiSym ’08). ACM, New York, NY, USA, Article 10, 5 pages.

[74] Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. 2017. Semantic Scene Completion from a Single Depth Image. Proc. of CVPR 2017 (2017).

[75] Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. ConceptNet 5.5: An Open Multilingual Graph of General Knowledge. In AAAI Conference on Artificial Intelligence. 4444–4451.

[76] Christian Spiekermann, Giuseppe Abrami, and Alexander Mehler. 2018. VAnnotatoR: a Gesture-driven Annotation Framework for Linguistic and Multimodal Annotation. In Proc. AREA 2018 (Miyazaki, Japan) (AREA).

[77] Systap LLC. 2015. BlazeGraph. https://blazegraph.com/. Accessed: 2020-02-15.

[78] Fuwen Tan, Song Feng, and Vicente Ordonez. 2019. Text2Scene: Generating Compositional Scenes from Textual Descriptions. In Proc. of CVPR 2019.

[79] Manfred Thüring, Jörg M Haake, and Jörg Hannemann. 1991. What’s Eliza doing in the Chinese room? Incoherent hyperdocuments—and how to avoid them. In Proc. HT 91. 161–177.

[80] Morgan Ulinski, Bob Coyne, and Julia Hirschberg. 2019. SpatialNet: A Declarative Resource for Spatial Relations. In Proc. of SpLU and RoboNLP 2019. 61–70.

[81] Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proc. of CVPR 2015. 3156–3164.

[82] Melissa Le-Hoa Võ, Sage EP Boettcher, and Dejan Draschkow. 2019. Reading scenes: How scene grammar guides attention and aids perception in real-world environments. Current opinion in psychology (2019).

[83] K Wolf, M Funk, R Khalil, and P Knierim. 2017. Using Virtual Reality for Prototyping Interactive Architecture. In Proc. of MUM 17. ACM, New York, 457–464.

[84] Song-Hai Zhang, Shao-Kui Zhang, Yuan Liang, and Peter Hall. 2019. A survey of 3D indoor scene synthesis. Computer Science and Technology 34, 3 (2019), 594–608.

[85] C Lawrence Zitnick, Devi Parikh, and Lucy Vanderwende. 2013. Learning the visual interpretation of sentences. In Proc. of IVVC 2013. 1681–1688.

[^1]: https://www.vive.com/de/product/#viveseries [^2]: https://www.oculus.com/rift/ [^3]: https://developers.google.com/ar [^4]: For videos introducing VAnnotatoR, see https://tinyurl.com/w4jctvv [^5]: Distributed UIMA Cluster Computing. [^6]: https://3dtextures.me/ [^7]: For a demonstration of VAnnotatoR for creating spatial hypertexts, see https://tinyurl.com/y87wtveq.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime