Geo-Contextualizaton and Aggregation of Information Resources
This paper explores the addition of spatial features to geographic objects described in textual and photographic information resources and the resulting research possibilities. We briefly describe one of the systems used by the BHMPI to identify geographic objects and the process of adding spatial features from open data providers. This is followed by a discussion of the possibilities of bridging information domains by using the spatial feature and the usefulness of geo-contextualization for named entity normalization. Finally, we explore the usefulness of geo-queries for research in art histo

Geo-Contextualizaton and Aggregation of Information Resources

Authors: Klaus E. Werner

Published in HT '23: 34th ACM Conference on Hypertext and Social Media, Rome, Italy, September 4-8, 2023 · DOI: 10.1145/3603163.3609045 · License: © Copyright held by the owner/author(s).

Abstract

This paper explores the addition of spatial features to geographic objects described in textual and photographic information resources and the resulting research possibilities. We briefly describe one of the systems used by the BHMPI to identify geographic objects and the process of adding spatial features from open data providers. This is followed by a discussion of the possibilities of bridging information domains by using the spatial feature and the usefulness of geo-contextualization for named entity normalization. Finally, we explore the usefulness of geo-queries for research in art history and related fields.

spatial features, spatial context, named entities, knowledge graph

1 GEOGRAPHIC ENTITIES IN TEXTUAL AND PHOTOGRAPHIC DATA

The BHMPI's library and phototeca hold an enormous amount of textual and photographic data. The research library offers researchers more than 300 thousand volumes and the photographic collection more than one million images. I will talk mainly about the library data, but many aspects - especially the use of the GND as an identifier - apply to both collections. Like all German-speaking libraries and photographic collections, the BHMPI uses the GND authority file system, which covers physical persons, corporations, keywords, and geographic entities. GND IDs are manually assigned to titles: monographs, series, or articles. The GND system is also used by the Photographic Collection. Mappings are available to the more multilingual VIAF system, which opens up the European and international context. And fortunately, a good percentage of GND geographic IDs are reflected in Wikidata. The current statistics (2023) show 324,000 geographic GND IDs made available by the DNB, of which 31,000 cover Italy. These are reflected in the actual library items (monographs, articles) of the BHMPI: of the 1.015.000 available titles, almost 25% or 235.000 titles have at least one geographic GND applied, 145.500 of these titles cover Italy 1.

1A few of these (4.900) come with a geofeature, but unfortunately it is a POINT geometry for the administrative borders of municipalities.

Figure or visual from the source paper.

Apart from these titles where a GND ID is applied to an entire book or article, additional fine-grained data will come from the digitization projects of the BHMPI, where 2,500 books about Rome and Naples have already been digitized and 10,000 more books will be added in 2023/24. All the textual data (about 1,300,000 pages) will be subjected to Named Entity Recognition and Normalization (again using the GND, of course) in order to make them usable for research. Similarly, 1,000,000 photographs will be digitized in 2023. There will be a massive amount of data with geospatial relevance, but still without actual spatial context. Only the fusion of digitized textual data and actual geospatial settings will provide the geocontextualization of the digitized information.

2 APPLYING SPATIAL CONTEXT TO GEOGRAPHIC ENTITIES

First, a note about the type of spatial context we want to apply to our data. A balanced approach would be to collect and apply POINT features when the geographic context is one of a broader overview or the queries do not require anything more sophisticated anyway. Say you're interested in the distribution of fountains in a particular area. Or the nationwide distribution of religious services. But the situation changes when you enter a localized context - say, cities, towns, finer-grained territory - and deal with more complex geo-queries. In these cases, the provision of POLYGON data is an absolute necessity. There is more than one way to actually get the spatial extent for named entities. For POINT data, Wikidata itself is the first choice as a provider. While providers like Geonames, the Getty CONA or DARE/Peliagos only provide POINT data for selected locations or selected application scenarios, Wikidata covers almost all geographic features, including individual buildings. Therefore, whenever POINT data is sufficient, a simple SPARQL query to Wikidata will return both data and geo-features in one go. Accuracy varies, but should not be a problem for simple requirements. For any more sophisticated approach, more advanced LINE/ POLYGON geometries are required: we don't want to reduce geographic objects like territorial features, rivers and roads to a POINT geometry. These POLYGON data are provided by OpenStreetMap, but also by various national agencies. OSM data has global coverage and is generally sufficiently detailed. However, due to its nature - it's a community project where data is updated daily by its users - it does not provide a stable ID. It compensates for this deficiency by offering Wikidata properties, which facilitate the combination of data from Wikidata and geo-features from OSM. But POLYGON data is also increasingly provided by national agencies. In Italy, for example, there are excellent offerings from ISTAT and Protezione Civile: The former provides data for municipalities, provinces and regions, the latter provides precise POLYGON data for each and every aggregate building on its territory (DNAS). Both datasets are licensed under a CC BY 4.0 license and can therefore be combined with other data sources and re-published without license restrictions. They don't contain Wikidata relationships, but they do provide stable IDs.

Figure 2: The center of Rome based on geofeatures from DNAS

Once the geographic features are combined with the datasets provided by Wikidata 2, the full range of related information becomes available: the Wikidata property, of course, but also the GND ID, the data from Arachne, ArchInform, Geonames, Getty CONA, Topostext, etc., just to name a few.

3 IMPROVING NAMED ENTITIES NORMALIZATION

The geo-contextualization of objects is in itself the basis for the recognition and normalization of named entities in digitized texts. In fact, the geographic context is absolutely necessary for a successful named entity normalization. Let's take the example of church buildings in Italy. There are more than 250 churches named "Santa Maria delle Neve" in Italy. The correct assessment of which actual church is being described in a given text can only be made by taking into account the geographical context: is the specific text about the city of Bologna? It is Q25057032. With the villages in the mountains of Sirente Velino? Then it is Q83169288. Or with the Reggio Emila area? Then it would be Q106081594. Automatic recognition of geographic context when trying to assign normalized values to recognized named entities - be it a GND ID or a Wikidata property - is still in the early stages of research. However, any successful approach must combine the evaluation of the spatial context of the text and the geo-features of potentially relevant named entities.

4 BRIDGING INFORMATION DOMAINS USING SPATIAL FEATURES

The geocontextualization of named entities also adds a new method for connecting data from different information domains. Consider combining data from information domains such as Wikidata with information from national cadastral surveys. By

2For data grabbing tools like Overpass Turbo (for OSM) or the Wikidata Query Service can be used. For combining data you would use a JOIN query on the data ingested in PostgreSQL.

Figure 2: The center of Rome based on geofeatures from DNAS

Figure 3: Fig.2 Different geofeatures inside a STDWithin query in PostgreSQL

design, they cannot be linked in the traditional way, i.e. by ID-based mapping. But by overlaying the projected geographic features, we can easily establish a 1:1 identity. Here is an example of the Pantheon in Rome, Italy:

• The historic building referenced by the GND and Wikidata, mapped to polygon geometry from OSM, has the GND ID 4115786-2 and the Wikidata property Q99309.

On this polygon geometry, we can aggregate - you'll probably want to use a geo-query like STDWithin or STCOVERS - the following information domains:

• The archaeological remains referenced in the ArcheoSitar system with its detailed polygon geometry: SITAR ID 15019;

• The national dataset of buildings, again with its own polygon geometry: DNAS ID 12058091000011509800;

• The Getty Cultural Objects Names Authority, though only with point geometry: CONA 700000158;

Aggregating data by their common spatial relationship, unlike aggregating data by mapping IDs, is not tied to any authority file. This is a huge advantage because it allows data to be added from completely unrelated information domains: it's enough for records from different information domains to reference the same geospatial feature. It does not even have to be the same geometry like POINT, LINE, POLYGON, and it does not even have to be the exact spatial location: the relationship of the specific objects to be compared can even be defined with an error level, i.e. with a certain (user-defined) distance from each other (STDWithin). Spatially implicit relationships can thus combine unrelated information domains for extended data grabbing and implicit linking without sharing the same identification system.

5 QUERYING GEO-CONTEXTUALIZED OBJECTS AND ARTWORKS

The potential of having geo-contextualized entities is revealed in the range of possible geo-queries: we can query for topological relations like equals, disjoint, intersects, touches, within, contains, overlaps, crosses; for directional relations like south, north or left, right; and finally for distance relations like at, near, far. In all of these scenarios, a special property of geospatial queries comes into play: while queries for IDs or properties are mostly boolean - a given object has a property or it doesn't - geospatial queries are relational: objects are in some distance or other relationship to each other, and that relationship is subject to interpretation. This allows superior flexibility in creating relevant queries. Here are just a few examples of such queries to facilitate cultural heritage research

• Locations with dense clusters of historic palaces and medieval patrician towers (to study the evolution of historic centers)

• Places of worship located at a significant distance from village centers (to study the ritual or spatial parameters for choosing a location)

• Religious buildings other than Roman Catholic and their relationship to specific areas of a city (to study the spatial relationship between different religious communities).

• Landscapes with Greek or Arabic cultural influence (to study building patterns and artistic idiosyncrasies).

We can even combine layers from different GIS and build queries such as

• Churches built on top of Roman or Greek temples (to study possible ritual and artistic traces of preceding cults)

Geoqueries are useful for studying travel literature and guidebooks, which are part of the BHMPI library. We can query

• monuments along a city tour according to historical guides (to study the spatial perception of their urban setting)

• the sequence of important civil/religious/public buildings along historical routes (to study the evolution of specific pilgrim structures).

But we can also relate objects to the natural environment. So we can build queries like

• Places of worship far from villages (to study their spatial relationship to the environment, e.g. rock churches)

• Palaces in a specific landscape situation (to study the resulting spatial-artistic qualities, e.g. along a coastline, at the entrance of a valley).

And by using Digital Elevation Models (DEM) for altimetric data, we can even perform queries such as

• Isolated monasteries at high elevations (to study their explicit panoramic nature over the surrounding villages).

For all these queries, the BHMPI will have the spatial, textual and photographic data.

Figure 3: Fig.2 Different geofeatures inside a STDWithin query in PostgreSQL

6 WEBGIS AND INSTITUTIONAL KNOWLEDGE GRAPH

The most obvious interface for all data (and geo-features) is a simple WebGIS for users who don't need a complex query mechanism. A nice side effect is the ease of use for cross-checking and identifying missing IDs during data collection. Buildings, villages, administrative or cultural areas are displayed in their specific spatial relationship to each other, but also in their relationship to natural features. Such an interface can be developed using the available web map technologies, as has been done successfully by the BHMPI 3 . Instead, an approach based on a knowledge graph based on a GraphDB has a steeper learning curve for first-time users, but much more interaction possibilities. All kinds of queries can be formulated and executed, and institutional cross-queries - for example, between the library and the photography department - can be scripted. Additional data can be easily added, even from external domains and even on-the-fly. In both cases, however, the necessary basic data - the fusion of data from Wikidata and the DNB with the geo-features from OSM or DNAS - must be prepared beforehand in a repository, where it is ingested by PostGreSQL for use in a WebMap or accessed using SPARQL.

3For a webmap built on Mapbox: Spatial Data

Figure 4: WebGIS for the city of L'Aquila, Italy, with information dashboard

Abbreviations

Arachne: iDAI.objects arachne ArcheoSitar: Sistema Informativo Territoriale Archeologico di Roma

ArchInform: International Architecture Database BHMPI: Biblioteca Hertziana - Max-Planck Institute for Art History, Rome

CONA: Cultural Objects Name Authority, Getty Research Institute

DARE: Digital Atlas of the Roman Empire DNAS: Dataset nazionale degli aggregati strutturali italiani DNB: Deutsche Nationalbibliothek GeoNames: GeoNames GND: Gemeinsame Normdatei ISTAT: Istituto Nazionale di Statistica Protezione Civile: Dipartimento della Protezione Civile OSM: OpenStreetMap Peliagos: Peliagos Network ToposText: ToposText Wikidata: Wikidata

Figure 4: WebGIS for the city of L'Aquila, Italy, with information dashboard

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime