SpaceE: Knowledge Graph Embedding by Relational Linear Transformation in the Entity Space
Translation distance based knowledge graph embedding (KGE) methods, such as _TransE_ and _RotatE_, model the relation in knowledge graphs as translation or rotation in the vector space. Both translation and rotation are injective; that is, the translation or rotation of different vectors results in different results. In knowledge graphs, different entities may have a relation with the same entity; for example, many actors starred in one movie. Such a non-injective relation pattern cannot be well modeled by the translation or rotation operations in existing translation distance based KGE methods. To tackle the challenge, we propose a translation distance-based KGE method called **SpaceE** to model relations as linear transformations. The proposed SpaceE embeds both entities and relations in knowledge graphs as matrices and SpaceE naturally models non-injective relations with singular line

SpaceE: Knowledge Graph Embedding by Relational Linear Transformation in the Entity Space


ACM source attribution. Complete text, figures, tables, captions, and references were converted from the authorized ACM archival PDF for 10.1145/3511095.3531284. © 2022 Association for Computing Machinery.

Conversion note. Text blocks follow inspected source coordinates: title/front matter in top-to-bottom, left-to-right order; thereafter each ACM two-column page is read down the left column and then down the right column. The terminal Visual-Meta page is excluded.

Page 1

SpaceE: Knowledge Graph Embedding by Relational Linear Transformation in the Entity Space

Jinxing Yu yujinxing@baidu.com Cognitive Computing Lab, Baidu Research Beijing, No.10 Xibeiwang East Road, Beijing 100193, China

Yunfeng Cai caiyunfeng@baidu.com Cognitive Computing Lab, Baidu Research Beijing, No.10 Xibeiwang East Road, Beijing 100193, China

Mingming Sun sunmingming01@baidu.com Cognitive Computing Lab, Baidu Research Beijing, No.10 Xibeiwang East Road, Beijing 100193, China

Ping Li liping11@baidu.com Cognitive Computing Lab, Baidu Research Bellevue, Washington, 10900 NE 8th ST. Bellevue, Washington 98004, USA

ABSTRACT

Translation distance based knowledge graph embedding (KGE) methods, such as TransE and RotatE, model the relation in knowl- edge graphs as translation or rotation in the vector space. Both translation and rotation are injective; that is, the translation or ro- tation of different vectors results in different results. In knowledge graphs, different entities may have a relation with the same entity; for example, many actors starred in one movie. Such a non-injective relation pattern cannot be well modeled by the translation or rota- tion operations in existing translation distance based KGE methods. To tackle the challenge, we propose a translation distance-based KGE method called SpaceE to model relations as linear transfor- mations. The proposed SpaceE embeds both entities and relations in knowledge graphs as matrices and SpaceE naturally models non-injective relations with singular linear transformations. We theoretically demonstrate that SpaceE is a fully expressive model with the ability to infer multiple desired relation patterns, includ- ing symmetry, skew-symmetry, inversion, Abelian composition, and non-Abelian composition. Experimental results on link pre- diction datasets illustrate that SpaceE substantially outperforms many previous translation distance based knowledge graph em- bedding methods, especially on datasets with many non-injective relations. The code is available based on the PaddlePaddle deep learning platform https://www.paddlepaddle.org.cn/.

CCS CONCEPTS

• Computing methodologies →Reasoning about belief and knowledge.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. HT ’22, June 28-July 1, 2022, Barcelona, Spain © 2022 Association for Computing Machinery. ACM ISBN 978-1-4503-9233-4/22/06...$15.00 https://doi.org/10.1145/3511095.3531284

KEYWORDS

Knowledge Graph Embedding, Link Prediction, Non-Injective Rela- tions

ACM Reference Format: Jinxing Yu, Yunfeng Cai, Mingming Sun, and Ping Li. 2022. SpaceE: Knowl- edge Graph Embedding by Relational Linear Transformation in the Entity Space. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT ’22), June 28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA, 9 pages. https://doi.org/10.1145/3511095.3531284

1 INTRODUCTION

Extracting entities and relationships from web texts [3, 9, 22, 32, 35–37, 42, 44, 47, 49] and reasoning using the extracted informa- tion [8, 25] is the long-term goal of data-mining and natural lan- guage processing. From the philosophy of representation learning, entities and relations (either natural texts or symbols) can be repre- sented by embedding representations, and the reasoning can be per- formed by applying the mathematical operations on these embed- ding vectors. This methodology has been successfully applied in the field of knowledge graph embedding in both symbolic knowledge- bases [6, 38, 46] and open-domain textual fact corpus [12], for the applications of knowledge graph completion [27], question answer- ing [13], topic modeling [20] and so on. The core theoretical problems of this methodology are to find appropriate embedding form and the mathematical operation for reasoning. These problems have been mostly studied in the field of knowledge graph embedding for the task of knowledge graph completion. The background of knowledge graph completion is that: although the large scale knowledge graphs such as Yago [34] and Freebase [5] store vast amounts of fact triples about the re- lationships between entities, knowledge graphs are incomplete and have missing relationships between entities. As the facts in knowledge graphs can be represented by (head,relation,tail), the task of knowledge graph completion is to answer two queries: (head,relation, ?) and (?,relation,tail). Knowledge graph embed- ding (KGE) methods learn embedding representations of the entities and relations and use distance score functions to measure the plau- sibility of candidate fact triples. There are three major types of KGE methods: Translation distance-based methods model the relations as geometric transformations from the head entity vector to the tail entity vector and use norms as the score functions; Bilinear

Page 2

semantic matching methods model the plausibility of fact triples by bilinear semantic matching functions; Deep learning methods use deep neural networks to model the interaction between the head entity, the relation, and the tail entity of a fact triple. The intuition behind the effectiveness of various design choices in KGE methods is to enable reasoning over the knowledge graph and to infer many relation patterns, including symmetry, skew- symmetry, inversion, and composition [38]. Furthermore, non- injective relations are prevalent in many knowledge graphs. For example, 98% of the relations in FB15k-237 dataset are non-injective (1-N, N-1, N-N). Modeling the non-injective relations is challenging, especially when the N-side of the relation is large, because the KGE model has to handle the uncertainty in the knowledge graph completion. Among the translation distance based KGE methods, RotatE [38] is an expressive model. Although it achieves excellent performances, the non-injective relations are still challenging to model for Ro- tatE. RotatE embeds the entities and relations as complex vec- tors, h,r,t ∈Ck, ∥ri ∥= 1. It is expected that t = h ◦r, where ◦denotes Hadamard product (elementwise). For two fact triples (h,r,t1), (h,r,t2) of a 1-N relation r, the embeddings of different tail entities t1,t2 tend to be the same, that is t1 = h ◦r = t2. The reason for this undesired tendency is that rotation in complex vector space is injective. BoxE [1] is a recently proposed state-of-the-art translation dis- tance based KGE method. It represents relations as a set of hyper- rectangles (or boxes). It has the ability to model non-injective rela- tions well by using the box representation. Nonetheless, it cannot explicitly model the compositions of relations [1].

Contributions. In this paper, we present a translation distance based knowledge graph embedding method called SpaceE that mod- els relations as linear transformations in the entity space. Our con- tributions are summarized as follows: 1) We propose SpaceE, a novel knowledge graph embedding method to better model non-injective relations with non-injective linear transformations; 2) We theoreti- cally demonstrate that SpaceE can infer relation patterns, including symmetry, skew-symmetry, inversion, Abelian composition, and non-Abelian composition; 3) We conduct extensive experiments on benchmarks and achieve comparable or better performances than previous translation distance based knowledge graph embedding methods. Experimental results demonstrate our model’s superior- ity in modeling non-injective relations and its capability to infer various relation patterns.

2 RELATED WORK

The task of link prediction in knowledge graphs has been exten- sively studied in the literature, and many methods have been pro- posed. Traditional approaches use rule-based logics [31], or collect path features and use logistic regression on the features [18, 19] for link prediction. Knowledge graph embedding methods later become popular given their simplicity, scalability, and better per- formance than traditional approaches. A few recent studies [30, 48] combine the rule-based Markov Logic Networks (MLNs) and knowl- edge graph embedding methods and achieve promising results. We briefly review some knowledge graph embedding methods and discuss their connections to our work.

Translation Distance Based Methods. TransE [6] represents entities and relations as vectors and models the relation as a transla- tion from the head entity to the tail entity, i.e., vec(head)+vec(relation) = vec(tail). Along the line, TorusE [11] and RotatE [38] are proposed to model relation as translation on a compact Lie group and rotation in a complex vector space, respectively. The relational operations in previous translation distance based methods are injective. The non- injective property of relations is discussed in several extensions of TransE, such as TransH [43] and TransR [21]. They project the vectors of entities into a subspace and then perform relational trans- lation between entities in the subspace. Different entity vectors of non-injective relations could be the same in the subspace. Despite the capability of these methods to model non-injective relations, their overall performance on benchmark datasets lags behind the recent state-of-the-art methods such as RotatE [38] and ConvE [10], and they can only model part of the relation patterns. MQuadE [46] addresses the problem of learning the non-injective relationships using a quadruple matrix representation for fact triples, in which the entity embedding matrices are required to be symmetric. The MQuadE has good theoretical properties and performs well in real- world tasks. Our proposed method - SpaceE has similar theoretical properties as MQuadE, and it does not impose symmetry on the entity embedding matrices (in fact, the entity embedding matrices are not necessarily square), which makes SpaceE more flexible and achieves comparable or even better performance than MQuadE in real-world tasks.

Bilinear Semantic Matching Methods. RESCALE [29] is the first bilinear model that uses matrices to represent relations. Dist- Mult [45] simplifies RESCALE and employs a diagonal matrix for relation modeling. As DistMult uses a symmetric score function, it cannot model skew-symmetry relations. Trouillon et al. [41] pro- posed ComplEx to model skew-symmetry relations.

Deep Learning Methods. Neural networks such as neural ten- sor networks [33] and convolutional neural networks [10] were leveraged for knowledge graph completion. Although they have strong expressiveness, they lack the interpretability of relation reasoning. Moreover, the high performances of some recent neu- ral network methods [26] can be attributed to the inappropriate evaluation protocols [39].

Tensor Decomposition Methods. The knowledge graph triples can be regarded as a 3-order tensor. Canonical Polyadic (CP) tensor decomposition is leveraged in Kazemi and Poole [15], Lacroix et al. [17] for knowledge graph completion. Balazevic et al. [2] proposed TuckER based on Tucker tensor decomposition. TuckER benefits from multi-task learning by sharing parameters through the core tensor.

3 RELATION MODELING BY LINEAR TRANSFORMATION 3.1 Notations

We use bold upper case letters to denote matrices. We denote the identity matrix as I, the Frobenius norm of a matrix A as ∥A∥F , the element-wise product of two matrices A and B as A ⊙B, and the Hadamard product of two complex vectors a and b as a ◦b.

Page 3

3.2 The SpaceE Model

Let E = {e1,e2, · · · ,en} be the set of entities, R = {r1,r2, · · · ,rm} be the set of relations, and T = {(hi,ri,ti)} be the collection of fact triples, where hi ∈E is the head entity, and ti ∈E is the tail entity, ri ∈R is the relation. SpaceE represents entities as p ×q matrices and relations as q ×q square matrices. Let H,R,T denote the matrix of the head entity h, the relation r, and the tail entity t, respectively. SpaceE uses the following score function to measure the plausibility of a fact triple (h,r,t): d(h,r,t) = ∥HR −T ∥2 F . (1)

The idea behind the score function is that the relation between two entities corresponds to a linear transformation of entity matrices. It is expected that HR ≈T when the fact triple (h,r,t) holds, while HR should be far away from T otherwise. Since (h,r,t) is equivalent to (t, ˆr,h) where ˆr is the reverse rela- tion of relationr, we can adopt the reciprocal learning approach [17] and develop following score function:

f (h,r,t) = ∥T ˆR −H ∥2 F , (2)

where ˆR is the representation of ˆr. The score function d(h,r,t) is utilized for head entity prediction while the score function f (h,r,t) is utilized for tail entity prediction during training and inference. As a result, a tail entity may have many head entities when R is singular, and a head entity may have many tail entities when ˆR is singular. In other words, the non-injective property of relation is preserved. See Section 4 for more details.

3.3 Training

The training process samples a mini-batch of fact triples from train- ing data and randomly corrupts the head entities and tail entities to obtain negative samples for head entity prediction and tail en- tity prediction, respectively. It optimizes a loss function to make the scores of true triples lower than that of negative triples. After training, the scores are utilized to rank candidate entities to answer the link prediction queries. We use the self-adversarial negative sampling loss function [24, 38] to learn the parameters:

L = −logσ(γ −d(h,r,t)) −logσ(γ −f (h,r,t))

k Õ

i=1 p(ˆhi,r,t) logσ(d(ˆhi,r,t) −γ)

−

k Õ

i=1 p(h,r, ˆti) logσ(f (h,r, ˆti) −γ),

−

where σ is the sigmoid function, γ is the fixed margin, ˆhi is the i-th sampled negative head entity, and ˆti is the i-th sampled negative tail entity, p(ˆhi,r,t) and p(h,r, ˆti) are the negative sample weights in the self-adversarial loss [38] defined as:

p(ˆhi,r,t) = expα(γ −d(ˆhi,r,t))

Ík j=1 expα(γ −d(ˆhj,r,t)) ,

p(h,r, ˆti) = expα(γ −f (h,r, ˆti))

Ík j=1 expα(γ −f (h,r, ˆtj)) ,

where α is the self-adversarial temperature. The weight correspond- ing to a negative sample is large when the model predicts the nega- tive sample as a true fact triple. The orthogonal constraint can stabilize the training and ease overfitting [4]. Orthogonal linear transformation has the nice prop- erty of preserving the norm. We propose a regularization term to encourage the relation matrix to be nearly orthogonal. Note that orthogonal matrices RT R = I and they are non-singular. We design a regularization term ∥(RT R) ⊙(RT R) −RT R∥2 F to encourage the elements in RT R to be either 1 or 0. The relation matrix satisfying the regularization constraint is nearly orthogonal and could be singular. The regularized loss function is

Lreд = L + λ(∥(RT R) ⊙(RT R) −RT R∥2 F +∥( ˆRT ˆR) ⊙( ˆRT ˆR) −ˆRT ˆR∥2 F ),

where λ is the regularization hyper-parameter.

4 MODEL PROPERTY ANALYSIS 4.1 Infer Relation Properties

Theorem 4.1. SpaceE can model non-injective, symmetric, skew- symmetric, inversion, Abelian composition, and non-Abelian com- position relations with different relation matrices, as summarized in Table 1.

Relation Property Relation Matrix

Non-injective R or ˆR is singular Symmetric R2 = I Skew-symmetric R2 , I r2 inversion of r1 R1R2 = R2R1 = I r3 = r1 ⊗r2 R3 = R1R2 r1 ⊗r2 Abelian R1R2 = R2R1 r1 ⊗r2 Non-Abelian R1R2 , R2R1

Table 1: The ability of SpaceE to model various relation prop- erties with different relation matrices.

Proof. Non-injective Relation. A relation r is non-injective, iff there exist multiple fact triples (h1,r,t), · · · , (hk,r,t), where h1, · · · ,hk are different or multiple fact triples (h,r,t1), · · · , (h,r,tk), where t1, · · · ,tk are different. In SpaceE, when R is singular, the equation XR = T can have different solutions X = H1, · · · , Hk ; when ˆR is singular, the equation Y ˆR = H can have different solu- tions Y = T1, · · · ,Tk .

Symmetric Relation. For a symmetric relation r, for any two entities h,t ∈E, the fact triple (h,r,t) holds ⇐⇒the fact triple (t,r,h) holds. In SpaceE, it requires

HR = T ⇐⇒ TR = H.

It follows from R2 = I.

Skew-symmetric Relation. For a skew-symmetric relation r, for any two entities h,t ∈E, the fact triple (h,r,t) is true =⇒the fact triple (t,r,h) is false. This requires in SpaceE that

HR = T =⇒ TR , H.

Page 4

It follows from R2 , I.

Inversion Relation. A relation r2 is the inversion of relation r1 iff for any two entities h,t ∈E, the fact triple (h,r1,t) is true ⇐⇒ the fact triple (t,r2,h) is true. It requires in SpaceE,

HR1 = T ⇐⇒TR2 = H.

It follows from R1R2 = R2R1 = I.

Relation Composition. A relation r3 = r1 ⊗r2 is the composi- tion of two relations r1 and r2 iff for any three entities a,b,c ∈E, the facts (a,r1,b), (b,r2,c) are true =⇒the fact (a,r3,c) is true. This requires in SpaceE that

AR1 = B, BR2 = C =⇒AR3 = C

It follows from R3 = R1R2.

Abelian Composition and Non-Abelian Composition. Let r3 = r1 ⊗r2 be the composition of the relation r1 and r2, r4 = r2 ⊗r1 be the composition of the relation r2 and r1. In SpaceE, the matrix representations of the two composite relations r3 and r4 can be written as R3 = R1R2, R4 = R2R1.

When R1R2 = R2R1, the composition of r1 and r2 is Abelian; other- wise, the composition is non-Abelian.

4.2 Connection to RotatE

We show that our model SpaceE can be regarded as an extension of RotatE [38].

Theorem 4.2. SpaceE subsumes RotatE with the special block diagonal matrix representations of complex vectors.

Proof. RotatE uses complex vectors to represent entities and re- lations and models the relation between two entities as the rotation in complex vector space. The score function of a fact triple (h,r,t) in RotatE is

s(h,r,t) = ∥h ◦r −t∥,

ri

= 1,

(3)

where h,r,t ∈CK are complex vector representations of head entity, relation, tail entity respectively.

A complex number z = a + bi can be represented by a 2 × 2

matrix a −b b a

 . The product of two complex numbers z1 = a + bi

and z2 = c + di is z1z2 = ac −bd + (bc + ad)i, which corresponds to the product of two matrices: a −b b a

 c −d d c

 = ac −bd −ad −bc bc + ad −bd + ac

 .

A complex vector z = (z1,z2, · · · ,zK) ∈CK can be represented as a 2K × 2K block diagonal matrix Z = diag{Z1,Z2, · · · ,ZK } whose i-th block component Zi is the 2 × 2 matrix representation of zi. Let z,s ∈CK be two complex vectors and Z, S be their corresponding block diagonal matrices respectively. The Hadamard product of z and s is

z ◦s = (z1s1,z2s2, · · · ,zKsK).

It is equivalent to the product of matrices Z and S,

ZS = diag{Z1S1,Z2S2, · · · ,ZKSK }.

As a result, the score function of RotatE corresponds to the score function of SpaceE if we change the complex vector representations of entities and relations to the corresponding block diagonal matrix representations.

4.3 Time and Space Complexity

RotatE and TransE use vectors of dimension n, the time and space complexity is O(n). SpaceE uses p × q and q × q matrices for entity and relation representations, respectively. Its space complexity is O(pq + q2); its time complexity is O(pq2). Assume that there are ne entities and nr relations, the total number of parameters in SpaceE is nepq + nrq2. In experiments, we set p ≈√n,q ≈√n and use the number of parameters not more than RotatE for a fair comparison. For example, in the FB15K dataset, RotatE uses 1000 dimensions complex vectors; it has 20000 parameters of a complex vector; 1000 for the real part, 1000 for the imaginary part. SpaceE uses 45 × 45 = 2025 dimensions embedding matrices. In the Yago3- 10 dataset, RotatE has 500 × 2 = 1000 parameters in a complex vector; SpaceE has 20 × 40 = 800 parameters in an embedding matrix.

5 EXPERIMENTS 5.1 Experimental Setup

Datasets. We conduct experiments on five benchmark datasets: FB15k, WN18, FB15K-237, WN18RR, and YAGO3-10. The statistics of the datasets are summarized in Table 2. FB15k [6] is a subset of FreeBase [5], a large-scale knowledge graph about the general world knowledge. FB15k samples 15k enti- ties and their relations such as /location/country/capital from Free- Base. WN18 [6] is a subset of the WordNet, a lexical database for the English language that groups synonymous words into synsets. WN18 contain relations between words such as hypernym and similarto. FB15k-237 [40] and WN18RR [10] are subsets of FB15k and WN18, respectively, with the inverse relation deleted to resolve the test set leakage problem and to examine the relation composi- tion modeling ability. YAGO3-10 [23] is a subset of YAGO3 whose entities have at least ten relations. Most of the relations in YAGO3-10 are descriptive attributes of people such as wasBornIn, worksAt, and graduatedFrom, which are non-injective. As shown in Table 2, the average number of head per tail of these relations is 913.

Evaluation Metrics and Protocols. We use the mean recipro- cal rank (MRR) and HIT@1, 3, 10 as the metrics to evaluate different models. MRR measures the average of the inverse rank of correct entities in the list predicted by the model. HIT@k measures the average percentage of correct entities that are ranked in the top k by the model. Following Bordes et al. [6], we use the filtered evalu- ation setting where the triplets that appear either in the training, validation, or test set (except the test triplet of interest) are removed from the list of corrupted triplets. To deal with the case that the

Page 5

Dataset #entity #relation #train #valid #test #tphr #hptr

FB15k 14951 1345 483142 50000 59071 9.3 19.7 WN18 40943 18 141442 5000 5000 4.2 4.2 FB15K-237 14541 237 272115 17535 20466 7.9 65.9 WN18RR 40943 11 86835 3034 3134 4.5 2.9 YAGO3-10 123182 37 1079040 5000 5000 3.5 913.1

Table 2: Statistics about the datasets. Among all the triplets of a relation r, let tphr denote the average number of tail entities per head entity, hptr denote the average number of head entities per tail entity. #tphr denotes the average of tphr for all relations, #hptr denotes the average of hptr for all relations.

model may predict the same score for different fact triples, we adopt the random evaluation protocol suggested by Sun et al. [39].

Baselines. We compare our model with representative state of the art models, including distance translation based methods (TransE [6], SE [7], TransH [43], RotatE [38], BoxE [1]), bi-linear se- mantic matching methods (DistMult [45], ComplEx [41], TuckER [2]), and a deep learning method ConvE [10].

Implementation Details. The Adam optimizer [16] is utilized for model training. We tune the hyper-parameters with grid search and select the model that have highest MRR on the validation dataset. The hyper-parameters are selected from following config- urations: the dimension of entity and relation matrices p,q ∈{10, 20, 40, 45}, the self-adversarial temperature α ∈{0.5, 1.0}, the fixed margin γ ∈{3, 6, 9, 12, 24}, the number of negative samples k ∈{256, 512, 1024}, the regularization coefficient λ ∈{0.005, 0.01, 0.05, 0.1, 0.3, 0.6}, the batch size b ∈{512, 1024}, the initial learning rate lr ∈ {1e-4, 2e-4}. The best hyper-parameters on each dataset are given in Table 3. The entity and relation matrix parameters are randomly initialized from the normal distribution N(0, 0.01). Our algorithms are implemented on the PaddlePaddle deep learning framework https://www.paddlepaddle.org.cn/.

FB15k WN18 FB15k-237 WN18RR YAGO3-10

p 45 45 45 10 20 q 45 45 45 40 40 b 1024 512 1024 512 1024 α 1 0.5 1 0.5 1 γ 24 12 9 3 12 k 256 1024 256 1024 256 lr 1e-4 1e-4 1e-4 2e-4 2e-4 λ 0.6 0.05 0.05 0.1 0.005

Table 3: The best hyper-parameters on benchmark datasets. Hyper-parameters p,q,b,α,γ,k,lr, λ, respectively, represent the number of rows of the embedding matrix, the number of columns of the embedding matrix, the batch size, the self- adversarial temperature in negative sampling, the fix mar- gin in the loss function, the number of negative samples, the initial learning rate of the optimizer, the regularization co- efficient.

5.2 Main Results

We highlight the best results on the model category with a bold font and mark the best overall results by underlines in Tables 4, 5, 6. Table 4 summarizes the results on the FB15k and WN18 datasets. The results of TransE are taken from later work’s implementa- tion [28]. The results show that SpaceE can get comparative results on all the evaluation metrics with the state of the art baselines. TransH explicitly models non-injective relations but does not get overall competitive results. As discussed in Dettmers et al. [10], Sun et al. [38], many test triples in the two datasets appear as a recipro- cal form of the training samples, and the test set leakage through inverse relation makes the datasets less challenging. The main rela- tion patterns of the two datasets are symmetry, skew-symmetry, and inversion. The competitive performance of SpaceE demonstrates its capability to model the three relation patterns. The difference between the results of RotatE and SpaceE is minor. This could be explained by our analysis of the connection between SpaceE and RotatE in Section 4.2. From the results in Table 5, we observe that on FB15k-237, SpaceE achieves the best performance among the translation distance based knowledge graph embedding methods on all evaluation metrics. Its performance is competitive with the state-of-the-art semantic matching method TuckER, especially on Hits@10. The main relation pattern in Fb15k-237 is composition. The results demonstrate the effectiveness of SpaceE to model the relation composition. The results in Table 6 show that SpaceE gets the state-of-the-art performance on Hits@1, Hits@3, and Hits@10 metrics on YAGO3-

    The results are promising since YAGO3-10 is the largest dataset

among the three datasets. It contains 123182 entities and more than one million fact triples. Most of the relations in YAGO3-10 are non-injective descriptive attributes of people such as wasBornIn and graduatedFrom. RotatE does not perform well on this dataset due to its limited ability for non-injective relation modeling. The high performances of SpaceE demonstrate its strong capability of non-injective relation modeling. As shown in Table 5, SpaceE gets competitive results with RotatE on WN18RR. RotatE performs well on this dataset, although not well on FB15k-237 and YAGO3-10 datasets. We find two major reasons for this phenomenon. First, the non-injective relations which are difficult for RotatE are not prevalent in WN18RR. #hptr of relations in WN18RR is 2.9. Second, symmetry is prevalent in WN18RR. RotatE can take advantage of its bias - the composition of two symmetric relations is (incorrectly) symmetric [1]. One might wonder whether adding the augmented inverse rela- tions and their embeddings can improve RotatE. We conduct the

Page 6

FB15k WN18

MRR H@1 H@3 H@10 MRR H@1 H@3 H@10

DistMult[❤] 0.798 - - 0.893 0.797 - - 0.946 ComplEx 0.692 0.599 0.759 0.840 0.941 0.936 0.945 0.947

ConvE 0.657 0.558 0.723 0.831 0.943 0.935 0.946 0.956

SE - - - 0.398 - - - 0.805 TransE[■] 0.463 0.297 0.578 0.749 0.495 0.113 0.888 0.943 TransH - - - 0.585 - - - 0.867 RotatE 0.797 0.746 0.830 0.884 0.949 0.944 0.952 0.959

SpaceE 0.791 0.736 0.830 0.883 0.947 0.941 0.951 0.959

Table 4: Results on FB15k and WN18 datasets. Results with suffix [■] and [❤] are taken from Nickel et al. [28] and [14] respectively. Others are obtained from the original papers.

FB15k-237 WN18RR

MRR H@1 H@3 H@10 MRR H@1 H@3 H@10

DistMult 0.24 0.155 0.263 0.419 0.43 0.39 0.44 0.49 ComplEx 0.247 0.158 0.275 0.428 0.44 0.41 0.46 0.51 TuckER 0.358 0.266 0.394 0.544 0.470 0.443 0.482 0.526

ConvE 0.325 0.237 0.356 0.501 0.43 0.40 0.44 0.52

TransE 0.294 - - 0.465 0.226 - - 0.501 RotatE 0.338 0.24 0.375 0.533 0.476 0.428 0.492 0.571

BoxE 0.337 - - 0.538 0.451 - - 0.541 SpaceE 0.351 0.253 0.389 0.544 0.473 0.423 0.496 0.570

Table 5: Results on FB15k-237 and WN18RR datasets. The results of TransE, DistMult, and ConvE are taken from Sun et al. [38], and others are obtained from the original papers. SE [7] and TransH [43] papers do not have results on the two datasets.

YAGO3-10

MRR H@1 H@3 H@10

DistMult 0.340 0.240 0.380 0.540 ComplEx 0.360 0.260 0.40 0.550 TuckER 0.527 0.446 0.576 0.676

ConvE 0.52 0.45 0.56 0.66

RotatE 0.495 0.402 0.550 0.670 BoxE 0.560 - - 0.691 MQuadE 0.536 0.449 0.592 0.689 SpaceE 0.549 0.463 0.604 0.702

Table 6: Results on YAGO3-10 datasets. The results of Dist- Mult is taken from Sun et al. [38]. Others are obtained from the origin papers.

experiments and get almost the same performances as RotatE; for example, the MRR metrics on FB5k-237 and YAGO3-10 are 0.338 and 0.506, respectively.

5.3 Results on Non-injective Relations

We compare our model with RotatE, the state-of-the-art translation distance based method on non-injective relations. Following previ- ous works [38, 43], we study the ability to model the non-injective relations by categorizing relations into 1-to-1, 1-to-N, N-to-1, and N-to-N relations and report the performances of the four relation groups.

The FB15k-237 test set has 74 1-to-1 relation triples, 67 1-to- N relation triples, 1710 N-to-1 relation triples, and 18615 N-to-N relation triples. The Yago3-10 test set has 556 N-to-1 relation triples and 4444 N-to-N relation triples; it does not contain 1-to-1 and 1-to-N relation triples. The results in Table 7 and Table 8 show that SpaceE consistently outperforms RotatE on N-N relations for both head and tail entity prediction, on 1-to-N relations for tail entity prediction, and on N-to-1 relations for head entity prediction. The number of head entities attached to (t, r) (denoted by hptr) or the number of tail entities attached to (h, r) (denoted by tphr) can reflect the non-injective degree of the relation in a fact triple. The Yago3-10 dataset has the largest average number of heads per tail and relations (#hptr) among the five datasets, so we use it to inves- tigate the relationship between the head prediction performance and hptr. There are 18 triples in the test set which contain entities that do not appear in the training set and can not be predicted by any model. We filter them to avoid noises in the investigation. We compare the head prediction performances of RotatE and SpaceE on test triples with different hptr. The test triples are cat- egorized into five groups by their hptr: [0, 10], [11, 50], [51, 100], [101, 1000], and [1001, 100000]. The average number of hptr of the triples in the five groups is 4, 28, 74, 273, and 46678. The number of test triples in the five categories is 987, 913, 692, 2016, and 374. From the results in Table 9, we observe that: 1) SpaceE consis- tently outperforms RotatE in terms of all evaluation metrics on the five hptr triple categories; 2) when hptr ≤10, SpaceE beats RotatE

Page 7

Metric Model 1-1 1-N N-1 N-N

H T H T H T H T

MRR RotatE 1 1 0.885 0.168 0.095 0.847 0.243 0.389

SpaceE 1 1 0.657 0.267 0.152 0.831 0.259 0.41

H@10 RotatE 1 1 0.985 0.239 0.159 0.911 0.441 0.610

SpaceE 1 1 0.985 0.582 0.228 0.905 0.455 0.624

Table 7: Results of RotatE and SpaceE on different relation types of the FB15k-237 dataset. H denotes head prediction perfor- mances, T denotes tail prediction performances.

Metric Model N-to-1 N-to-N

H T H T

MRR RotatE 0.01 0.66 0.38 0.64

SpaceE 0.02 0.65 0.45 0.69

H@10 RotatE 0.04 0.79 0.61 0.79

SpaceE 0.04 0.80 0.65 0.81

Table 8: Head prediction and tail prediction results of Ro- tatE and SpaceE on different relation types of the Yago3-10 dataset. H and T denote head prediction and tail prediction, respectively.

hptr group 1 2 3 4 5

MRR RotatE 0.35 0.45 0.42 0.32 2e-3

SpaceE 0.37 0.52 0.53 0.41 3e-3

H@1 RotatE 0.27 0.34 0.28 0.18 0

SpaceE 0.30 0.42 0.41 0.28 0

H@3 RotatE 0.38 0.53 0.52 0.37 0

SpaceE 0.40 0.59 0.60 0.47 0

H@10 RotatE 0.48 0.64 0.69 0.60 3e-3

SpaceE 0.49 0.67 0.73 0.64 0.01

Table 9: The head entity prediction performances of RotatE and SpaceE on triples with different hptr in the YAGO3-10 dataset.

with more than 0.01 absolute improvement; when 100 ≤hptr ≤ 1000, the improvement of SpaceE over RotatE is more substantial, with about 0.10 absolute improvement on MRR and Hits@1, 3.

5.4 Case Studies on Relation Patterns

The relation matrix representation of SpaceE provides insights into the property of the relation.

Symmetry/skew-symmetry. We plot the contours of |R2| for a relationr. Figure 1a and Figure 1b show that for symmetric relations similarto and verbgroup in the WN18 dataset, the matrix |R2| is similar to the identity matrix. Figure 1c and Figure 1d show that for skew-symmetric relations hypernym and hyponym, the matrix looks different from the identity matrix.

Inversion. We plot the contours of |R1R2| for two relations r1 and r2 that are the inversions of each other. As shown in Figure 1e and Figure 1f, |R1R2| approximates the identity matrix for inversion

relation pairs hypernym ⊗hyponym and haspart ⊗partof in the WN18 dataset. It confirms that SpaceE is able to model the inversion of the relation by the inversion of the relation matrix.

Composition. For three relations for2, winner, for1 in FB15k- 237 dataset, for2 = winner ⊗for1, we denote their relation matrices as R3,R1,R2. Let c for2 be the augmented inverse relation of for2 and ˆR3 be its relation matrix. We plot the contours of | ˆR3R3| and

ˆR3R1R2

in Figure 1g and Figure 1h, respectively. It shows that the

matrices | ˆR3R3| and | ˆR3R1R2| look similar to the identity matrix. The values in the diagonal positions are larger than the values in other positions. Thus, R3 ≈R1R2.

6 CONCLUSION

The general philosophy of representation learning is to automati- cally discover the representations of input data in order to make appropriate decisions. Two fundamental problems in representation learning are (i) the form of representations and (ii) the decision function. The mathematical properties of the representation form and the decision function determine the learning ability of the learning machine.

In this paper, we study the ability of representation learning al- gorithms in knowledge graph reasoning. We show that the rela- tionships in knowledge graphs vary in their logical properties, including injective vs. non-injective, Abelian vs. Non-Abelian, etc. Furthermore, the reasoning procedure in the knowledge graph builds logical connections between relations, such as inversion and composition. These logical properties and connections exert math- ematical constraints on the representation form and the decision function. Improper design of the representation form and the deci- sion function may fail in implementing those logical properties and connections. We demonstrate that the theoretical failures indeed happened in existing methods.

As the solution to implementing all these logical properties and connections, we propose a translation distance-based knowledge graph embedding method called SpaceE using the idea of mod- eling relations as linear transformations in the entity space. We theoretically demonstrate the ability of SpaceE to model various relation properties, including injective, non-injective, symmetry, skew-symmetry, inversion, Abelian composition, and non-Abelian composition. Qualitative case studies show that the property of learned relation matrices can reflect the symmetry, skew-symmetry, inversion, and composition of relations. Experiments on five bench- mark datasets show that our model obtains competitive results on

Page 8

45

45

0.6

0.5

40

40

0.5

35

35

0.4

30

30

0.4

25

25

0.3

20

20

0.3

15

15

0.2

10

10

0.2

0.1

5

5

0.1

10 20 30 40

10 20 30 40

(a) similarto

(b) verbgroup

45

45

0.3

0.3

40

40

0.25

0.25

35

35

30

30

0.2

0.2

25

25

20

20

0.15

0.15

15

15

10

10

0.1

0.1

5

5

0.05

0.05

10 20 30 40

10 20 30 40

(c) hypernym

(d) hyponym

45

45

0.5

0.5

40

40

35

35

0.4

0.4

30

30

25

25

0.3

0.3

20

20

15

15

0.2

0.2

10

10

0.1

0.1

5

5

10 20 30 40

10 20 30 40

(e) hypernym ⊗hyponym

(f) haspart ⊗partof

45

45

40

40

0.8

0.5

35

35

0.7

0.4

30

30

0.6

25

25

0.5

0.3

20

20

0.4

15

15

0.2

0.3

10

10

0.2

5

5

0.1

0.1

10 20 30 40

10 20 30 40

(g) c for2 ⊗for2

(h) c for2 ⊗winner ⊗for1

Figure 1: Contours of relation matrices. for2, winner, for1 represent the relation awardnominations./nominatedfor, win- ners./awardwinner, nominees./nominatedfor, respectively. Fig- ure (1a, 1b): |R2| of symmetric relations; (1c, 1d): |R2| of skew-symmetric relations; (1e, 1f): |R1R2| for two inver- sion relations r1 and r2; (1g): | ˆRR| of the relation for2; (1h):

ˆR3R1R2

,r3 = r1 ⊗r2. Matrices are normalized by their L2

norms for visualization.

all benchmarks. On two datasets, FB15k-237 and Yago3-10, which contain many non-injective relations, SpaceE substantially outper- forms previous translation distance-based KGE methods, especially on highly non-injective relation triples.

Although our method in this paper is proposed and tested only in knowledge graph reasoning, our design can potentially be applied to other fields that involve complex logical properties and connections, such as many problems in causal discovery. It is a long road to find a general representation learning solution for all kinds of logical relations. It is admitted that we have only made a small step forward down the road.

REFERENCES

[1] Ralph Abboud, İsmail İlkan Ceylan, Thomas Lukasiewicz, and Tommaso Salva- tori. 2020. BoxE: A Box Embedding Model for Knowledge Base Completion. In Advances in Neural Information Processing Systems (NeurIPS). virtual.

[2] Ivana Balazevic, Carl Allen, and Timothy M. Hospedales. 2019. TuckER: Ten- sor Factorization for Knowledge Graph Completion. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China, 5184–5193.

[3] Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew Broadhead, and Oren Etzioni. 2007. Open Information Extraction from the Web. In Proceedings of the 20th International Joint Conference on Artificial Intelligence (IJCAI). Hyderabad, India, 2670–2676.

[4] Nitin Bansal, Xiaohan Chen, and Zhangyang Wang. 2018. Can We Gain More from Orthogonality Regularizations in Training Deep Networks?. In Advances in Neural Information Processing Systems (NeurIPS). Montréal, Canada, 4266–4276.

[5] Kurt D. Bollacker, Colin Evans, Praveen K. Paritosh, Tim Sturge, and Jamie Taylor. 2008. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD). Vancouver, Canada, 1247–1250.

[6] Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. In Advances in Neural Information Processing Systems (NIPS). Lake Tahoe, NV, 2787–2795.

[7] Antoine Bordes, Jason Weston, Ronan Collobert, and Yoshua Bengio. 2011. Learn- ing Structured Embeddings of Knowledge Bases. In Proceedings of the Twenty-Fifth AAAI Conference on Artificial Intelligence (AAAI). San Francisco, CA.

[8] Xiaojun Chen, Shengbin Jia, and Yang Xiang. 2020. A review: Knowledge reason- ing over knowledge graph. Expert Syst. Appl. 141 (2020).

[9] Zeyu Dai, Hongliang Fei, and Ping Li. 2019. Coreference Aware Representation Learning for Neural Named Entity Recognition. In Proceedings of the Twenty- Eighth International Joint Conference on Artificial Intelligence (IJCAI). Macao, China, 4946–4953.

[10] Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. 2018. Convolutional 2D Knowledge Graph Embeddings. In Proceedings of the Thirty- Second AAAI Conference on Artificial Intelligence (AAAI). New Orleans, LA, 1811– 1818.

[11] Takuma Ebisu and Ryutaro Ichise. 2018. TorusE: Knowledge Graph Embedding on a Lie Group. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI). New Orleans, LA, 1819–1826.

[12] Swapnil Gupta, Sreyash Kenkre, and Partha P. Talukdar. 2019. CaRe: Open Knowledge Graph Embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China, 378–388.

[13] Xiao Huang, Jingyuan Zhang, Dingcheng Li, and Ping Li. 2019. Knowledge Graph Embedding Based Question Answering. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (WSDM). Melbourne, Australia, 105–113.

[14] Rudolf Kadlec, Ondrej Bajgar, and Jan Kleindienst. 2017. Knowledge Base Comple- tion: Baselines Strike Back. In Proceedings of the 2nd Workshop on Representation Learning for NLP (Rep4NLP@ACL). Vancouver, Canada, 69–74.

[15] Seyed Mehran Kazemi and David Poole. 2018. SimplE Embedding for Link Prediction in Knowledge Graphs. In Advances in Neural Information Processing Systems (NeurIPS). Montréal, Canada, 4289–4300.

[16] Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimiza- tion. In Proceedings of the 3rd International Conference on Learning Representations (ICLR). San Diego, CA.

[17] Timothée Lacroix, Nicolas Usunier, and Guillaume Obozinski. 2018. Canonical Tensor Decomposition for Knowledge Base Completion. In Proceedings of the 35th International Conference on Machine Learning (ICML). Stockholmsmässan, Stockholm, Sweden, 2869–2878.

[18] Ni Lao and William W. Cohen. 2010. Relational retrieval using a combination of path-constrained random walks. Mach. Learn. 81, 1 (2010), 53–67.

[19] Ni Lao, Tom M. Mitchell, and William W. Cohen. 2011. Random Walk Inference and Learning in A Large Scale Knowledge Base. In Proceedings of the 2011 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP). Edinburgh,

Page 9

UK, 529–539.

[20] Dingcheng Li, Siamak Zamani, Jingyuan Zhang, and Ping Li. 2019. Integration of Knowledge Graph Embedding Into Topic Modeling with Hierarchical Dirichlet Process. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT). Minneapolis, MN, 940–950.

[21] Yankai Lin, Zhiyuan Liu, Maosong Sun, Yang Liu, and Xuan Zhu. 2015. Learning Entity and Relation Embeddings for Knowledge Graph Completion. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence (AAAI). Austin, TX, 2181–2187.

[22] Guiliang Liu, Xu Li, Jiakang Wang, Mingming Sun, and Ping Li. 2020. Extracting Knowledge from Web Text with Monte Carlo Tree Search. In Proceedings of the Web Conference (WWW). Taipei, 2585–2591.

[23] Farzaneh Mahdisoltani, Joanna Biega, and Fabian M. Suchanek. 2015. YAGO3: A Knowledge Base from Multilingual Wikipedias. In Proceedings of the Seventh Biennial Conference on Innovative Data Systems Research (CIDR). Asilomar, CA.

[24] Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean.

    Distributed Representations of Words and Phrases and their Composition-

ality. In Advances in Neural Information Processing Systems (NIPS). Lake Tahoe, NV, 3111–3119.

[25] Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019. OpenDi- alKG: Explainable Conversational Reasoning with Attention-based Walks over Knowledge Graphs. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL). Florence, Italy, 845–854.

[26] Deepak Nathani, Jatin Chauhan, Charu Sharma, and Manohar Kaul. 2019. Learn- ing Attention-based Embeddings for Relation Prediction in Knowledge Graphs. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL). Florence, Italy, 4710–4723.

[27] Dat Quoc Nguyen. 2020. A survey of embedding models of entities and rela- tionships for knowledge graph completion. In Proceedings of the Graph-based Methods for Natural Language Processing (TextGraphs). Barcelona, Spain (Online), 1–14.

[28] Maximilian Nickel, Lorenzo Rosasco, and Tomaso A. Poggio. 2016. Holographic Embeddings of Knowledge Graphs. In Proceedings of the Thirtieth AAAI Confer- ence on Artificial Intelligence (AAAI). Phoenix, AZ, 1955–1961.

[29] Maximilian Nickel, Volker Tresp, and Hans-Peter Kriegel. 2011. A Three-Way Model for Collective Learning on Multi-Relational Data. In Proceedings of the 28th International Conference on Machine Learning (ICML). Bellevue, WA, 809–816.

[30] Meng Qu and Jian Tang. 2019. Probabilistic Logic Neural Networks for Reason- ing. In Advances in Neural Information Processing Systems (NeurIPS). Vancouver, Canada, 7710–7720.

[31] Matthew Richardson and Pedro M. Domingos. 2006. Markov logic networks. Mach. Learn. 62, 1-2 (2006), 107–136.

[32] Alisa Smirnova and Philippe Cudré-Mauroux. 2019. Relation Extraction Using Distant Supervision: A Survey. ACM Comput. Surv. 51, 5 (2019), 106:1–106:35.

[33] Richard Socher, Danqi Chen, Christopher D. Manning, and Andrew Y. Ng. 2013. Reasoning With Neural Tensor Networks for Knowledge Base Completion. In Advances in Neural Information Processing Systems (NIPS). Lake Tahoe, NV, 926– 934.

[34] Fabian M. Suchanek, Gjergji Kasneci, and Gerhard Weikum. 2007. Yago: a core of semantic knowledge. In Proceedings of the 16th International Conference on World Wide Web (WWW). Banff, Alberta, Canada, 697–706.

[35] Mingming Sun, Wenyue Hua, Zoey Liu, Xin Wang, Kangjie Zheng, and Ping Li. 2020. A Predicate-Function-Argument Annotation of Natural Language for Open-Domain Information eXpression. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Online, 2140–2150.

[36] Mingming Sun, Xu Li, and Ping Li. 2018. Logician and Orator: Learning from the Duality between Language and Knowledge in Open Domain. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP). Brussels, Belgium, 2119–2130.

[37] Mingming Sun, Xu Li, Xin Wang, Miao Fan, Yue Feng, and Ping Li. 2018. Logician: A Unified End-to-End Neural Approach for Open-Domain Information Extraction. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining (WSDM). Marina Del Rey, CA, 556–564.

[38] Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowl- edge Graph Embedding by Relational Rotation in Complex Space. In Proceedings of the 7th International Conference on Learning Representations (ICLR). New Or- leans, LA.

[39] Zhiqing Sun, Shikhar Vashishth, Soumya Sanyal, Partha P. Talukdar, and Yiming Yang. 2020. A Re-evaluation of Knowledge Graph Completion Methods. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). Online, 5516–5522.

[40] Kristina Toutanova, Danqi Chen, Patrick Pantel, Hoifung Poon, Pallavi Choud- hury, and Michael Gamon. 2015. Representing Text for Joint Embedding of Text and Knowledge Bases. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP). Lisbon, Portugal, 1499–1509.

[41] Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. 2016. Complex Embeddings for Simple Link Prediction. In Proceedings of the 33nd International Conference on Machine Learning (ICML). New York City, NY, 2071–2080.

[42] Xin Wang, , Minlong Peng, Mingming Sun, and Ping Li. 2022. OIE@OIA: an Adaptable and Efficient Open Information Extraction Framework. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). Dublin, Ireland.

[43] Zhen Wang, Jianwen Zhang, Jianlin Feng, and Zheng Chen. 2014. Knowledge Graph Embedding by Translating on Hyperplanes. In Proceedings of the Twenty- Eighth AAAI Conference on Artificial Intelligence (AAAI). Québec City, Canada, 1112–1119.

[44] Vikas Yadav and Steven Bethard. 2018. A Survey on Recent Advances in Named Entity Recognition from Deep Learning models. In Proceedings of the 27th In- ternational Conference on Computational Linguistics (COLING). Santa Fe, NM, 2145–2158.

[45] Bishan Yang, Wen-tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. 2015. Em- bedding Entities and Relations for Learning and Inference in Knowledge Bases. In Proceedings of the 3rd International Conference on Learning Representations (ICLR). San Diego, CA.

[46] Jinxing Yu, Yunfeng Cai, Mingming Sun, and Ping Li. 2021. MQuadE: a Unified Model for Knowledge Fact Embedding. In Proceedings of the Web Conference (WWW). Virtual Event / Ljubljana, Slovenia, 3442–3452.

[47] Jingyuan Zhang, Mingming Sun, Yue Feng, and Ping Li. 2020. Learning Inter- pretable Relationships between Entities, Relations and Concepts via Bayesian Structure Learning on Open Domain Facts. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). Online, 8045–8056.

[48] Yuyu Zhang, Xinshi Chen, Yuan Yang, Arun Ramamurthy, Bo Li, Yuan Qi, and Le Song. 2020. Efficient Probabilistic Logic Reasoning with Graph Neural Networks. In Proceedings of the 8th International Conference on Learning Representations (ICLR). Addis Ababa, Ethiopia.

[49] Yue Zhang, Hongliang Fei, and Ping Li. 2021. ReadsRE: Retrieval-Augmented Distantly Supervised Relation Extraction. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). Virtual Event, Canada, 2257–2262.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime