Role-Aware Neural Convex Divergence Heads for Asymmetric Representation Learning

arXiv cs.LG Papers

Summary

The paper introduces role-aware neural convex divergence heads that apply source and target role projections before evaluating an input-convex neural Bregman divergence, enabling structured and interpretable asymmetric distance learning for tasks like lexical entailment, sentence entailment, and ontology hierarchy. Experiments show consistent improvements in directional accuracy over plain ICNN-Bregman heads across semantic and ontology benchmarks.

arXiv:2607.01762v1 Announce Type: new Abstract: Many representation learning problems involve directed relations, such as lexical entailment, sentence entailment, ontology hierarchy, and citation links. Standard Euclidean, cosine, and Mahalanobis heads are symmetric, while generic neural scorers can model directionality but provide limited geometric structure. This paper proposes a role-aware neural convex divergence head for asymmetric representation learning. The head applies source- and target-role projections before evaluating an input-convex neural Bregman divergence, yielding a nonnegative structured score in the role-projected space. We characterize its projected-space identity, source-role convexity, directional-gap decomposition, and Hessian-based local curvature. Experiments on lexical, sentence, ontology, and directed graph benchmarks compare symmetric distances, unstructured asymmetric scorers, order/hyperbolic baselines, plain ICNN-Bregman heads, and the proposed role-aware variant. Across ten random seeds on the main semantic and ontology benchmarks, role-aware projections consistently improve directional accuracy over plain ICNN-Bregman heads while preserving zero observed negative divergence rate. The results also identify a boundary case: on large fixed-feature citation prediction, specialized symmetric or hyperbolic baselines remain stronger in ranking accuracy. Overall, the proposed head is best understood as a structured and interpretable plug-in distance module for tasks where directional relations matter.
Original Article
View Cached Full Text

Cached at: 07/03/26, 05:43 AM

# Role-Aware Neural Convex Divergence Heads for Asymmetric Representation Learning
Source: [https://arxiv.org/html/2607.01762](https://arxiv.org/html/2607.01762)
Lu Shen222He Huang and Lu Shen contributed equally to this work\.Yunfeng HuangLi QiSchool of Mathematics and Statistics, Chongqing Technology and Business University, Chongqing, 400067, ChinaChongqing Key Laboratory of Statistical Intelligent Computing and Monitoring, Chongqing Technology and Business University, Chongqing, 400067, ChinaSchool of Food Science, Chongqing Technology and Business University, Chongqing, 400067, ChinaFaculty of Electrical Engineering and Information Technology, TU Dortmund University, Dortmund, Germany

###### Abstract

Many representation learning problems involve directed relations, such as lexical entailment, sentence entailment, ontology hierarchy, and citation links\. Standard Euclidean, cosine, and Mahalanobis heads are symmetric, while generic neural scorers can model directionality but provide limited geometric structure\. This paper proposes a role\-aware neural convex divergence head for asymmetric representation learning\. The head applies source\- and target\-role projections before evaluating an input\-convex neural Bregman divergence, yielding a nonnegative structured score in the role\-projected space\. We characterize its projected\-space identity, source\-role convexity, directional\-gap decomposition, and Hessian\-based local curvature\. Experiments on lexical, sentence, ontology, and directed graph benchmarks compare symmetric distances, unstructured asymmetric scorers, order/hyperbolic baselines, plain ICNN\-Bregman heads, and the proposed role\-aware variant\. Across ten random seeds on the main semantic and ontology benchmarks, role\-aware projections consistently improve directional accuracy over plain ICNN\-Bregman heads while preserving zero observed negative divergence rate\. The results also identify a boundary case: on large fixed\-feature citation prediction, specialized symmetric or hyperbolic baselines remain stronger in ranking accuracy\. Overall, the proposed head is best understood as a structured and interpretable plug\-in distance module for tasks where directional relations matter\.

###### keywords:

asymmetric metric learning , Bregman divergence , input\-convex neural networks , role\-aware projection , directed representation learning , interpretability

††journal:Neurocomputing## 1Introduction

Distance functions are a central component of representation learning\. They are used to compare examples in metric learning, retrieve nearest neighbors, score candidate links, classify examples by prototypes, and regularize embedding spaces\. A large fraction of practical systems still rely on symmetric distances such as Euclidean distance, cosine distance, or Mahalanobis distance\. These choices are effective when the relation of interest is similarity\. However, many relations studied in machine learning are not symmetric\. Hypernymy is directed from a specific concept to a more general concept; entailment is directed from a premise to a hypothesis; ontology edges are directed from child terms to parent terms; and citation links point from a citing paper to a cited paper\.

The importance of asymmetry in similarity judgments is not new: Tversky’s feature\-based theory showed that similarity can be directional and context\-dependent, challenging purely metric views of comparison\[[22](https://arxiv.org/html/2607.01762#bib.bib28)\]\. In representation learning, one response is to abandon distance structure and use a generic neural scorer, such as a multilayer perceptron on concatenated pair features\. This provides expressive asymmetry but weakens the geometric meaning of the learned score\. Another response is to use structured asymmetric geometry, including order embeddings\[[23](https://arxiv.org/html/2607.01762#bib.bib29)\]or hyperbolic embeddings\[[15](https://arxiv.org/html/2607.01762#bib.bib20)\]\. These approaches are well matched to hierarchical data, but their assumptions can be restrictive and they are usually implemented as a complete embedding model rather than as a plug\-in head for arbitrary encoders\.

This paper studies a middle ground\. We ask whether an asymmetric head can be both neural and geometrically structured\. Our starting point is the Bregman divergence, originally introduced in convex programming\[[6](https://arxiv.org/html/2607.01762#bib.bib3)\],

Dϕ​\(x,y\)=ϕ​\(x\)−ϕ​\(y\)−∇ϕ​\(y\)⊤​\(x−y\),D\_\{\\phi\}\(x,y\)=\\phi\(x\)\-\\phi\(y\)\-\\nabla\\phi\(y\)^\{\\top\}\(x\-y\),\(1\)whereϕ\\phiis a differentiable convex potential\. Bregman divergences are generally asymmetric and nonnegative, and they include many classical divergences as special cases\[[2](https://arxiv.org/html/2607.01762#bib.bib4)\]\. Their properties are grounded in classical convex analysis and convex optimization\[[18](https://arxiv.org/html/2607.01762#bib.bib23),[5](https://arxiv.org/html/2607.01762#bib.bib5)\]\. Recent work has explored neural parameterizations of Bregman divergences and convex potentials\[[1](https://arxiv.org/html/2607.01762#bib.bib1),[8](https://arxiv.org/html/2607.01762#bib.bib8),[12](https://arxiv.org/html/2607.01762#bib.bib17)\]\. We build on this line of work, but target a specific problem: using neural convex divergences as plug\-in distance heads for asymmetric representation learning\.

The proposed method, called the role\-aware neural convex divergence head, applies two learnable role projections before computing a neural Bregman divergence:

Dϕ,Ps,Pt​\(x,y\)=Dϕ​\(Ps​x,Pt​y\),D\_\{\\phi,P\_\{s\},P\_\{t\}\}\(x,y\)=D\_\{\\phi\}\(P\_\{s\}x,P\_\{t\}y\),\(2\)wherePsP\_\{s\}andPtP\_\{t\}map an input embedding into source and target roles\. This construction turns a known asymmetric divergence into a practical, role\-aware, encoder\-agnostic head\. The contribution lies in the combination of a plug\-in ICNN potential, role\-specific projections, directed training losses, and interpretation tools for asymmetric pairwise learning\.

This work makes three main contributions\.

First, we propose a role\-aware neural convex divergence head for asymmetric representation learning\. The head can be attached to fixed embeddings or learned encoders and returns a nonnegative Bregman divergence in a source\-target projected space\.

Second, we provide a theoretical and diagnostic characterization of the head\. We show how classical Bregman properties are retained after role projection, derive a quadratic special\-case decomposition of the directional gap, and connect this decomposition to Hessian\-based local curvature analysis\.

Third, we conduct a systematic empirical study across lexical entailment, sentence entailment, ontology hierarchy, and directed citation tasks\. The experiments compare symmetric distances, unstructured asymmetric scorers, order/hyperbolic baselines, plain ICNN\-Bregman heads, and the proposed role\-aware variant, with ablations, significance tests, and interpretability diagnostics\.

## 2Related Work

### 2\.1Metric learning and symmetric distance heads

Classical metric learning methods learn distances that preserve neighborhood or constraint structure, including large\-margin nearest\-neighbor learning\[[25](https://arxiv.org/html/2607.01762#bib.bib31)\]and information\-theoretic metric learning\[[9](https://arxiv.org/html/2607.01762#bib.bib12)\]\. The broader literature is reviewed byKulis \[[11](https://arxiv.org/html/2607.01762#bib.bib16)\], who emphasizes metric learning as the problem of adapting the geometry of a representation space to task\-specific comparison constraints\. In modern deep learning pipelines, learned encoders are often paired with Euclidean, cosine, or Mahalanobis heads\. These heads are stable, efficient, and interpretable for similarity, but they satisfyd​\(x,y\)=d​\(y,x\)d\(x,y\)=d\(y,x\)and therefore collapse direction\-specific information by design\.

### 2\.2Asymmetric representation learning

Directed relations have motivated specialized embedding geometries\. Order embeddings model partial\-order relations through coordinate\-wise inequalities and have been used for visual\-semantic hierarchy and entailment\[[23](https://arxiv.org/html/2607.01762#bib.bib29)\]\. Hyperbolic and Poincare embeddings exploit negative curvature to represent trees and hierarchies compactly\[[15](https://arxiv.org/html/2607.01762#bib.bib20)\]\. Knowledge\-graph models such as TransE and RotatE provide additional examples of structured directional scoring for relational data\[[3](https://arxiv.org/html/2607.01762#bib.bib7),[20](https://arxiv.org/html/2607.01762#bib.bib27)\]\. These models are powerful, but they often impose task\-specific geometry or relation\-specific scoring assumptions\.

### 2\.3Input\-convex networks and neural Bregman divergences

Input\-convex neural networks \(ICNNs\) parameterize convex scalar functions by constraining selected weights to be nonnegative\[[1](https://arxiv.org/html/2607.01762#bib.bib1)\]\. Convex neural potentials connect to a longer tradition of convex analysis, Bregman distances, and projection methods\[[6](https://arxiv.org/html/2607.01762#bib.bib3),[7](https://arxiv.org/html/2607.01762#bib.bib9),[18](https://arxiv.org/html/2607.01762#bib.bib23)\]\. Deep divergence learning and neural Bregman divergence models show that neural potentials can learn flexible dissimilarities while retaining structure inherited from convexity\[[8](https://arxiv.org/html/2607.01762#bib.bib8),[19](https://arxiv.org/html/2607.01762#bib.bib26),[12](https://arxiv.org/html/2607.01762#bib.bib17)\]\. Our work differs in emphasis: we treat the divergence as a reusable head for directed representation learning and evaluate whether role\-specific projections improve directional behavior without abandoning Bregman structure\.

### 2\.4Benchmarks for directed semantic and graph relations

HyperLex evaluates graded lexical entailment and hypernymy\[[24](https://arxiv.org/html/2607.01762#bib.bib30)\]\. WordNet provides large lexical taxonomies\[[14](https://arxiv.org/html/2607.01762#bib.bib19)\]\. SICK and SNLI are standard sentence\-pair entailment resources\[[13](https://arxiv.org/html/2607.01762#bib.bib18),[4](https://arxiv.org/html/2607.01762#bib.bib6)\]\. Gene Ontology provides directed biological concept relations\[[21](https://arxiv.org/html/2607.01762#bib.bib14)\]\. OGB link\-prediction datasets evaluate directed graph prediction under larger\-scale fixed\-feature regimes\[[10](https://arxiv.org/html/2607.01762#bib.bib15)\]\. These resources allow us to test whether the proposed head is useful beyond one narrow dataset\.

## 3Method

### 3\.1Problem setting

Letx,y∈ℝdx,y\\in\\mathbb\{R\}^\{d\}be two embeddings produced by an encoder or directly provided as fixed input features\. A directed pair\(x,y\)\(x,y\)indicates thatxxshould be related toyyin a source\-to\-target direction\. The datasets used in this paper do not provide ground\-truth divergence valuesD​\(x,y\)D\(x,y\)\. Instead, they provide directed relation labels or graded relation scores, such as child\-to\-parent ontology edges, premise\-to\-hypothesis entailment labels, or citing\-to\-cited paper links\. The divergence is therefore a learned scoring function\. The goal is to learn a scores​\(x,y\)s\(x,y\)or distance\-like quantityD​\(x,y\)D\(x,y\)such that positive directed pairs have smaller divergence than corrupted negative pairs, and such that the forward direction can be distinguished from the reverse direction when the task requires it\.

The head is designed to be plug\-in in two senses\. Mathematically, it only consumes embeddings and returns a pairwise divergence, so it can be placed after different encoders\. In software, it is implemented as a module with the same interface as ordinary distance heads: given two batches of embeddings, it returns one scalar per pair\.

### 3\.2Input\-convex potential

We parameterize a differentiable strongly convex potentialϕθ:ℝk→ℝ\\phi\_\{\\theta\}:\\mathbb\{R\}^\{k\}\\rightarrow\\mathbb\{R\}using an ICNN\[[1](https://arxiv.org/html/2607.01762#bib.bib1)\]\. A typical layer has the form

zℓ\+1=σ​\(Wℓ\(z\)​zℓ\+Wℓ\(u\)​u\+bℓ\),z\_\{\\ell\+1\}=\\sigma\(W\_\{\\ell\}^\{\(z\)\}z\_\{\\ell\}\+W\_\{\\ell\}^\{\(u\)\}u\+b\_\{\\ell\}\),\(3\)whereuuis the input,σ\\sigmais a convex nondecreasing activation, andWℓ\(z\)W\_\{\\ell\}^\{\(z\)\}is constrained to be element\-wise nonnegative\. To improve numerical stability and ensure strong convexity, we use a quadratic term:

ϕθ​\(u\)=gθ​\(u\)\+λ2​∥u∥22,λ\>0,\\phi\_\{\\theta\}\(u\)=g\_\{\\theta\}\(u\)\+\\frac\{\\lambda\}\{2\}\\lVert u\\rVert\_\{2\}^\{2\},\\quad\\lambda\>0,\(4\)wheregθg\_\{\\theta\}is the ICNN output\.

The induced Bregman divergence is

Dϕθ​\(u,v\)=ϕθ​\(u\)−ϕθ​\(v\)−∇ϕθ​\(v\)⊤​\(u−v\)\.D\_\{\\phi\_\{\\theta\}\}\(u,v\)=\\phi\_\{\\theta\}\(u\)\-\\phi\_\{\\theta\}\(v\)\-\\nabla\\phi\_\{\\theta\}\(v\)^\{\\top\}\(u\-v\)\.\(5\)This quantity is nonnegative whenϕθ\\phi\_\{\\theta\}is convex, and it is generally asymmetric\.

### 3\.3Role\-aware projected divergence

Plain Bregman divergence is asymmetric, but in practice the learned direction may be weak if both arguments are passed through the same representation role\. We therefore introduce role projections:

u=Ps​x,v=Pt​y,u=P\_\{s\}x,\\quad v=P\_\{t\}y,\(6\)wherePsP\_\{s\}andPtP\_\{t\}are learnable source and target maps\. The proposed head is

Dθ,Ps,Pt​\(x,y\)=Dϕθ​\(Ps​x,Pt​y\)\.D\_\{\\theta,P\_\{s\},P\_\{t\}\}\(x,y\)=D\_\{\\phi\_\{\\theta\}\}\(P\_\{s\}x,P\_\{t\}y\)\.\(7\)The projections can be linear maps, shallow MLPs, or constrained affine maps\. The experiments mainly use linear maps because they give the clearest interpretation: the model learns which directions of the embedding space matter when an item acts as a source and which directions matter when it acts as a target\.

Figure[1](https://arxiv.org/html/2607.01762#S3.F1)summarizes the system\-level role of the proposed head\. A pair of items is encoded into two embeddings, the proposed head scores the ordered pair, and the resulting divergence can be used for ranking, link prediction, or interpretation\. In the experiments, this scoring module is trained through triplet\-style pairwise ranking, described below\. Although the reported experiments use fixed embeddings to isolate the head, the module itself has the same input\-output interface as ordinary distance heads and can be placed after a learned encoder\.

directed pairencoderembeddingsxxyyrole\-awareplug\-in headDDtrainexplainzoomed view of the proposed headzxz\_\{x\}zyz\_\{y\}PsP\_\{s\}PtP\_\{t\}source roletarget roleuuvvinput\-convex potentialDDlossdiagnosisThe same head can be attached after fixed embeddings or learned encoders\.

Figure 1:Neural architecture view of the proposed role\-aware plug\-in divergence head\. A directed pair is encoded into two embeddings, the proposed head applies source\- and target\-role projections, and an ICNN\-induced Bregman divergence returns an ordered\-pair score\. The score is used for pairwise ranking losses and for geometric diagnostics such as directional gaps and local curvature\.
### 3\.4Training objectives

For directed link prediction and pairwise relation learning, the experiments use a triplet\-style pairwise ranking objective, following the large\-margin metric\-learning view that supervision can be given as relative comparison constraints rather than absolute distances\[[25](https://arxiv.org/html/2607.01762#bib.bib31),[11](https://arxiv.org/html/2607.01762#bib.bib16)\]\. Each observed directed relation\(x,y\+\)\(x,y^\{\+\}\)is treated as a positive source\-target pair\. During training, one corrupted targety−y^\{\-\}is sampled for each positive pair from a candidate pool that excludes annotated positive targets of the same source\. This negative\-sampling protocol is also common in embedding\-based link prediction, where models are trained to score observed edges above corrupted edges\[[3](https://arxiv.org/html/2607.01762#bib.bib7)\]\. We then minimize a margin objective:

ℒr​a​n​k=max⁡\{0,m\+D​\(x,y\+\)−D​\(x,y−\)\},\\mathcal\{L\}\_\{rank\}=\\max\\\{0,m\+D\(x,y^\{\+\}\)\-D\(x,y^\{\-\}\)\\\},\(8\)where\(x,y\+\)\(x,y^\{\+\}\)is a positive directed pair andy−y^\{\-\}is a sampled corrupted target\. Thus a minibatch contains triples\(x,y\+,y−\)\(x,y^\{\+\},y^\{\-\}\), but the learned score remains an ordered\-pair divergenceD​\(x,y\)D\(x,y\)\. The loss does not regress to a known numerical divergence; it only enforces that the learned divergence ranks the annotated relation above sampled non\-relations\. Across epochs, the same positive pair can be compared with different corrupted targets\. For direction supervision, we add a forward\-reverse margin:

ℒd​i​r=max⁡\{0,md\+D​\(x,y\)−D​\(y,x\)\}\.\\mathcal\{L\}\_\{dir\}=\\max\\\{0,m\_\{d\}\+D\(x,y\)\-D\(y,x\)\\\}\.\(9\)The full objective is

ℒ=ℒr​a​n​k\+α​ℒd​i​r,\\mathcal\{L\}=\\mathcal\{L\}\_\{rank\}\+\\alpha\\mathcal\{L\}\_\{dir\},\(10\)whereα\\alphacontrols the ranking\-direction trade\-off\.

Thus, the main training signal is relational: positive pairs should be closer than corrupted pairs, and annotated forward directions should be preferred over their reversals\. In the reported experiments, the input representations are held fixed and only the comparison head is optimized\. This design isolates the behavior of the proposed head and allows the same training protocol to be applied to heterogeneous datasets whose labels are edges, entailment judgments, or graded relation strengths rather than explicit metric distances\.

## 4Theoretical Analysis

### 4\.1Nonnegativity and identity in projected space

Proposition 1\.Ifϕ\\phiis convex and differentiable, the role\-aware head in Eq\. \([7](https://arxiv.org/html/2607.01762#S3.E7)\) is nonnegative\. Ifϕ\\phiis strictly convex, its zero set is characterized by equality in the role\-projected space\.

Proof\.The first\-order supporting hyperplane property of convex functions implies that for allu,vu,v\[[18](https://arxiv.org/html/2607.01762#bib.bib23),[5](https://arxiv.org/html/2607.01762#bib.bib5)\],

Dϕ​\(u,v\)≥0\.D\_\{\\phi\}\(u,v\)\\geq 0\.\(11\)Therefore Eq\. \([7](https://arxiv.org/html/2607.01762#S3.E7)\) satisfies

Dθ,Ps,Pt​\(x,y\)≥0D\_\{\\theta,P\_\{s\},P\_\{t\}\}\(x,y\)\\geq 0\(12\)for allx,yx,y\. Ifϕ\\phiis strictly convex, thenDϕ​\(u,v\)=0D\_\{\\phi\}\(u,v\)=0if and only ifu=vu=v\. Consequently,

Dθ,Ps,Pt​\(x,y\)=0⟺Ps​x=Pt​y\.D\_\{\\theta,P\_\{s\},P\_\{t\}\}\(x,y\)=0\\quad\\Longleftrightarrow\\quad P\_\{s\}x=P\_\{t\}y\.\(13\)The projected head therefore no longer claims thatD​\(x,x\)=0D\(x,x\)=0in the original input space\. Instead, it has a role\-diagonal identity: a pair has zero divergence when the source representation of the first item equals the target representation of the second item\. This is appropriate for directed relations, where source and target roles are semantically distinct\.□\\square

### 4\.2Quadratic special case and relation to projected metrics

The proposed head contains familiar distance heads as special cases\. This helps clarify which part of the model creates directionality\.

Proposition 2\.If the potential is quadratic,

ϕ​\(u\)=12​u⊤​H​u,H⪰0,\\phi\(u\)=\\frac\{1\}\{2\}u^\{\\top\}Hu,\\quad H\\succeq 0,\(14\)then the role\-aware Bregman head becomes a role\-aware Mahalanobis distance in the projected space:

Dϕ​\(Ps​x,Pt​y\)=12​\(Ps​x−Pt​y\)⊤​H​\(Ps​x−Pt​y\)\.D\_\{\\phi\}\(P\_\{s\}x,P\_\{t\}y\)=\\frac\{1\}\{2\}\(P\_\{s\}x\-P\_\{t\}y\)^\{\\top\}H\(P\_\{s\}x\-P\_\{t\}y\)\.\(15\)IfH=IH=I, this reduces to a squared Euclidean distance between the source\-role representation ofxxand the target\-role representation ofyy\.

Proof\.For a quadratic potential,∇ϕ​\(v\)=H​v\\nabla\\phi\(v\)=Hv\. Substituting into the Bregman definition gives

Dϕ​\(u,v\)\\displaystyle D\_\{\\phi\}\(u,v\)=12​u⊤​H​u−12​v⊤​H​v−v⊤​H​\(u−v\)\\displaystyle=\\frac\{1\}\{2\}u^\{\\top\}Hu\-\\frac\{1\}\{2\}v^\{\\top\}Hv\-v^\{\\top\}H\(u\-v\)=12​\(u−v\)⊤​H​\(u−v\)\.\\displaystyle=\\frac\{1\}\{2\}\(u\-v\)^\{\\top\}H\(u\-v\)\.\(16\)Takingu=Ps​xu=P\_\{s\}xandv=Pt​yv=P\_\{t\}ygives Eq\. \([15](https://arxiv.org/html/2607.01762#S4.E15)\)\.□\\square

This special case shows that the proposed model is not merely an unconstrained asymmetric neural scorer\. It generalizes a role\-aware projected metric by replacing the fixed quadratic potential with a learned convex potential\.

### 4\.3Source of directional asymmetry

Although Bregman divergences can be asymmetric even without role projections, the experiments show that this asymmetry does not always align with the annotated source\-to\-target relation\. The role\-aware formulation makes the source of directionality explicit\. Let

a=Ps​x,b=Pt​y,c=Ps​y,d=Pt​x\.a=P\_\{s\}x,\\quad b=P\_\{t\}y,\\quad c=P\_\{s\}y,\\quad d=P\_\{t\}x\.\(17\)The directional gap can be written as

G​\(x,y\)=Dϕ​\(a,b\)−Dϕ​\(c,d\)\.G\(x,y\)=D\_\{\\phi\}\(a,b\)\-D\_\{\\phi\}\(c,d\)\.\(18\)
Proposition 3\.Under the quadratic potential in Proposition 2, the directional gap has the exact decomposition

G​\(x,y\)\\displaystyle G\(x,y\)=12​\(Ps​x−Pt​y\)⊤​H​\(Ps​x−Pt​y\)\\displaystyle=\\frac\{1\}\{2\}\(P\_\{s\}x\-P\_\{t\}y\)^\{\\top\}H\(P\_\{s\}x\-P\_\{t\}y\)−12​\(Ps​y−Pt​x\)⊤​H​\(Ps​y−Pt​x\)\.\\displaystyle\\quad\-\\frac\{1\}\{2\}\(P\_\{s\}y\-P\_\{t\}x\)^\{\\top\}H\(P\_\{s\}y\-P\_\{t\}x\)\.\(19\)
Equation \([19](https://arxiv.org/html/2607.01762#S4.E19)\) shows that directionality is produced by an interaction between role\-specific projections and the geometry induced byHH\. IfPs=PtP\_\{s\}=P\_\{t\}andHHis symmetric positive semidefinite, the two terms become identical after swappingxxandyy, so the gap vanishes in the quadratic case\. Thus, in this interpretable limit, role separation is necessary for directional preference\. With a neural convex potential, additional asymmetry can also arise from the non\-quadratic shape ofϕ\\phi, but the role maps still determine which representation is evaluated as source and which is evaluated as target\.

### 4\.4Convexity and local curvature

For a fixed second argumentvv,Dϕ​\(u,v\)D\_\{\\phi\}\(u,v\)is convex in the first argumentuubecause it is the sum of the convex functionϕ​\(u\)\\phi\(u\)and a linear term inuu\. For the role\-aware head, this means that the divergence is convex in the source\-role variablePs​xP\_\{s\}xwhen the target\-role variablePt​yP\_\{t\}yis fixed\.

The local geometry is controlled by the Hessian of the potential\. Ifu=v\+δu=v\+\\deltaandϕ\\phiis twice differentiable, Taylor expansion aroundvvgives

Dϕ​\(v\+δ,v\)≈12​δ⊤​∇2ϕ​\(v\)​δ\.D\_\{\\phi\}\(v\+\\delta,v\)\\approx\\frac\{1\}\{2\}\\delta^\{\\top\}\\nabla^\{2\}\\phi\(v\)\\delta\.\(20\)For the role\-aware head,δ=Ps​x−Pt​y\\delta=P\_\{s\}x\-P\_\{t\}ydescribes the local displacement between the source\-role and target\-role representations\. The trace and eigenvalues of∇2ϕ​\(Pt​y\)\\nabla^\{2\}\\phi\(P\_\{t\}y\)therefore quantify local curvature around the target role\. In our experiments, the minimum eigenvalues remain positive because the potential includes a strongly convex quadratic component\. This supports the interpretation that the learned head does not behave like an unconstrained black\-box scorer\.

### 4\.5Directional gap

We define the directional gap

G​\(x,y\)=D​\(x,y\)−D​\(y,x\)\.G\(x,y\)=D\(x,y\)\-D\(y,x\)\.\(21\)For a directed positive pair\(x,y\)\(x,y\), a negative gap means the learned divergence prefers the annotated forward direction\. Unlike ranking accuracy, which compares positive and corrupted pairs, the directional gap directly measures whether the head distinguishes the observed relation from its reversal\. This quantity is also useful for case studies: pairs with the largest negative gaps are examples where the model expresses strongest directional confidence\.

## 5Experiments

### 5\.1Datasets

We evaluate on five main directed semantic and ontology benchmarks\. HyperLex contains graded lexical entailment pairs represented with GloVe embeddings\[[16](https://arxiv.org/html/2607.01762#bib.bib21)\]\. WordNet contains lexical hierarchy edges also represented with GloVe embeddings\. SICK and SNLI contain sentence entailment pairs represented with Sentence\-BERT embeddings\[[17](https://arxiv.org/html/2607.01762#bib.bib22)\]\. Gene Ontology contains directed biological concept relations represented with Sentence\-BERT embeddings of term names and definitions\.

Across these datasets, the notation\(x,y\)\(x,y\)always denotes an ordered pair rather than an unordered similarity pair\. In HyperLex,xxis the more specific lexical item andyyis the more general lexical item; the dataset provides a graded lexical entailment score, which we use as relation strength and for filtering/evaluation\. In WordNet,xxis a child synset andyyis a hypernym synset; the label is the existence of a hypernym edge\. In SICK and SNLI,xxis the premise sentence andyyis the hypothesis sentence; entailment pairs are treated as positive directed relations\. In Gene Ontology,xxis a child GO term andyyis a parent GO term connected by anis\_aorpart\_ofrelation\. In all cases, the numerical divergenceD​\(x,y\)D\(x,y\)is not observed in the dataset; it is learned from these relation labels by ranking annotated targets above sampled corrupted targets\.

We also evaluate OGBL\-Citation2 as a larger fixed\-feature stress test in which the distance head is applied to fixed node features without a graph neural encoder\.

For citation graphs,xxis the citing paper andyyis the cited paper, and the label is whether the directed citation edge exists\. Node features are fixed bag\-of\-words or provided node attributes\. Positive pairs are observed citation edges; negative targets are sampled papers that are not observed as cited by the same source paper\.

### 5\.2Baselines

We compare the proposed role\-aware Bregman head with symmetric Euclidean, cosine, and Mahalanobis distances; an unstructured MLP scorer; a bilinear asymmetric scorer; a plain ICNN\-Bregman head without role projections; order embeddings; and a Poincare\-style hyperbolic baseline\. The comparison is designed to separate three factors: symmetry versus asymmetry, structured versus unstructured asymmetry, and plain Bregman asymmetry versus role\-aware Bregman asymmetry\.

### 5\.3Evaluation metrics

Ranking accuracy measures whether a positive pair is scored better than sampled negatives\. Direction accuracy measures whether the annotated direction is preferred to the reversed pair\. For link\-prediction and parent\-retrieval evaluation, we also report AUC, average precision, mean reciprocal rank \(MRR\), and Hits atKK\(Hits@K\), following common ranking\-based evaluation practice for link prediction\[[10](https://arxiv.org/html/2607.01762#bib.bib15)\]\. MRR is the average reciprocal rank of the true target among candidate targets: if the annotated target is ranked at positionrr, the contribution is1/r1/r\. Hits atKKis the fraction of queries for which the annotated target appears in the topKKranked candidates\. Negative divergence rate measures the fraction of evaluated pairs assigned a negative value; for a valid divergence this should be zero\. Directional gap statistics and Hessian traces are reported for interpretability\.

All main semantic and ontology results use ten random seeds\. We report mean and standard deviation\. For the key comparison between the proposed role\-aware Bregman head and the plain ICNN\-Bregman head, we use paired seed\-level tests and bootstrap confidence intervals over the seed differences\.

## 6Results and Analysis

### 6\.1Main directed semantic and ontology results

Table[1](https://arxiv.org/html/2607.01762#S6.T1)summarizes the central results\. The proposed role\-aware Bregman head consistently improves direction accuracy over the plain ICNN\-Bregman head on all five main datasets\. The improvements are large on HyperLex and SNLI, moderate on SICK and WordNet, and smaller but consistent on Gene Ontology\. At the same time, the proposed head keeps a zero observed negative divergence rate, unlike the bilinear scorer, which often produces negative values\.

Table 1:Main ten\-seed results\. Values are mean±\\pmstandard deviation\. R\-Acc is ranking accuracy; D\-Acc is direction accuracy\. Bold entries mark the proposed role\-aware Bregman head and its key direction/validity measurements, rather than best\-in\-column values\.
### 6\.2Statistical significance

Table[2](https://arxiv.org/html/2607.01762#S6.T2)reports paired ten\-seed comparisons between the role\-aware and plain ICNN\-Bregman heads\. Direction accuracy improves significantly on every main dataset\. The same analysis also confirms that the role\-aware Bregman head has a significantly lower negative\-value rate than the bilinear scorer\.

Table 2:Paired seed\-level significance tests for direction accuracy: role\-aware Bregman minus plain ICNN\-Bregman\. CI denotes bootstrap confidence interval over seed differences;pp\-values are from an exact two\-sided sign test over the ten paired seeds\.Note: with ten paired seeds, the exact two\-sided sign test is discrete\. If all ten seed\-level differences favor the same method, the minimum attainablepp\-value is2/210=0\.0019532/2^\{10\}=0\.001953, which explains why the reportedpp\-values are identical across datasets\.

### 6\.3Projection and direction\-weight ablations

Table[3](https://arxiv.org/html/2607.01762#S6.T3)reports the ablation evidence\. The projection ablation uses HyperLex, SNLI, and WordNet over five seeds and compares plain ICNN\-Bregman, shared projection, source\-only projection, target\-only projection, and the full source\-target role\-aware head\. The results show that role separation is the key design choice\. Shared projection variants behave similarly to plain ICNN\-Bregman heads and often fail to recover directionality\. Source\-only and target\-only projections improve direction accuracy, showing that the improvement does not come merely from adding parameters\. This empirical pattern agrees with the quadratic gap decomposition in Eq\. \([19](https://arxiv.org/html/2607.01762#S4.E19)\): when source and target roles are not separated, the projected quadratic limit cannot express a directional preference between a pair and its reversal\. The full source\-target variant is a robust default, while target\-only projection can be especially strong on hierarchy\-like target roles\.

The direction\-weight sweep in Table[3](https://arxiv.org/html/2607.01762#S6.T3)shows a clear ranking\-direction trade\-off\. Increasing the weight of the forward\-reverse margin improves direction accuracy up to a moderate range\. On SNLI and WordNet, very large weights provide little additional direction gain and slightly reduce ranking quality\. This supports the interpretation that directionality is not obtained for free; it must be balanced against positive\-negative ranking\.

Table 3:Projection and direction\-weight ablations over five seeds\. The upper panel reports direction accuracy for projection variants\. The lower panel reports ranking accuracy / direction accuracy for different direction\-loss weights using the full role\-aware projected Bregman head\.
### 6\.4Interpretability

Interpretability is one of the main reasons for using a structured divergence head instead of an unconstrained asymmetric scorer\. We therefore analyze the learned head at three levels\. The first level is*validity*: because the score is a Bregman divergence in the role\-projected space, negative distance values should not occur\. The second level is*directional preference*: the directional gapG​\(x,y\)=D​\(x,y\)−D​\(y,x\)G\(x,y\)=D\(x,y\)\-D\(y,x\)indicates whether the model prefers the annotated direction or its reversal\. The third level is*local geometry*: the Hessian of the convex potential describes how sensitive the divergence is to perturbations around the target\-role representation\. These diagnostics operationalize the theoretical analysis in Section 4: the gap measures the observable consequence of the source\-target decomposition, while the Hessian summarizes the local convex geometry through which role\-projected differences are evaluated\.

Table[4](https://arxiv.org/html/2607.01762#S6.T4)and Figure[2](https://arxiv.org/html/2607.01762#S6.F2)summarize these diagnostics\. The directional gap becomes more consistently negative for annotated forward pairs after role\-aware projection\. On HyperLex, the fraction of pairs with forward\-preferred gap increases from0\.33580\.3358for plain ICNN\-Bregman to0\.91240\.9124for the role\-aware head\. This is not merely a ranking improvement against sampled negatives; it means that the head learns a directional preference between an observed pair and its reversed counterpart\. WordNet and Gene Ontology already have stronger directionality under the plain Bregman head, but role\-aware projection still increases the forward\-preferred rate and makes the average gap more negative\.

The case\-level view in Figure[2](https://arxiv.org/html/2607.01762#S6.F2)shows how this diagnostic can be used in practice\. For the HyperLex pair celery→\\rightarrowfood, the learned divergence isD​\(celery,food\)=0\.16D\(\\text\{celery\},\\text\{food\}\)=0\.16in the annotated direction andD​\(food,celery\)=1\.93D\(\\text\{food\},\\text\{celery\}\)=1\.93in the reverse direction, producing a gap of−1\.77\-1\.77\. This gives an interpretable answer to the question “why did the model prefer this direction?”: the source\-role representation of the specific concept is close to the target\-role representation of the general concept, while the reverse role assignment is much farther away\. Such examples can be inspected directly, sorted by gap magnitude, or compared with annotation scores\.

The Hessian analysis complements the gap analysis\. A generic MLP scorer can produce a directional score, but it does not provide a convex potential whose curvature can be inspected\. In contrast, the role\-aware Bregman head allows local curvature summaries such as the Hessian trace, top eigenvalue, and minimum eigenvalue\. The positive minimum eigenvalues observed in the case\-level example and aggregate analysis are expected from the strongly convex quadratic component, and they provide a sanity check that the head remains in a convex\-divergence regime\. The trace values are larger for WordNet and Gene Ontology than for HyperLex, suggesting that the potential uses stronger local curvature on denser or more heterogeneous hierarchy structures\. We do not treat the trace as a causal explanation of a prediction, but as a geometric diagnostic of how the learned divergence organizes the projected space\.

Together, these diagnostics distinguish the proposed head from both symmetric distances and unconstrained asymmetric scorers\. Symmetric distances cannot produce nonzero directional gaps by construction\. MLP and bilinear scorers can express directionality, but their scores are not guaranteed to be nonnegative divergences and they do not yield a convex potential for Hessian\-based inspection\. The role\-aware Bregman head therefore provides a middle ground: its predictions are still learned neural scores, but they can be audited through directional gaps, divergence validity, and local convex geometry\. In this sense, the interpretability analysis is not an auxiliary visualization step; it is the empirical counterpart of the gap decomposition and curvature characterization\.

Table 4:Directional gap and curvature diagnostics\. A higherG<0G<0rate means the model more often prefers the annotated forward direction\.![Refer to caption](https://arxiv.org/html/2607.01762v1/figures/case_interpretability_hyperlex.png)Figure 2:Case\-level interpretability on HyperLex\. For the pair celery→\\rightarrowfood, the role\-aware Bregman head assigns a much smaller forward divergence than reverse divergence, producing a negative directional gap\. The same example also reports local Hessian diagnostics of the convex potential\.
### 6\.5OGB citation stress test

Table[5](https://arxiv.org/html/2607.01762#S6.T5)reports OGBL\-Citation2 results with fixed node features and no graph neural encoder\. Poincare and cosine baselines achieve stronger ranking accuracy and MRR than the proposed head\. The role\-aware Bregman head slightly improves ranking accuracy and AUC over plain ICNN\-Bregman, but does not dominate this large citation setting\. The result suggests that large\-scale citation prediction requires either stronger input representations or integration with a graph encoder\.

Table 5:OGBL\-Citation2 fixed\-feature stress test over five seeds\.

## 7Discussion

The experiments support three conclusions\. First, plain Bregman asymmetry is not sufficient in practice\. Although Bregman divergences are mathematically asymmetric, the learned head can still fail to align that asymmetry with the annotated relation direction\. Role\-aware projections address this mismatch by allowing an embedding to take different forms when it acts as a source or target\. This conclusion is consistent with the theoretical gap decomposition: the direction score depends not only on the convex potential, but also on how a pair is assigned to source and target roles\.

Second, the proposed method should be understood as a structured alternative to generic asymmetric scorers\. MLP scorers often achieve strong direction accuracy and can be stronger on pure retrieval metrics\. However, they do not provide nonnegative divergence values, projected\-space identity, or Hessian\-based curvature diagnostics\. Conversely, symmetric distances can retrieve semantically or topologically related items well, but they remain direction\-blind\. The role\-aware Bregman head is therefore most attractive when the user wants both directional performance and a distance\-like geometric explanation\.

Third, the method is not a universal ranking winner\. Symmetric distances and hyperbolic baselines remain highly competitive when the input embeddings already encode the task well or when the evaluation mostly rewards undirected neighborhood quality\. This is visible in SICK, SNLI, Gene Ontology, and OGB\. The contribution is therefore not that role\-aware Bregman heads replace all scorers, but that they provide a principled option for directed relations where geometric validity and interpretation matter\.

## 8Limitations and Future Work

The current experiments use fixed embeddings, which isolates the behavior of the comparison head but does not test full end\-to\-end representation learning\. Joint encoder and head training may improve ranking performance, especially on graph benchmarks where the OGB stress test currently uses fixed node features rather than a graph neural encoder\. The role projections are mostly linear; nonlinear role maps could increase expressiveness but may weaken interpretability\. The empirical study also focuses on head\-level comparisons rather than complete task\-specific systems\. Future work should evaluate the head inside stronger encoders, study robustness to noisy directed labels, and develop generalization guarantees for convex\-divergence heads\.

## 9Conclusion

This paper presented a role\-aware neural convex divergence head for asymmetric representation learning\. By combining source\-target role projections with an ICNN\-induced Bregman divergence, the method preserves structured nonnegative divergence values while improving directional discrimination over plain ICNN\-Bregman heads\. The theoretical analysis characterizes how the head retains projected\-space Bregman properties and how directional gaps arise from the interaction between role separation and convex geometry\. Experiments across lexical, sentence, ontology, and graph benchmarks show that the method is especially useful when asymmetric relations are central and interpretability is desired\. The ablation and interpretability analyses further connect the theory to practice: role separation drives directional gains, while gap and Hessian diagnostics make the learned asymmetry inspectable\. The results also clarify its limitations: unstructured MLPs and specialized geometric baselines can outperform it on pure ranking metrics\. Overall, role\-aware neural convex divergence heads provide a practical and interpretable plug\-in distance head for directed representation learning\.

## CRediT authorship contribution statement

He Huang: Conceptualization, Methodology, Software, Formal analysis, Investigation, Writing – original draft, Funding acquisition, Project administration\. Lu Shen: Methodology, Investigation, Validation, Resources, Writing – review and editing, Funding acquisition\. Yunfeng Huang: Software, Validation, Formal analysis, Visualization, Writing – review and editing\. Li Qi: Supervision, Funding acquisition, Resources, Writing – review and editing\.

## Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.

## Data availability

## Declaration of generative AI and AI\-assisted technologies in the manuscript preparation process

During the preparation of this work, the authors used OpenAI ChatGPT/Codex to assist with language editing, formatting, code review, and manuscript organization\. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article\.

## Funding

This work was supported by the Science and Technology Research Program of Chongqing Municipal Education Commission \[grant numbers KJQN202500828, KJQN202500837\] and the Chongqing Technology and Business University High Level Talent Research Project \[grant numbers 2556012, 2556008\]\.

## References

- \[1\]B\. Amos, L\. Xu, and J\. Z\. Kolter\(2017\)Input convex neural networks\.InProceedings of the International Conference on Machine Learning,pp\. 146–155\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.1),[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1),[§3\.2](https://arxiv.org/html/2607.01762#S3.SS2.p1.1)\.
- \[2\]A\. Banerjee, S\. Merugu, I\. S\. Dhillon, and J\. Ghosh\(2005\)Clustering with bregman divergences\.Journal of Machine Learning Research6,pp\. 1705–1749\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.1)\.
- \[3\]A\. Bordes, N\. Usunier, A\. Garcia\-Duran, J\. Weston, and O\. Yakhnenko\(2013\)Translating embeddings for modeling multi\-relational data\.InAdvances in Neural Information Processing Systems,pp\. 2787–2795\.Cited by:[§2\.2](https://arxiv.org/html/2607.01762#S2.SS2.p1.1),[§3\.4](https://arxiv.org/html/2607.01762#S3.SS4.p1.2)\.
- \[4\]S\. R\. Bowman, G\. Angeli, C\. Potts, and C\. D\. Manning\(2015\)A large annotated corpus for learning natural language inference\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 632–642\.Cited by:[§2\.4](https://arxiv.org/html/2607.01762#S2.SS4.p1.1)\.
- \[5\]S\. Boyd and L\. Vandenberghe\(2004\)Convex optimization\.Cambridge University Press,Cambridge\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.1),[§4\.1](https://arxiv.org/html/2607.01762#S4.SS1.p2.1)\.
- \[6\]L\. M\. Bregman\(1967\)The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming\.USSR Computational Mathematics and Mathematical Physics7\(3\),pp\. 200–217\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.2),[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1)\.
- \[7\]Y\. Censor and A\. Lent\(1981\)An iterative row\-action method for interval convex programming\.Journal of Optimization Theory and Applications34\(3\),pp\. 321–353\.Cited by:[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1)\.
- \[8\]H\. K\. Cilingir, R\. Manzelli, and B\. Kulis\(2020\)Deep divergence learning\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 2027–2037\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.1),[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1)\.
- \[9\]J\. V\. Davis, B\. Kulis, P\. Jain, S\. Sra, and I\. S\. Dhillon\(2007\)Information\-theoretic metric learning\.InProceedings of the International Conference on Machine Learning,pp\. 209–216\.Cited by:[§2\.1](https://arxiv.org/html/2607.01762#S2.SS1.p1.1)\.
- \[10\]W\. Hu, M\. Fey, M\. Zitnik, Y\. Dong, H\. Ren, B\. Liu, M\. Catasta, and J\. Leskovec\(2020\)Open graph benchmark: datasets for machine learning on graphs\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 22118–22133\.Cited by:[§2\.4](https://arxiv.org/html/2607.01762#S2.SS4.p1.1),[§5\.3](https://arxiv.org/html/2607.01762#S5.SS3.p1.5)\.
- \[11\]B\. Kulis\(2013\)Metric learning: a survey\.Foundations and Trends in Machine Learning5\(4\),pp\. 287–364\.Cited by:[§2\.1](https://arxiv.org/html/2607.01762#S2.SS1.p1.1),[§3\.4](https://arxiv.org/html/2607.01762#S3.SS4.p1.2)\.
- \[12\]F\. Lu, E\. Raff, and F\. Ferraro\(2023\)Neural bregman divergences for distance learning\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.1),[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1)\.
- \[13\]M\. Marelli, S\. Menini, M\. Baroni, L\. Bentivogli, R\. Bernardi, and R\. Zamparelli\(2014\)A sick cure for the evaluation of compositional distributional semantic models\.InProceedings of the Ninth International Conference on Language Resources and Evaluation,pp\. 216–223\.Cited by:[§2\.4](https://arxiv.org/html/2607.01762#S2.SS4.p1.1)\.
- \[14\]G\. A\. Miller\(1995\)WordNet: a lexical database for english\.Communications of the ACM38\(11\),pp\. 39–41\.Cited by:[§2\.4](https://arxiv.org/html/2607.01762#S2.SS4.p1.1)\.
- \[15\]M\. Nickel and D\. Kiela\(2017\)Poincare embeddings for learning hierarchical representations\.InAdvances in Neural Information Processing Systems,pp\. 6341–6350\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.01762#S2.SS2.p1.1)\.
- \[16\]J\. Pennington, R\. Socher, and C\. D\. Manning\(2014\)GloVe: global vectors for word representation\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 1532–1543\.Cited by:[§5\.1](https://arxiv.org/html/2607.01762#S5.SS1.p1.1)\.
- \[17\]N\. Reimers and I\. Gurevych\(2019\)Sentence\-bert: sentence embeddings using siamese bert\-networks\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 3982–3992\.Cited by:[§5\.1](https://arxiv.org/html/2607.01762#S5.SS1.p1.1)\.
- \[18\]R\. T\. Rockafellar\(1970\)Convex analysis\.Princeton University Press,Princeton, NJ\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p3.1),[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1),[§4\.1](https://arxiv.org/html/2607.01762#S4.SS1.p2.1)\.
- \[19\]A\. Siahkamari, X\. Xia, V\. Saligrama, D\. A\. Castanon, and B\. Kulis\(2020\)Learning to approximate a bregman divergence\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 3603–3612\.Cited by:[§2\.3](https://arxiv.org/html/2607.01762#S2.SS3.p1.1)\.
- \[20\]Z\. Sun, Z\. Deng, J\. Nie, and J\. Tang\(2019\)RotatE: knowledge graph embedding by relational rotation in complex space\.InInternational Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2607.01762#S2.SS2.p1.1)\.
- \[21\]The Gene Ontology Consortium\(2019\)The gene ontology resource: 20 years and still going strong\.Nucleic Acids Research47\(D1\),pp\. D330–D338\.Cited by:[§2\.4](https://arxiv.org/html/2607.01762#S2.SS4.p1.1)\.
- \[22\]A\. Tversky\(1977\)Features of similarity\.Psychological Review84\(4\),pp\. 327–352\.Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p2.1)\.
- \[23\]I\. Vendrov, R\. Kiros, S\. Fidler, and R\. Urtasun\(2016\)Order\-embeddings of images and language\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.01762#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.01762#S2.SS2.p1.1)\.
- \[24\]I\. Vulic, D\. Gerz, D\. Kiela, F\. Hill, and A\. Korhonen\(2017\)HyperLex: a large\-scale evaluation of graded lexical entailment\.Computational Linguistics43\(4\),pp\. 781–835\.Cited by:[§2\.4](https://arxiv.org/html/2607.01762#S2.SS4.p1.1)\.
- \[25\]K\. Q\. Weinberger and L\. K\. Saul\(2009\)Distance metric learning for large margin nearest neighbor classification\.Journal of Machine Learning Research10,pp\. 207–244\.Cited by:[§2\.1](https://arxiv.org/html/2607.01762#S2.SS1.p1.1),[§3\.4](https://arxiv.org/html/2607.01762#S3.SS4.p1.2)\.

Similar Articles

Predictive Divergence Masks for LLM RL

Hugging Face Daily Papers

Proposes predictive divergence masks for LLM reinforcement learning that improve upon PPO's direction criterion by predicting whether the next policy gradient step will increase or decrease the divergence used by the trust region, leading to better alignment and improved RL training across model scales.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Hugging Face Daily Papers

Proposes Asymmetric Mutual Variational Learning (AMVL) to resolve train-inference mismatch in multimodal continuous reasoning by using bidirectional calibration to prevent answer leakage and improve latent-space stability, achieving significant gains on the BLINK benchmark.