Tag
The paper introduces role-aware neural convex divergence heads that apply source and target role projections before evaluating an input-convex neural Bregman divergence, enabling structured and interpretable asymmetric distance learning for tasks like lexical entailment, sentence entailment, and ontology hierarchy. Experiments show consistent improvements in directional accuracy over plain ICNN-Bregman heads across semantic and ontology benchmarks.
This paper reveals that Mirror Descent with non-quadratic regularizers can be exponentially more sensitive to initialization than Gradient Descent, even under well-conditioned settings, which has implications for reproducibility in RL and LLM post-training.