Tag
This paper introduces an open benchmark and a specialized bi-encoder model for natural language code retrieval in the 1C:Enterprise ecosystem, addressing the lack of domain-specific resources for Russian-language code search.
Tevatron-Elastic presents a unified abstraction for training elastic retrievers and rerankers, enabling a single checkpoint to serve multiple model sizes across depth, token, and width axes. It generalizes prior methods like Matryoshka embeddings and early exit, and introduces Matryoshka LTC for joint token-compression training.
QuixiAI releases embeddinggemma.c, a fast cross-platform embedding inference engine written in C, supporting multiple backends (CPU, Metal, CUDA, ROCm, SYCL) and Matryoshka embeddings with a standard HTTP API.
CAMMAR introduces a representation learning framework that organizes Arabic metaphorical meaning into nested lexical, cultural, and metaphorical subspaces using a staged semantic curriculum, achieving strong metaphor detection (AUC up to 0.84) on a new span-annotated dataset.