This paper investigates why language models fail at two-hop generalization, showing that models succeed when the second hop follows training distribution but fail when it deviates, and proposes a recurrent-style training strategy to improve out-of-distribution two-hop reasoning.
The paper introduces Stoicheia, a 405M-parameter character-level masked diffusion encoder for Ancient Greek that unifies textual restoration, parsing, and metrical scansion in a single model, outperforming prior systems like Ithaca on benchmark tasks.
本文引入了“结晶化问题”,用于评估文本到SQL系统中的可复用记忆,表明将经过验证的修正查询存储在每个数据库的库中,可以将BIRD上留出集的首次尝试准确率提高4.34个百分点,捕获按需修复所提供的44.4%的提升空间。受控干预措施识别出数据库特定内容是主要驱动因素。
本文提出使用线性探针和RFM概念向量,从LLM内部激活中测量文本的概念内容,并将其应用于ESG分类。最佳线性探针在无需任务特定微调的情况下接近微调分类器的准确率,且优于模型自身输出,表明激活携带了响应之外的概念内容。
本文提出了HNR-DAC,一种用于引用论文科学声明验证的两阶段框架,结合了难负样本重排序与分布对齐分类。该方法在NLPCC 2026 Task 10 Track 2上表现优异,以最高的Macro-F1位列排行榜第三名。
This paper proposes a hybrid knowledge graph generation pipeline that combines top-down grounding in Wikidata with bottom-up agentic synthesis to handle noisy, multilingual HR skill declarations, producing a scalable and self-healing skills taxonomy.
The paper investigates whether retrieving more evidence helps visual retrieval-augmented generation with diffusion language models, finding that unconditionally expanding evidence hurts accuracy due to semantic conflict, and proposes a training-free Entropy-Based Candidate Filter (ECF) to selectively admit evidence, improving accuracy across benchmarks.
本文介绍了 GPTKB 2.0,这是一个由 LLM 派生的大规模去歧义知识库,包含 160 万个规范化实体上的 3840 万条三元组。它提供了一个网页界面,用于浏览、SPARQL/自然语言查询,以及审计事实来源和消歧决策。
这篇预印本评估了六个大型语言模型在160个提示词中对提示框架和偏见提示的响应方式,发现LLMs即使在事实性语境下也会系统性地调整其响应以与提示框架保持一致,从而可能强化用户的偏见。
This paper introduces PHASE-Tree, a multi-timescale character-state representation for long-horizon role-playing dialogue, along with the LongEvoRoleBench benchmark to evaluate evolved-state generation. The proposed approach outperforms baselines on character-level, semantic, and embedding metrics across long-dialogue corpora.
介绍了Ekphrasis,一个包含400个任务的基准,用于衡量纯文本LLM的视觉创意构思,将有用性、表现力与新颖性分离开来。论文通过跨模态锚定研究验证了该基准,表明文本层面的视觉构思排序在渲染后基本保持不变。
This paper investigates how epistemic stance (qualifiers, attributions) survives memory compression in AI agent memory systems. It finds that making the stance explicit as a labelled field improves retention significantly, while merely lengthening the text does not.
This paper investigates how molecular generative models internally organize molecular identity in their latent spaces, revealing piecewise-constant regions and coarse-to-fine boundaries across three architectures.
本文回答了Hanneke、Moran和Waknine提出的一个开放问题,证明了直接和的不可知PAC学习曲线并非仅由单实例学习曲线和因子数量决定,从而提供了一种速率分离。
ELMZip 是一种新颖的卫星图像压缩框架,利用极限学习机(ELM)和区域分解实现高效的星上神经表示,仅传输紧凑的输出权重以减少下行数据量,同时保持较高的重建保真度。
本文介绍了包含来自1,196名受访者的29,870条步行适宜性评分的数据集,并提出了一种用户条件多模态深度学习框架,将视觉特征与个体评分者属性相结合,以捕捉步行感知中的主观变异性。该模型在排名一致性上比仅基于图像的基线提高了65%,表明评估环境的主体至关重要。
Ask-E 是一个新的基准测试和训练环境,用于评估和训练模型生成针对特定技能水平校准的问题,该技能水平由两个现有语言模型的能力定义。前沿模型在校准上的得分低于 50%,而在 Ask-E 上进行训练可以提高下游数学基准测试的成绩,且无需新的数学数据或基于正确性的奖励。
MiCoPro 提出了一种用于混合精度量化的端到端软硬件协同设计框架,利用硬件感知代理模型在延迟约束下搜索最优的逐层位宽,并直接部署到边缘硬件,在准确率下降不到3%的情况下,实现了最高40%的延迟降低。
本文提出将ZCA白化作为WEAT的几何预处理步骤,以解决嵌入各向异性问题,结果表明校准会改变超过30%结果的显著性状态,未经校准的偏差测量可能不可靠。
本文介绍了对生物标本记录中发现的历史和方言地名(这些地名不在当前地名录中)进行地理参照的方法,比较了确定性方法、概率方法和基于LLM的方法。概率推理实现了最高精度,而LLMs提供了有竞争力但精度稍低的估计。