Tag
This preregistered replication tests whether the monotonicity effect on label agreement in NLI generalizes from selected low-agreement items to unselected populations, finding that the effect reverses and is small, suggesting the earlier finding was conditional on selection.
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
This paper measures how much formal semantic structure explains human label variation in natural language inference (NLI) using ChaosNLI data, finding group-level effects on entropy but item-level ceilings and null composition effects.
ProvenanceGuard是一种用于MCP驱动的LLM代理的源感知事实性验证器,它通过分解回答为原子声明、路由到特定源证据、检查支持并验证归因,解决了跨源混淆问题。在医疗领域的评估中,它达到了0.802的块F1和0.858的源准确率。
StepGap is a hybrid NLI-LLM decision tree that detects step-level evidence gaps in multi-hop QA, labeling them as Contradicted Claim, Irrelevant Evidence, or Missing Bridge. It achieves competitive F1 while providing a decomposable structure that improves downstream QA performance when used as a process reward for reinforcement learning.
This paper identifies a systematic gap between legal interpretation and formal logic in AI legal reasoning, proposes a neuro-symbolic approach to bridge it, and demonstrates substantial label shifts when re-annotating legal NLI data under strict formal entailment.