nli

Tag

Cards List
#nli

Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

arXiv cs.CL · 2026-07-22 Cached

This preregistered replication tests whether the monotonicity effect on label agreement in NLI generalizes from selected low-agreement items to unselected populations, finding that the effect reverses and is small, suggesting the earlier finding was conditional on selection.

0 favorites 0 likes
#nli

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

arXiv cs.CL · 2026-07-20 Cached

This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.

0 favorites 0 likes
#nli

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI

arXiv cs.CL · 2026-07-20 Cached

This paper measures how much formal semantic structure explains human label variation in natural language inference (NLI) using ChaosNLI data, finding group-level effects on entropy but item-level ceilings and null composition effects.

0 favorites 0 likes
#nli

ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

arXiv cs.AI · 2026-06-17 Cached

ProvenanceGuard是一种用于MCP驱动的LLM代理的源感知事实性验证器,它通过分解回答为原子声明、路由到特定源证据、检查支持并验证归因,解决了跨源混淆问题。在医疗领域的评估中,它达到了0.802的块F1和0.858的源准确率。

0 favorites 0 likes
#nli

StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering

arXiv cs.CL · 2026-05-26 Cached

StepGap is a hybrid NLI-LLM decision tree that detects step-level evidence gaps in multi-hop QA, labeling them as Contradicted Claim, Irrelevant Evidence, or Missing Bridge. It achieves competitive F1 while providing a decomposable structure that improves downstream QA performance when used as a process reward for reinforcement learning.

0 favorites 0 likes
#nli

Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning

arXiv cs.AI · 2026-05-15 Cached

This paper identifies a systematic gap between legal interpretation and formal logic in AI legal reasoning, proposes a neuro-symbolic approach to bridge it, and demonstrates substantial label shifts when re-annotating legal NLI data under strict formal entailment.

0 favorites 0 likes
← Back to home

Submit Feedback