Tag
This paper presents HNR-DAC, a two-stage framework for scientific claim verification over cited papers, combining hard-negative reranking and distribution-aligned classification. It achieves strong results on NLPCC 2026 Task 10 Track 2, ranking third on the leaderboard with the highest Macro-F1.
Proposes Distribution-Aligned Self-Distillation (DASD), which dynamically filters tokens during self-distillation to preserve beneficial logical corrections while suppressing distributionally misaligned style noise, improving robust reasoning on math, code, and commonsense benchmarks.
The paper introduces PRISM, a method that inserts a distribution-alignment stage between supervised fine-tuning and reinforcement learning to mitigate distributional drift in multimodal models. It uses a black-box adversarial game with an MoE discriminator to improve RLVR performance on models like Qwen3-VL.