Tag
This paper introduces RENDEQ, a generator of render-equivalence sets for scientific figures, and measures how well model agreement across semantics-preserving re-renderings tracks correctness in open-weight VLMs. It finds that agreement certifies correctness only above a threshold and that fine-tuning on self-consensus can hurt accuracy.
This paper proposes a framework to elicit intrinsic hallucinations in LLMs using semantically equivalent adversarial perturbations, showing that state-of-the-art models degrade significantly in contextual faithfulness even with meaning-preserving query variations.
ModelEquivBench is a certifying multi-relational evaluation system for LLM-generated optimization models, reporting per-pair semantic profiles across seven equivalence relations instead of a single accuracy score. It evaluates GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B on a fixed benchmark, revealing stage-wise failures that coarse baselines miss.
This paper proposes a structural and dynamical framework for modeling cognitive processes using iterative state transformations and semantic equivalence, integrating dynamical systems, category theory, and feedback mechanisms to model cognition as a process evolving toward stable interpretations.
Researchers from University of Edinburgh propose a self-play framework using Liquid Haskell for formal verification to train LLMs on semantic equivalence reasoning, releasing OpInstruct-HSx dataset (28k programs) and achieving 13.3pp accuracy gains on EquiBench.