semantic-equivalence

Tag

Cards List
#semantic-equivalence

When Does Consensus Mean Correctness? Measuring the Agreement-Accuracy Coupling with Semantics-Preserving Re-Rendering

arXiv cs.LG · 2026-08-07 Cached

This paper introduces RENDEQ, a generator of render-equivalence sets for scientific figures, and measures how well model agreement across semantics-preserving re-renderings tracks correctness in open-weight VLMs. It finds that agreement certifies correctness only above a threshold and that fine-tuning on self-consensus can hurt accuracy.

0 favorites 0 likes
#semantic-equivalence

Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

arXiv cs.CL · 2026-08-06 Cached

This paper proposes a framework to elicit intrinsic hallucinations in LLMs using semantically equivalent adversarial perturbations, showing that state-of-the-art models degrade significantly in contextual faithfulness even with meaning-preserving query variations.

0 favorites 0 likes
#semantic-equivalence

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

arXiv cs.AI · 2026-08-03 Cached

ModelEquivBench is a certifying multi-relational evaluation system for LLM-generated optimization models, reporting per-pair semantic profiles across seven equivalence relations instead of a single accuracy score. It evaluates GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B on a fixed benchmark, revealing stage-wise failures that coarse baselines miss.

0 favorites 0 likes
#semantic-equivalence

A Dynamical Framework for Cognitive Processes Based on Transformations and Semantic Equivalence

arXiv cs.AI · 2026-05-26 Cached

This paper proposes a structural and dynamical framework for modeling cognitive processes using iterative state transformations and semantic equivalence, integrating dynamical systems, category theory, and feedback mechanisms to model cognition as a process evolving toward stable interpretations.

0 favorites 0 likes
#semantic-equivalence

Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification

arXiv cs.CL · 2026-04-21 Cached

Researchers from University of Edinburgh propose a self-play framework using Liquid Haskell for formal verification to train LLMs on semantic equivalence reasoning, releasing OpInstruct-HSx dataset (28k programs) and achieving 13.3pp accuracy gains on EquiBench.

0 favorites 0 likes
← Back to home

Submit Feedback