evaluation-protocol

Tag

Cards List
#evaluation-protocol

Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

arXiv cs.CL · 2d ago Cached

The paper introduces SCEval, a diagnostic evaluation protocol that applies structural corruptions to test the fragility of omni-modal large language models, revealing that clean accuracy does not ensure reliable cross-modal reasoning.

0 favorites 0 likes
#evaluation-protocol

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

arXiv cs.CL · 2026-08-21 Cached

The paper introduces Hear2Act, a unified benchmark for evaluating how prosodic cues affect task-oriented dialogue decisions, finding that prosody matters when lexical evidence is insufficient and that audio LLMs benefit from explicit concern representation for actions.

0 favorites 0 likes
#evaluation-protocol

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

Hugging Face Daily Papers · 2026-07-14 Cached

This paper presents a practical evaluation protocol for assessing AI pentesting agents in realistic, complex targets rather than simplified benchmarks. It uses LLM-based semantic matching, bipartite resolution, and continuous ground-truth to score vulnerabilities discovered, and releases expert-annotated ground truth and code.

0 favorites 0 likes
#evaluation-protocol

When Does Continual Learning Require Learning

arXiv cs.LG · 2026-07-10 Cached

This paper proposes a unified framework for continual learning in LLMs, disentangling change along space (new domains) and time (data drift). It evaluates various methods including prompting, supervised learning, reinforcement learning, and context compression under realistic sequential settings.

0 favorites 0 likes
#evaluation-protocol

When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory

arXiv cs.AI · 2026-05-11 Cached

This paper introduces a scale-conditioned evaluation protocol for agent memory, analyzing how reliability degrades as irrelevant sessions accumulate. It identifies specific failure regimes and usable-scale boundaries across different memory interfaces and LLMs.

0 favorites 0 likes
← Back to home

Submit Feedback