evaluation-protocols

Tag

Cards List
#evaluation-protocols

Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition

arXiv cs.LG · 2026-07-31 Cached

This paper examines how evaluation protocols affect reported accuracy in EEG emotion recognition, using a DGCNN on SEED and SEED-IV datasets. It demonstrates that subject-dependent, subject-disjoint, and cross-session evaluations answer different questions, and that checkpoint selection and test-set reuse can inflate accuracy.

0 favorites 0 likes
#evaluation-protocols

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

arXiv cs.CL · 2026-07-22 Cached

A comprehensive survey on multimodal humor understanding using large language models, covering methods, datasets, evaluation protocols, and challenges in interpreting humor in memes, cartoons, and comics.

0 favorites 0 likes
#evaluation-protocols

Rethinking the Evaluation of Harness Evolution for Agents

arXiv cs.AI · 2026-07-15 Cached

This paper re-evaluates the methodology of automatic harness evolution for LLM agents, highlighting that its gains may stem from additional test-time search rather than improved harness design, and that evaluation on the same benchmark risks overfitting. Experiments show that harness evolution does not consistently outperform simpler test-time scaling methods.

0 favorites 0 likes
#evaluation-protocols

Adversarial Graph Neural Network Benchmarks: Towards Practical and Fair Evaluation

arXiv cs.LG · 2026-05-08 Cached

This paper presents a comprehensive benchmark for evaluating adversarial attacks and defenses in Graph Neural Networks, highlighting the need for standardized and fair experimental protocols.

0 favorites 0 likes
#evaluation-protocols

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity

arXiv cs.CL · 2026-05-08 Cached

This paper introduces a paired-prompt protocol to measure 'evaluation-context divergence' in open-weight LLMs, finding that models behave differently depending on whether prompts are framed as evaluations or live deployments. The study highlights heterogeneity across models, with some being 'eval-cautious' and others 'deployment-cautious', raising concerns about the validity of safety benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback