evaluation-framework

Tag

Cards List
#evaluation-framework

Concept Drift from a Causal Perspective

arXiv cs.LG · yesterday Cached

This paper proposes a causal framework for understanding concept drift in data streams using Structural Causal Models, with a taxonomy and generator for simulating and evaluating drift events in non-stationary environments.

0 favorites 0 likes
#evaluation-framework

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

arXiv cs.AI · yesterday Cached

TicTacBench is a new benchmark for evaluating coding agents' timing closure capabilities in RTL design, revealing that current agents have significant room for improvement and proposing TicTacSkill to enhance performance.

0 favorites 0 likes
#evaluation-framework

A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare

arXiv cs.LG · 2d ago Cached

This paper introduces OpTFM, a comparative evaluation framework for tabular foundation models in healthcare, assessing models across six clinically meaningful dimensions like generalization and fairness, and applies it to two use cases to demonstrate context-dependent rankings.

0 favorites 0 likes
#evaluation-framework

Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B

arXiv cs.CL · 2d ago Cached

This paper introduces LLM endognostics, a white-box framework for auditing latent knowledge in language models, showing that behavioral evaluations can fail to reflect internal knowledge distinctions, as demonstrated on Minerva-7B-Instruct-v1.0.

0 favorites 0 likes
#evaluation-framework

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

arXiv cs.CL · 2d ago Cached

The Situated Identity Test (SIT) is an architecture-independent framework introduced to evaluate whether an AI agent's behavior is functionally attributable to a specific developmental lineage, addressing the problem of persona imitation in large language models.

0 favorites 0 likes
#evaluation-framework

EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

arXiv cs.AI · 3d ago Cached

EnterpriseVal introduces a comprehensive evaluation system for generative AI in enterprises, addressing the measurement gap with a use-case-level framework that includes specifications, metrics, and a grading protocol, demonstrated through a pilot study in banking.

0 favorites 0 likes
#evaluation-framework

From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

arXiv cs.CL · 3d ago Cached

This paper introduces a productivity-oriented framework for evaluating human-AI collaboration based on outcome quality relative to interaction cost, showing that identical quality ratings can differ significantly in interaction costs and that subjective user ratings are not reliable for measuring productivity.

0 favorites 0 likes
#evaluation-framework

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

arXiv cs.AI · 6d ago Cached

EconSkills introduces a skill library and evaluation framework for web agents to transfer and retrieve procedural knowledge for live economic data retrieval, showing improved efficiency in controlled transfer and competitive performance at library scale.

0 favorites 0 likes
#evaluation-framework

Google DeepMind launches institute to widen the AGI debate

TechCrunch AI · 6d ago Cached

Google DeepMind has launched a new institute to advance discussions on AGI, featuring essays on AI transparency, evaluation frameworks, and safety principles.

0 favorites 0 likes
#evaluation-framework

Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues

arXiv cs.CL · 2026-09-17 Cached

This paper evaluates whether LLM-generated cyberbullying dialogues faithfully reproduce the social dynamics of authentic interactions, finding that while high-level structures are preserved, finer-grained details are systematically distorted in a model-dependent manner.

0 favorites 0 likes
#evaluation-framework

Skill-based Agentic Evaluation for Real-time Data Science Tasks

arXiv cs.AI · 2026-09-16 Cached

This paper introduces a ground-truth-as-code framework for evaluating data-science agents on continuously updated data, achieving improved agreement with human evaluators and reduced token consumption.

0 favorites 0 likes
#evaluation-framework

Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice

arXiv cs.CL · 2026-09-15 Cached

This research explores sycophancy in large language models when providing romantic relationship advice, revealing that perspective-driven query framing influences behavior more than grammatical mood, and that models tend to increase sycophancy over dialogue turns.

0 favorites 0 likes
#evaluation-framework

Verifiable Social Reasoning for LLM Assistants

Hugging Face Daily Papers · 2026-09-15 Cached

The paper introduces Fuse, a multi-agent simulation framework for evaluating social reasoning in LLM assistants by providing verifiable ground truth through user-mediated interactions, validated with a human study and applied to 12 LLMs.

0 favorites 0 likes
#evaluation-framework

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

arXiv cs.AI · 2026-09-12 Cached

This paper introduces LogiMed-RoB, a benchmark for evaluating large language models' hierarchical logical consistency in medical risk-of-bias assessment, revealing that high atomic consistency can conceal critical reasoning flaws in clinical deployment.

0 favorites 0 likes
#evaluation-framework

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

arXiv cs.AI · 2026-09-11 Cached

This paper presents a systematic black-box framework for evaluating agentic AI systems, introducing a taxonomy of risks and automated red teaming methods. Empirical validation across agent architectures reveals critical vulnerabilities, with high rates of governance and privacy risks.

0 favorites 0 likes
#evaluation-framework

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

arXiv cs.AI · 2026-09-11 Cached

ContractEval is a diagnostic framework that evaluates procedural conformance in LLM agents by making active obligations explicit and matching them against evidence, detecting failures missed by traditional judges.

0 favorites 0 likes
#evaluation-framework

The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods

arXiv cs.CL · 2026-09-11 Cached

This paper critiques optimization-based balancing strategies for multimodal sentiment analysis, showing they fail due to conflating fitting speed with discriminative importance, and proposes a new research agenda focusing on held-out modality valuation.

0 favorites 0 likes
#evaluation-framework

From Repetition to Recognition: Inductive Discovery of Disinformation Narratives

arXiv cs.CL · 2026-09-11 Cached

The paper introduces a three-tier evaluation framework for unsupervised narrative label generation and compares clustering-based and graph-community-based methods for discovering disinformation narratives, releasing human-validated labels to support taxonomy development.

0 favorites 0 likes
#evaluation-framework

CriticGen: Generation-Aware Evaluation as Actionable Feedback

arXiv cs.AI · 2026-09-10 Cached

CriticGen is a fine-grained, generation-aware evaluation framework for large language models that generates sample-specific evaluation criteria to provide actionable feedback for improving answer quality.

0 favorites 0 likes
#evaluation-framework

Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations

Hugging Face Daily Papers · 2026-09-06 Cached

This paper introduces Deep Persona, a psychologically grounded architecture for role-playing agents, and proposes an evaluation framework. It evaluates LLMs and finds systematic limitations in emotional expression despite high pragmatic fluency.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback