Tag
This paper proposes a causal framework for understanding concept drift in data streams using Structural Causal Models, with a taxonomy and generator for simulating and evaluating drift events in non-stationary environments.
TicTacBench is a new benchmark for evaluating coding agents' timing closure capabilities in RTL design, revealing that current agents have significant room for improvement and proposing TicTacSkill to enhance performance.
This paper introduces OpTFM, a comparative evaluation framework for tabular foundation models in healthcare, assessing models across six clinically meaningful dimensions like generalization and fairness, and applies it to two use cases to demonstrate context-dependent rankings.
This paper introduces LLM endognostics, a white-box framework for auditing latent knowledge in language models, showing that behavioral evaluations can fail to reflect internal knowledge distinctions, as demonstrated on Minerva-7B-Instruct-v1.0.
The Situated Identity Test (SIT) is an architecture-independent framework introduced to evaluate whether an AI agent's behavior is functionally attributable to a specific developmental lineage, addressing the problem of persona imitation in large language models.
EnterpriseVal introduces a comprehensive evaluation system for generative AI in enterprises, addressing the measurement gap with a use-case-level framework that includes specifications, metrics, and a grading protocol, demonstrated through a pilot study in banking.
This paper introduces a productivity-oriented framework for evaluating human-AI collaboration based on outcome quality relative to interaction cost, showing that identical quality ratings can differ significantly in interaction costs and that subjective user ratings are not reliable for measuring productivity.
EconSkills introduces a skill library and evaluation framework for web agents to transfer and retrieve procedural knowledge for live economic data retrieval, showing improved efficiency in controlled transfer and competitive performance at library scale.
Google DeepMind has launched a new institute to advance discussions on AGI, featuring essays on AI transparency, evaluation frameworks, and safety principles.
This paper evaluates whether LLM-generated cyberbullying dialogues faithfully reproduce the social dynamics of authentic interactions, finding that while high-level structures are preserved, finer-grained details are systematically distorted in a model-dependent manner.
This paper introduces a ground-truth-as-code framework for evaluating data-science agents on continuously updated data, achieving improved agreement with human evaluators and reduced token consumption.
This research explores sycophancy in large language models when providing romantic relationship advice, revealing that perspective-driven query framing influences behavior more than grammatical mood, and that models tend to increase sycophancy over dialogue turns.
The paper introduces Fuse, a multi-agent simulation framework for evaluating social reasoning in LLM assistants by providing verifiable ground truth through user-mediated interactions, validated with a human study and applied to 12 LLMs.
This paper introduces LogiMed-RoB, a benchmark for evaluating large language models' hierarchical logical consistency in medical risk-of-bias assessment, revealing that high atomic consistency can conceal critical reasoning flaws in clinical deployment.
This paper presents a systematic black-box framework for evaluating agentic AI systems, introducing a taxonomy of risks and automated red teaming methods. Empirical validation across agent architectures reveals critical vulnerabilities, with high rates of governance and privacy risks.
ContractEval is a diagnostic framework that evaluates procedural conformance in LLM agents by making active obligations explicit and matching them against evidence, detecting failures missed by traditional judges.
This paper critiques optimization-based balancing strategies for multimodal sentiment analysis, showing they fail due to conflating fitting speed with discriminative importance, and proposes a new research agenda focusing on held-out modality valuation.
The paper introduces a three-tier evaluation framework for unsupervised narrative label generation and compares clustering-based and graph-community-based methods for discovering disinformation narratives, releasing human-validated labels to support taxonomy development.
CriticGen is a fine-grained, generation-aware evaluation framework for large language models that generates sample-specific evaluation criteria to provide actionable feedback for improving answer quality.
This paper introduces Deep Persona, a psychologically grounded architecture for role-playing agents, and proposes an evaluation framework. It evaluates LLMs and finds systematic limitations in emotional expression despite high pragmatic fluency.