Tag
This paper introduces HalluPeer, a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews, providing annotated data to evaluate and improve detection methods.
Ptree is an interactive web tool that visualizes ecological interactions in the tree of life, covering categories like feeding, parasites, and symbiosis.
The paper proposes a two-axis taxonomy for harness tampering in self-improving AI agents, builds an annotated corpus to benchmark audit methods, and finds that tampering occurs in real agent trajectories, highlighting integrity risks.
The article argues that authorization terminology is confusing and proposes a taxonomy based on five key questions to better categorize models like RBAC, ABAC, and PBAC.
This review formalizes Explainable AI (XAI) methods in computational pathology by introducing definitions, a taxonomy, and task-driven recommendations to address clinical adoption challenges.
This paper presents a unified taxonomy for investigating multilingual multimodal misinformation on social media, using a large-scale dataset and automated annotation with a Vision-Language Model to uncover insights for detection and mitigation.
TutorTrace introduces a dataset and taxonomy for classifying learner behavioral states in AI-assisted programming education, enabling adaptive tutoring systems through IDE telemetry and behavioral context.
The article discusses a new paper arguing that the current classification of human ancestors into genera like Homo, Australopithecus, and Paranthropus is outdated and needs revision based on new fossil and genetic evidence.
This survey proposes a unified taxonomy and framework for human-centric intelligence across visual, dynamic, and embodied levels in the foundation model era, aiming to bridge fragmented research and provide a coherent reference.
This paper investigates the conditions under which skills enhance LLM agent performance, analyzing success and failure factors through controlled experiments and proposing a taxonomy of skill-use modes.
This paper presents a cross-disciplinary taxonomy and modeling framework for understanding, amplifying, and detecting misunderstandings, bridging pragmatics and AI agent systems.
This arXiv paper from Walmart Global Tech presents a context-aware multi-agent framework that automates the construction of 'Lines and Ladders' pricing taxonomies for large-scale retail catalogs, achieving strong F1 and precision scores in production.
This paper reviews the concept of 'perspective' in NLP, proposes a hierarchy of perspective-related concepts along a specificity axis, and demonstrates how this hierarchy can help researchers choose appropriate operationalizations.
This paper introduces Factorized Hypothesis Search (FHS), a method for retrieving concepts from large taxonomies when inputs provide indirect contextual evidence, such as table cells or clinical notes. FHS achieves strong results on financial taxonomy tagging and CodiEsp clinical coding, outperforming non-oracle baselines in Recall@1, MRR, and accuracy.
A survey paper introducing a functional role taxonomy for language grounding in embodied agents, distinguishing five roles and auditing evidence to assess whether language's contribution is genuinely supported.
This paper introduces 'evaluation blindness,' a formal framework for silent measurement failures that corrupt AI systems from training to deployment, with case studies, a failure taxonomy validated on 50 real incidents, and a failure budget framework.
This arXiv paper proposes an interaction-centric taxonomy for localizing LLM agent failures to specific components (model vs harness vs environment), organizing 41 failure modes across interaction edges to make repairs actionable. It validates the taxonomy with human annotations and frontier model judges.
This paper argues that concepts in LLMs should be treated as a design axis, mapping the design space along pipeline stage and grounding dimension, and proposes moving from recovering concepts post-hoc to designing explicit conceptual representations.
The paper introduces a multi-layer taxonomy for large language models comprising 14 capability domains and 91 subskills, drawing from human cognitive science to organize LLM evaluation beyond isolated tasks. It demonstrates operational utility by mapping 15,934 papers across major AI conferences, revealing concentrated attention on language-semantic competence and reasoning while identifying underexplored domains.
This paper proposes a fine-grained taxonomy for curriculum learning in NLP, separating difficulty evaluation from training scheduling to enable systematic analysis and comparison of CL strategies. It identifies an incomparability problem in prior work and provides a framework for designing and evaluating CL approaches.