All articles, most recently crawled first.
The paper proposes Evaluation-as-Search (EaS), an adaptive methodology for evaluating grounding failures in LLM-powered meeting assistants, and introduces MeetingProbe, a benchmark of over 3,000 annotated question-answer pairs to improve failure detection.
This paper introduces ImmigrationReason, a large-scale structured dataset of U.S. immigration appeals for legal reasoning research, addressing the gap in administrative adjudication data for NLP studies.
This paper introduces DreamBench-SWE, a benchmark for evaluating memory hygiene in multi-session software agents, and reports experimental results showing its ability to discriminate memory configurations and characterize performance.
Ansari is a deployed retrieval-grounded Islamic AI assistant that uses authenticated corpora to answer questions with citations, based on 140,000 conversations and evaluations demonstrating its effectiveness and insights for values-sensitive AI deployments.
This paper proposes an ontology-driven framework to ensure trustworthiness and auditability in large language model analytics for enterprise financial applications.
Intent Engine is an architecture that translates natural-language intents into validated Service-level Objective artifacts for compute-continuum service placement, using LLMs with retrieval augmentation to reduce hallucination and placement failures.
The paper presents a multi-criteria framework for evaluating socio-technical interventions, using misinformation as a case study, and reveals trade-offs between effectiveness, user acceptance, and implementation feasibility.
Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.
The paper introduces the Weighted Memory Tree, a hierarchical memory system for LLM agents that dynamically retains important information, improving task accuracy by 9.97% and reducing prompt token usage by 32.8%.
This paper explores using human–LLM disagreement to enhance checklist-based quality appraisal in systematic reviews, showing that analyzing disagreements can identify ambiguous items and improve checklist design for better agreement and study ranking preservation.
SAGE introduces a unified logical and physical framework for AI functions in SQL using three primitives (AI_SCALAR, AI_AGG, AI_JOIN) to optimize execution, significantly reducing model calls and costs.
The paper introduces Libra, a decoupled vision-language architecture for multimodal large language models that enables both image-to-text understanding and text-to-image generation, demonstrating strong performance on benchmarks.
The paper proposes a harness paradigm for AI agents in large enterprises, focusing on governance and standardization to make AI tools more manageable and compliant.
EditPPT introduces a multi-agent framework for accurate and faithful slide editing in long decks, using structured tool-using and dual-modal validators, and presents the DeckEdit-Bench benchmark.
The paper introduces XKV, a method for efficient latent space communication between heterogeneous language models in multi-agent systems, improving accuracy and speed over existing text and cache-based protocols.
This paper introduces TH-GNN, a heterogeneous temporal graph neural network that fuses graph structure and textual semantics to detect shilling attacks generated by LLM agents in recommender systems, achieving superior performance over existing methods.
GRAFT introduces a draft-tree construction framework for diffusion language model-based speculative decoding, optimizing edge selection and budget allocation to achieve 2.13×–6.36× speedup over autoregressive decoding with low overhead.
The paper presents ACES, a framework for continuous evaluation of AI agent skills through live trials, measuring Skill Lift to quantify added value, and demonstrating its effectiveness on enterprise repositories compared to scan-only gates.
This paper proposes VA-DPO, a method for controllable emotion generation in language models using continuous valence-arousal dimensions, which improves over prompting techniques without degrading model performance.
This paper proposes DASO, a tree-aware post-training method for generative recommendation that addresses difficulty mismatch in GRPO by profiling rollout groups and reallocating based on prefix-match depth, improving performance on public benchmarks.