Newest

All articles, most recently crawled first.

Cards List

Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants

arXiv cs.CL · 6h ago Cached

The paper proposes Evaluation-as-Search (EaS), an adaptive methodology for evaluating grounding failures in LLM-powered meeting assistants, and introduces MeetingProbe, a benchmark of over 3,000 annotated question-answer pairs to improve failure detection.

0 favorites 0 likes

ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research

arXiv cs.CL · 6h ago Cached

This paper introduces ImmigrationReason, a large-scale structured dataset of U.S. immigration appeals for legal reasoning research, addressing the gap in administrative adjudication data for NLP studies.

0 favorites 0 likes

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

arXiv cs.AI · 6h ago Cached

This paper introduces DreamBench-SWE, a benchmark for evaluating memory hygiene in multi-session software agents, and reports experimental results showing its ability to discriminate memory configurations and characterize performance.

0 favorites 0 likes

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

arXiv cs.CL · 6h ago Cached

Ansari is a deployed retrieval-grounded Islamic AI assistant that uses authenticated corpora to answer questions with citations, based on 140,000 conversations and evaluations demonstrating its effectiveness and insights for values-sensitive AI deployments.

0 favorites 0 likes

Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance

arXiv cs.AI · 6h ago Cached

This paper proposes an ontology-driven framework to ensure trustworthiness and auditability in large language model analytics for enterprise financial applications.

0 favorites 0 likes

Intent Engine: Natural-Language Intent Translation for Intent-Driven Orchestration in the Compute Continuum

arXiv cs.CL · 6h ago Cached

Intent Engine is an architecture that translates natural-language intents into validated Service-level Objective artifacts for compute-continuum service placement, using LLMs with retrieval augmentation to reduce hallucination and placement failures.

0 favorites 0 likes

Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions

arXiv cs.AI · 6h ago Cached

The paper presents a multi-criteria framework for evaluating socio-technical interventions, using misinformation as a case study, and reveals trade-offs between effectiveness, user acceptance, and implementation feasibility.

0 favorites 0 likes

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

arXiv cs.CL · 6h ago Cached

Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.

0 favorites 0 likes

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

arXiv cs.AI · 6h ago Cached

The paper introduces the Weighted Memory Tree, a hierarchical memory system for LLM agents that dynamically retains important information, improving task accuracy by 9.97% and reducing prompt token usage by 32.8%.

0 favorites 0 likes

Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal

arXiv cs.CL · 6h ago Cached

This paper explores using human–LLM disagreement to enhance checklist-based quality appraisal in systematic reviews, showing that analyzing disagreements can identify ambiguous items and improve checklist design for better agreement and study ranking preservation.

0 favorites 0 likes

SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL

arXiv cs.AI · 6h ago Cached

SAGE introduces a unified logical and physical framework for AI functions in SQL using three primitives (AI_SCALAR, AI_AGG, AI_JOIN) to optimize execution, significantly reducing model calls and costs.

0 favorites 0 likes

Decoupled Vision-Language System for Multimodal Understanding and Generation

arXiv cs.CL · 6h ago Cached

The paper introduces Libra, a decoupled vision-language architecture for multimodal large language models that enables both image-to-text understanding and text-to-image generation, demonstrating strong performance on benchmarks.

0 favorites 0 likes

Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work

arXiv cs.AI · 6h ago Cached

The paper proposes a harness paradigm for AI agents in large enterprises, focusing on governance and standardization to make AI tools more manageable and compliant.

0 favorites 0 likes

EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators

arXiv cs.CL · 6h ago Cached

EditPPT introduces a multi-agent framework for accurate and faithful slide editing in long decks, using structured tool-using and dual-modal validators, and presents the DeckEdit-Bench benchmark.

0 favorites 0 likes

Dual-Cache Latent Space Communication between Heterogeneous Language Models

arXiv cs.AI · 6h ago Cached

The paper introduces XKV, a method for efficient latent space communication between heterogeneous language models in multi-agent systems, improving accuracy and speed over existing text and cache-based protocols.

0 favorites 0 likes

TH-GNN: Heterogeneous Temporal Graph Neural Networks for LLM-Agent Shilling Attack Detection

arXiv cs.CL · 6h ago Cached

This paper introduces TH-GNN, a heterogeneous temporal graph neural network that fuses graph structure and textual semantics to detect shilling attacks generated by LLM agents in recommender systems, achieving superior performance over existing methods.

0 favorites 0 likes

GRAFT: Adaptive DLM-Based Draft Tree Construction with Target-Distilled Edge Scoring

arXiv cs.CL · 6h ago Cached

GRAFT introduces a draft-tree construction framework for diffusion language model-based speculative decoding, optimizing edge selection and budget allocation to achieve 2.13×–6.36× speedup over autoregressive decoding with low overhead.

0 favorites 0 likes

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

arXiv cs.AI · 6h ago Cached

The paper presents ACES, a framework for continuous evaluation of AI agent skills through live trials, measuring Skill Lift to quantify added value, and demonstrating its effectiveness on enterprise repositories compared to scan-only gates.

0 favorites 0 likes

VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

arXiv cs.CL · 6h ago Cached

This paper proposes VA-DPO, a method for controllable emotion generation in language models using continuous valence-arousal dimensions, which improves over prompting techniques without degrading model performance.

0 favorites 0 likes

Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

arXiv cs.AI · 6h ago Cached

This paper proposes DASO, a tree-aware post-training method for generative recommendation that addresses difficulty mismatch in GRPO by profiling rollout groups and reallocating based on prefix-match depth, improving performance on public benchmarks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback