conversational-agents

Tag

Cards List
#conversational-agents

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

arXiv cs.CL · 5d ago Cached

This paper benchmarks the serving cost of three agentic memory systems (Mem0, Hindsight, Mastra Observational Memory) against reference strategies across conversational backbones, finding that cost is driven by internal memory behavior, break-even points vary widely, and no system wins on both cost and accuracy.

0 favorites 0 likes
#conversational-agents

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv cs.AI · 2026-07-15 Cached

This paper introduces GenAI Evaluation, a governed and configuration-driven pipeline for scalable multi-dimensional evaluation of retail conversational agents. It achieves high accuracy using LLM-as-a-judge scoring with selective re-evaluation, validated against human-labeled data.

0 favorites 0 likes
#conversational-agents

ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration

arXiv cs.CL · 2026-07-07 Cached

ProACT introduces a breakdown-aware agent framework for multi-user collaboration, where the agent proactively detects collaboration breakdowns and decides whether to intervene. The paper also presents the first multi-user collaboration benchmark for evaluating such proactive agents.

0 favorites 0 likes
#conversational-agents

Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

arXiv cs.CL · 2026-06-25 Cached

This paper introduces a taxonomy of conversational memory types and a user-centric evaluation framework to study how different memory roles affect response quality in RAG-based conversational agents.

0 favorites 0 likes
#conversational-agents

Recently hired by an award winning SaaS. Documenting the agent runtime businesses actually need

Reddit r/AI_Agents · 2026-06-22

A developer documents the architecture of an AI agent runtime built for a SaaS company, focusing on safety, tool execution, state management, and separation of reasoning from execution.

0 favorites 0 likes
#conversational-agents

@ElevenLabsDevs: Call your Hermes Agent

X AI KOLs Following · 2026-06-04

ElevenLabs introduces the ability to call your Hermes Agent, enabling voice-based interaction with AI agents through their platform.

0 favorites 0 likes
#conversational-agents

SaliMory: Orchestrating Cognitive Memory for Conversational Agents

arXiv cs.CL · 2026-06-04 Cached

SaliMory is a framework that trains a single language model to manage cognitively-structured memory (user facts, preferences, and working memory) for conversational agents, using hierarchical stage-wise process rewards and reward-decomposed contrastive refinement. It reduces memory-attributed failures by one-third, outperforms state-of-the-art by over 10% in end-to-end accuracy, and more than doubles the Good Personalization rate.

0 favorites 0 likes
#conversational-agents

Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents

arXiv cs.CL · 2026-05-26 Cached

Proposes Structure-Aware RAG (SA-RAG), which uses tables as an intermediate structured representation to reduce noise in retrieval-augmented generation for conversational agents, with quality-aware metadata generation and two table generation methods, outperforming existing baselines on noisy real-world datasets.

0 favorites 0 likes
#conversational-agents

Evaluating multimodal emotion recognition in proactive conversational agents: A user study

arXiv cs.AI · 2026-05-22 Cached

This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.

0 favorites 0 likes
#conversational-agents

When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models

arXiv cs.CL · 2026-05-08 Cached

When2Speak is a synthetic dataset and pipeline for training LLMs to decide when to speak in multi-party conversations. Fine-tuning on this dataset significantly improves turn-taking, with reinforcement learning reducing missed interventions from 50% to ~20%.

0 favorites 0 likes
#conversational-agents

Ecom-RLVE: Adaptive Verifiable Environments for E-Commerce Conversational Agents

Hugging Face Blog · 2026-04-16 Cached

Huggingface introduces EcomRLVE-GYM, a framework providing eight verifiable environments for training reinforcement learning agents on complex e-commerce tasks. The tool features adaptive difficulty curricula and algorithmic rewards to improve task completion in shopping assistants, demonstrated by training a Qwen 3 8B model.

0 favorites 0 likes
← Back to home

Submit Feedback