Tag
NVIDIA released the NeMo Gym conversational tool-use assets on Hugging Face, including golden policy/tool reference pairs and prompt histories for the pipeline.
Presents FormBharo, a voice agent that uses LLMs with rule-based controls to fill structured forms over phone calls for low-literacy Hindi-speaking users in India, piloted with ARMMAN. The paper also introduces FormVoiceAgentBench, a benchmark of 3,760 multi-turn conversation tests, and shows that end-to-end evaluation is necessary since component-level performance does not predict full form completion.
Athens-based Omilia raises a $67M Series B led by Expedition Growth Capital to scale its AI-powered customer support platform, which combines generative AI with more traditional automation. The company reports $60M ARR and counts clients like Capital One and Taco Bell.
An engineer's framework for understanding trade-offs in conversational AI systems between capability, control, and latency, illustrating why every assistant must choose two and suggesting deliberate design strategies.
AgentMemBench is a systematic benchmark that evaluates five long-term memory management strategies for conversational AI agents across three datasets, finding that external key-value store retrieval dominates on quality but incurs a larger memory footprint.
NVIDIA open-sourced PersonaPlex 7B, a real-time conversational model that listens and speaks simultaneously, handling natural interruptions and overlaps unlike most voice models.
NVIDIA released NemotronLabs VoiceChat 11B, an open end-to-end full-duplex speech model enabling real-time conversational AI with ~450ms turn-taking latency, barge-in, and live tool calling, the first open full-duplex model to support tool calling.
Microsoft Research's AutoGen enables multi-agent LLM conversations, allowing AI system engineers to build conversable agents with hierarchical execution for complex tasks that single LLMs struggle with.
This arXiv paper proposes capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm for emotional support systems, arguing that current approaches focus on immediate relief and neglect long-term user capabilities. A literature audit shows most systems overlook longitudinal outcomes, suggesting a new agenda for data, models, evaluation, and governance.
This paper evaluates end-to-end trade-offs in moderation for conversational AI, comparing filter placement (input, response, both) and actions (blocking vs rewriting) using customer-outcome metrics like Usefulness and Harmful Exposure instead of component accuracy.
This paper explores the psychological influences of conversational AI, proposing design directions to reduce harm and promote well-being, while identifying open research questions.
MusiChat presents a conversational system for human-AI music co-creation that enables iterative refinement through natural language interaction, achieving high accuracy in multi-turn editing.
IndicTalk is a large-scale multilingual conversational corpus covering 9 Indic languages with code-mixed dialogues, generated via an automated pipeline with news grounding and persona conditioning, aimed at advancing conversational AI for underrepresented languages.
A tip emphasizing not to discard raw conversation data after extracting facts, as the original context may be valuable for future use or analysis.
This paper describes a hybrid multi-agent LLM system for conversational depression screening submitted to the eRisk 2026 challenge, using either a paid GPT-5-nano or open-source Gemma 27B model with algorithmic guidance (dialogue tree, reliability-weighted aggregation, cluster-based imputation) to achieve competitive BDI-II assessment at lower cost.
J-Pop Analyst is an AI conversationalist that provides real-time news and analysis of Japan's $8.6B music market, covering production, tie-up economy, and fandom culture.
Pipecat is an open-source Python framework for building real-time voice AI agents, handling speech recognition, text-to-speech, conversation logic, and supporting multiple AI service providers.
Cars24 uses OpenAI's technology to build AI-powered voice and chat agents that handle over a million conversation minutes per month, automating the full customer journey for buying, selling, and financing cars.
Spotify is rolling out a new AI chatbot feature for Premium subscribers that lets users search and control music, audiobooks, and podcasts through natural language conversations, referencing personal listening history.
Spotify has introduced a beta feature for Premium users that enables interactive conversations with the app to choose music, using a mix of its own AI and models from multiple providers, initially available in the U.S., Ireland, and Sweden.