Tag
Sierra announces a strategic partnership with SoftBank to deliver its conversational AI platform to Japanese enterprises, with SoftBank serving as exclusive sales partner and adopting the platform across its own brands.
A comparison of AI voice assistants ChatGPT-Live, Pi, Lucy OS1, and Gemini-Live focusing on which feels most natural to talk with, concluding that conversational quality is becoming the key differentiator as intelligence improves.
This paper analyzes 2,053 real patient-chatbot conversations to show that communication styles vary widely and can significantly alter triage outcomes, finding that patient simulators that model emotional state and conversational strategy produce conversations nearly indistinguishable from real ones in a Turing test.
OpenAI released new full-duplex voice models GPT-Live-1 and GPT-Live-1 mini for more natural live conversations, allowing simultaneous speaking and listening, with improvements in turn-taking and context handling, and replacing Advanced Voice Mode in ChatGPT.
OpenAI releases GPT-Live-1, an upgraded voice mode for ChatGPT that interrupts less, allows real-time translation, and can generate visual context for topics like weather and sports.
NapMem is a framework that treats long-term user memory as a structured action space rather than passive retrieval, using a multi-granularity memory pyramid and reinforcement learning to train agents to navigate memory. Experiments show competitive performance on memory-intensive tasks.
A developer shares techniques for making LLM-generated podcasts sound natural, including using constraints to force disagreement and pre-editing content before generation.
Proposes a Proactive Thinking framework that allows LLMs to pre-compute response elements during conversational pauses, improving interaction efficiency without sacrificing quality. Introduces a training-free baseline that speculatively anticipates future states, evaluated on time-aware benchmarks.
A new tool lets you create a custom conversational AI podcast with two hosts that you can interrupt to ask questions in real time.
A tweet highlights that prompt engineering remains relevant, recommending a read on how to enrich AI agent interactions through techniques like brainstorming and planning.
Google's Gemini Omni Flash model can edit videos through conversational prompts, using the Interactions API to generate new clips based on user descriptions.
This paper introduces Profile-guided Personalized Retrieval Optimization (PPRO), a framework that enhances long-term conversational agents by incorporating user profiles into memory retrieval and optimizing retrieval via reinforcement learning, achieving consistent improvements over existing methods.
This paper presents TRACE, a query processing framework that models conversational data as temporal evidence graphs to enable state-aware reasoning over evolving user states, improving temporal and multi-hop reasoning for long-conversation QA.
This paper proposes a reference-based evaluation protocol for assessing prosody and rhythm in speech-to-speech AI systems, using matched human conversation data to provide interpretable behavioral plausibility checks.
IMCBench is a new benchmark for evaluating multimodal LLMs on image-grounded medical conversations, pairing clinical images with synthetic patient profiles. Evaluations across safety, accuracy, and uncertainty show that even strong models like Claude Opus 4.6 have safety issues, highlighting the need for multi-dimensional evaluation.
Presents a systematic methodology for converting Hindi WordNet into 1.25 million instruction-response pairs to fine-tune a 12B-parameter language model using LoRA, demonstrating improved pedagogical effectiveness for specialized conversational systems in low-resource languages.
This paper explores using Nonviolent Communication (NVC) principles as lightweight prompt constraints to reduce conversational escalation in LLMs during conflict-prone interactions. Experiments across multiple instruction-tuned models show that NVC-constrained prompting consistently de-escalates dialogue and stabilizes interactions with highly resistant users.
This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.
Coval announces a Series A funding round to build infrastructure for testing and deploying conversational AI agents in enterprises, inspired by the rigor of autonomous vehicle testing.
This paper introduces a conversational voice agent system that uses a lightweight on-device 'Talker' model to start responding immediately, then incorporates knowledge from a frontier LLM 'Reasoner' as it becomes available, achieving 7-19x faster time-to-first-response while approaching frontier-level performance on a laptop.