user-simulators

Tag

Cards List
#user-simulators

KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

arXiv cs.AI · 18h ago Cached

This paper presents KnowSim, an evaluation framework that models user knowledge states to assess information calibration in LLM assistants, validated against human judgments and outperforming baseline simulators.

0 favorites 0 likes
#user-simulators

Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents

arXiv cs.AI · 2026-05-14 Cached

Introduces Persona Policies (PPol), a plug-and-play control layer that uses LLM-driven evolutionary program search to generate diverse, human-like user personas for evaluating LLM agents. Achieves 33–62% fitness gains over baseline, with human-likeness rated at 80.4%, and improves agent robustness with +17% task success.

0 favorites 0 likes
← Back to home

Submit Feedback