realistic-evaluations

Tag

Cards List
#realistic-evaluations

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

arXiv cs.CL · 2026-05-27 Cached

This paper proposes LURE (Live-Usage Replay Evaluations), a method for constructing realistic, deployment-like evaluations of large language models by replaying real agentic interaction trajectories and appending evaluation prompts, reducing the detectability of evaluations compared to existing benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback