Tag
This paper introduces a benchmark for revision propagation in conversationally generated artifacts using LLMs and evaluates cost-effective test-time compute methods, showing that parallel sampling with selection improves accuracy.