webshop

Tag

Cards List
#webshop

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

arXiv cs.AI · 2026-05-20 Cached

This paper presents the first systematic study of credit assignment in multi-turn LLM agents, introducing SERL, a selective environment-reweighted learning framework. SERL uses environment feedback to sharpen the RL objective on causally relevant actions, achieving 90.0% and 80.1% success rates on ALFWorld and WebShop respectively.

0 favorites 0 likes
← Back to home

Submit Feedback