Tag
Built a customer-support AI agent named SupportMemory using Hindsight for persistent memory to recall and retain customer context, improving personalization in interactions.
TurnSight introduces a turn-level hindsight self-distillation framework for tool-integrated reasoning, providing dense supervision via execution-conditioned hindsight and adaptive RL advantage modulation.
Introduces Hindsight Policy Optimization (HPO), a novel policy gradient method that uses an intent space and Wasserstein distance to reduce variance in long-horizon language agent training, showing improved stability over GRPO and PPO.
MeetMemory is a tool that gives AI permanent memory for meetings, built on Hindsight and Groq, allowing instant recall across all past conversations without manual search or note-taking.
HERO introduces a hindsight-enhanced self-distillation framework that uses environment observations as locally aligned feedback to improve multi-turn agent capabilities, outperforming existing methods on TauBench and WebShop, especially under limited turn budgets.
HINT-SD proposes a targeted self-distillation framework that selects failure-relevant actions from full trajectories to improve long-horizon LLM agent training, achieving up to 18.80% improvement and 2.26× speedup over dense feedback baselines.