Tag
The author shares a deterministic learning harness that lets multi-agent systems improve across episodes without fine-tuning or prompt edits, by promoting successful strategies into persistent playbooks. On the Mini Amusement Park benchmark, reward improved from 12,121 to 483,019, reaching #1 on the leaderboard.
MOTHRAG is a multi-hop RAG system that matches the performance of top GPU-dependent systems (HippoRAG 2, CoRAG, NeocorRAG) using only commodity API calls, with no GPU, no fine-tuning, and deployment via pip install plus API keys.
Poetiq's Meta-System, using recursive self-improvement via standard API access without fine-tuning, achieves new state-of-the-art results on the LiveCodeBench Pro coding benchmark, outperforming leading models like GPT 5.5.