Tag
A benchmark of three AI agents on 12 multi-app tasks shows Kimi K3 tied the most expensive model GPT-5.6 Sol at a fraction of the cost, though all three failed cross-app reconcile tasks, highlighting the need for verification in production.