have you checked out Hark Handoff it has scored better On eval than GPT 5.5 & opus 4.8 at 90% less cost

Reddit r/ArtificialInteligence Models

Summary

Hark Handoff reportedly outperforms GPT-5.5 and Opus 4.8 on several benchmarks at 90% lower cost, using SFT and asynchronous RL with GRPO on an undisclosed base model. The author expresses skepticism about latency in computer-use agents but is bullish on the demo.

97.7 on Online Mine 2Web 83.2 on internal 68.6 on WebTail Bench Best across board and at 2.37 dollars per million token 90% Than GPT5.5 ! how they have trained this. They are using an undisclosed base model and using SFT to accelerate time to market, combined with asynchronous reinforcement learning, especially leveraging the GRPO algorithm. If you don't know, this is similar to how DeepMind historically has trained their AlphaGo Even though they are talking about 1-3 sec latency the huge problem in computer use agents are page rendering and state resolution and there own data showcases it adds roughly 10 Secs so i am skeptical there but I don't think latency matter always and I am bullish on CUA I spend most of the time scrolling the web for silly things, and my mind was blown by the demo videosss Not associated with Any labs. I wish I was :) https://preview.redd.it/wf6n7usifiih1.png?width=812&format=png&auto=webp&s=52b08113306b0737127f0bf725f09cff7da0d485
Original Article

Similar Articles