have you checked out Hark Handoff it has scored better On eval than GPT 5.5 & opus 4.8 at 90% less cost
Summary
Hark Handoff reportedly outperforms GPT-5.5 and Opus 4.8 on several benchmarks at 90% lower cost, using SFT and asynchronous RL with GRPO on an undisclosed base model. The author expresses skepticism about latency in computer-use agents but is bullish on the demo.
Similar Articles
@PrajwalTomar_: A SMALLER TECH COMPANY just surpassed Claude and GPT 73% on OSWorld 2.0. Above Opus 5. Above GPT-5.6 Sol. At about two-…
Simular's computer-use agent sai_borg achieved 73% accuracy on the OSWorld 2.0 benchmark, surpassing Claude and GPT models at about two-thirds the cost per task, with all trajectories made public.
@orca_build: Anthropic’s new Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1… …but it’s noticeably better at UI tasks.…
Anthropic's Opus 4.8 scores 3.6% lower than GPT 5.5 on Terminal-Bench 2.1 but excels at UI tasks; Orca's orchestration enables Codex to delegate UI tasks to Claude Code.
@sashimikun_void: GPT-5.5 outperformed Claude Opus 4.8 on the DEEPSWE benchmark. Opus 4.8 takes twice as long, generates three times the …
GPT-5.5 outperforms Claude Opus 4.8 on the DEEPSWE benchmark, achieving higher scores with lower cost and less token bloat.
Tokens were eating my wallets, so i tried GLM5.2 and compare it to Fable 5 and Opus 4.8
A developer compares GLM 5.2 with Fable 5 and Opus 4.8 on OpenRouter, finding GLM 5.2 cost-effective for code audits and reviews but not for long unattended builds due to accumulating test-error-edit costs.
@dhh: I've been driving GPT5.5 on low reasoning for the last week+ and it's very good, very efficient. Haven't been tempted t…
DHH praises the performance and efficiency of GPT-5.5 on low reasoning settings, noting it surpasses Opus and Kimi.