@omarsar0: Interesting results here. This is why I expect more agent workloads to run on blended models. Pareto 26.9 from @TheUnbi…

X AI KOLs Timeline News

Summary

The article discusses a performance evaluation where Pareto 26.9, a blended AI model, ties with GPT-6 Astra in agent tasks at one-third the cost and faster completion than other models, suggesting potential for agent workloads on blended models.

Interesting results here. This is why I expect more agent workloads to run on blended models. Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer. In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task. It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.
Original Article
View Cached Full Text

Cached at: 09/25/26, 12:33 AM

Interesting results here. This is why I expect more agent workloads to run on blended models.

Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer.

In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task.

It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.

Composio (@composio): We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash.

Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task.

Here’s how all 6 models compared 🧵🧵🧵

Similar Articles

This is fun. I finally got to follow up on a RemindMe comment. Back on March 25 of this year, six months ago, no model was getting even 1% on ARC-AGI-3. A commenter asked if we could see 75% at $2 cost. Well, the cost is still high ($26.1k), but GPT-6-Astra-Max was able to get 62.7% (no harness!)

Reddit r/singularity

GPT-6-Astra-Max achieved 62.7% on the ARC-AGI-3 benchmark within six months, and with a memory adapter, the benchmark is saturated, though costs remain high, indicating a trend toward cheaper AI models.