Tag
Google has reduced prices for Gemini 3.7 Flash on OpenRouter by 75%, making it cheaper and outperforming DeepSeek models in price/performance based on Artificial Analysis.
A benchmark test of nine agentic web search tools using Claude Opus 4.8 agent found that lower-cost tools often outperformed more expensive ones, with Firecrawl ranking first for accuracy and Serper being the most cost-effective.
Databricks published an internal benchmark evaluating coding agents on their multi-million line codebase, revealing that harness choice can double cost savings and that open models like GLM 5.2 perform competitively at the highest difficulty levels.
Databricks shares results from an internal benchmark evaluating coding agents on their multi-million line codebase, revealing capability tiers and cost-performance tradeoffs, and highlighting the effectiveness of open models like GLM 5.2.
Compares the performance and cost of four AI models (Fable 5, GPT 5.5, Opus 4.8, GLM 5.2) on three HTML5 physics scene tasks. Fable 5 delivers the best quality but costs nearly 6x more than Opus; quality and price are not yet both attainable.
This paper introduces The Efficiency Frontier, a unified framework for cost–performance optimization in LLM context management that models context strategy selection as a deployment-aware optimization problem, achieving 25% reduction in token usage and over 50% lower token cost with amortized memory compression compared to full-context prompting.
A controlled study of compound LLM agent design in an adversarial POMDP (CybORG CAGE-2), systematically varying context, reasoning, and hierarchy across five model families. Key findings: programmatic state abstraction yields large returns per token, hierarchy without deliberation tools achieves best absolute performance, and context engineering is more cost-effective than deeper reasoning.
Benchmarks seven foundation models on Ukrainian legal text, finding tokenizer fertility varies 1.6×, few-shot prompting degrades performance, and cost-performance analysis shows NVIDIA Nemotron Super 3 outperforms larger models.