cost-analysis

Tag

Cards List
#cost-analysis

A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.

Reddit r/AI_Agents ↗ · yesterday

A developer built a custom benchmark for coding AI agents and found that a 35B-parameter model outperformed a 120B-parameter one when the harness was optimized, highlighting the importance of tailored evaluation over generic specifications.

0 favorites 0 likes
#cost-analysis

This is fun. I finally got to follow up on a RemindMe comment. Back on March 25 of this year, six months ago, no model was getting even 1% on ARC-AGI-3. A commenter asked if we could see 75% at $2 cost. Well, the cost is still high ($26.1k), but GPT-6-Astra-Max was able to get 62.7% (no harness!)

Reddit r/singularity ↗ · 2d ago

GPT-6-Astra-Max achieved 62.7% on the ARC-AGI-3 benchmark within six months, and with a memory adapter, the benchmark is saturated, though costs remain high, indicating a trend toward cheaper AI models.

0 favorites 0 likes
#cost-analysis

@omarsar0: Interesting results here. This is why I expect more agent workloads to run on blended models. Pareto 26.9 from @TheUnbi…

X AI KOLs Timeline ↗ · 2d ago Cached

The article discusses a performance evaluation where Pareto 26.9, a blended AI model, ties with GPT-6 Astra in agent tasks at one-third the cost and faster completion than other models, suggesting potential for agent workloads on blended models.

0 favorites 0 likes
#cost-analysis

@theo_the_dev: BREAKING: Opus 5.5 on Medium just CRUSHED GPT-6 Astra Ultra. $21,365 in tokens for @threejs game. Just compare these tw…

X AI KOLs Following ↗ · 3d ago Cached

A tweet claims that Opus 5.5 outperformed GPT-6 Astra Ultra in a three.js game development scenario, involving significant token costs and compute time, with human involvement.

0 favorites 0 likes
#cost-analysis

This interactive island was built in 8 hours with Opus 5.5

Reddit r/singularity ↗ · 3d ago

Dan Greenheck built an interactive island in 8 hours using Opus 5.5 with simple prompts, featuring animations and effects, at a token cost of $1,874.40.

0 favorites 0 likes
#cost-analysis

@mattshumer_: Jacob sent me a screenshot of the cost of a Claude run he did the other day. I almost feel bad saying just how much it …

X AI KOLs Following ↗ · 3d ago Cached

A user shares the high cost of using 26 Claude Opus 5.5 agents to build a complex multiplayer game overnight, demonstrating AI's application in game development.

0 favorites 0 likes
#cost-analysis

@paulg: Paweł Huryn tested models' ability to find bugs planted in code. Cost increases exponentially with performance (note th…

X AI KOLs Following ↗ · 4d ago Cached

Paweł Huryn tested AI models' ability to find planted bugs in code, revealing that performance improvements come with exponentially increasing costs, as shown in a chart with a log scale.

0 favorites 0 likes
#cost-analysis

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Hugging Face Daily Papers ↗ · 4d ago Cached

StudentBench study finds that AI tutoring yields learning gains equivalent to human tutoring for GRE questions, with one AI tutor achieving similar results at a significantly lower cost.

0 favorites 0 likes
#cost-analysis

Opus 5.5 dominates on all three performance metrics from ArtificialAnalysis.ai

Reddit r/artificial ↗ · 4d ago

Opus 5.5 dominates all three performance metrics from ArtificialAnalysis.ai and offers a cost-effective alternative to Fable 5.1.

0 favorites 0 likes
#cost-analysis

299 real user intents tested Jev against production base line. Here is the result.

Reddit r/AI_Agents ↗ · 4d ago

The article reports on a performance comparison between glm-4-flash and TypeSafe Jev on 299 real user intents, showing glm-4-flash's higher accuracy but TypeSafe Jev's faster speed and lower cost.

0 favorites 0 likes
#cost-analysis

Every Benchmark Comparison for Opus 5.5 and GPT-6 Astra/Sol

Reddit r/singularity ↗ · 4d ago

Opus 5.5 outperforms GPT-6 Astra and Sol in benchmark comparisons, but at a higher cost, as shown in the provided image.

0 favorites 0 likes
#cost-analysis

Opus 5.5 Cost vs Performance on Terminal-Bench 4.0

Reddit r/singularity ↗ · 5d ago

This article evaluates the cost-effectiveness and performance of the Opus 5.5 AI model on the Terminal-Bench 4.0 benchmark.

0 favorites 0 likes
#cost-analysis

@ProfTomYeh: Can you solve these agentic AI math problems by hand ? Getting a bit harder now. Download PDF: http://byhand.ai/tokens-…

X AI KOLs Timeline ↗ · 5d ago Cached

Prof. Tom Yeh shares a PDF with math problems on agentic AI topics like token costs and system prompts, encouraging manual solving for learning.

0 favorites 0 likes
#cost-analysis

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

arXiv cs.CL ↗ · 5d ago Cached

This paper empirically analyzes cost savings in context-compression gateways for multi-turn coding agents, revealing that tool-schema filtering provides fixed token savings, while content compression saves quadratically but can be offset by recalls, offering actionable insights for cost optimization.

0 favorites 0 likes
#cost-analysis

Swarm Scaling (13 minute read)

TLDR AI ↗ · 5d ago Cached

This article analyzes the capabilities and scaling dynamics of large AI agent swarms, citing OpenAI's recent examples, and discusses their potential as a new form of inference scaling.

0 favorites 0 likes
#cost-analysis

607 AI agent/use cases I found after Jev dropped

Reddit r/AI_Agents ↗ · 5d ago

The author collected 607 AI agent use cases after Jev's release, highlighting distinctions between demos and production-ready applications, and built ShipWithJev.com to catalog them.

0 favorites 0 likes
#cost-analysis

AI Doesn’t Live in the Cloud. It Lives on Earth.

Reddit r/artificial ↗ · 6d ago

The article critiques the oversight of AI's tangible environmental impact, emphasizing that AI relies on physical resources like energy, water, and hardware, and calls for integrating planetary stewardship into AI development.

0 favorites 0 likes
#cost-analysis

@10xmylife: Reply to the questions in the comment section - How do you read the game state? We didn't actually have Jev look at scr…

X AI KOLs Timeline ↗ · 2026-09-20 Cached

A developer explains using an AI agent named Jev to play Slay the Spire 2, employing a C# Mod to extract game data and a Python API for decision-making, while noting high token costs and mediocre performance.

0 favorites 0 likes
#cost-analysis

@corbin_braun: so logically the AI editor does take some time, maybe 5-6 hours a video. cost wise its like 30 USD of tokens, compare t…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

A tweet compares the time and cost of using an AI video editor, which takes 5-6 hours and $30 in tokens per video, to a human editor who would take 3-5 days and cost $250, noting the AI can handle any edit style.

0 favorites 0 likes
#cost-analysis

@yibie: Jev Ecosystem 72 Hours: From 46 to 160 Jev turned "judgment" into a primitive that's cheap enough to call directly in c…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

In 72 hours, the Jev ecosystem expanded from 46 to 160 projects, focusing on context compression, platform integrations, and new domains like financial trading, with debates on its novelty and implementation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback