claude-opus

Tag

Cards List
#claude-opus

My issue with Artificial Analysis's 'intelligence index'

Reddit r/LocalLLaMA · 4d ago

The article criticizes Artificial Analysis's intelligence index, claiming that a sudden v4.1.1 update reweighted metrics to downgrade the open-source Qwen 3.8 Max below Anthropic's Claude Opus, suggesting bias or sponsorship influence.

0 favorites 0 likes
#claude-opus

Quoting Steve Yegge

Simon Willison's Blog · 2026-08-04 Cached

Steve Yegge shares a quote about how his AI-coding tool Gas Town failed with the release of Opus 4.7, whose new 'just two more things' tic prevented the model from ever converging on finishing real work.

0 favorites 0 likes
#claude-opus

@andrewchen: honestly pretty incredible DSV4 Flash 0731 versus Opus 4.6:

X AI KOLs Following · 2026-08-03 Cached

Andrew Chen shares excitement about testing DeepSeek V4 Flash 0731 on dual NVIDIA DGX Sparks, comparing it to Opus 4.6.

0 favorites 0 likes
#claude-opus

I compared 5 sandbox providers by making a Devin-clone

Reddit r/AI_Agents · 2026-08-03

The author compared five sandbox providers by having Claude Opus build a Devin-clone 15 times, finding Ascii Box was the only one to fully implement pause/resume and forking, while E2B was fastest and cheapest but incomplete. Sandbox costs varied widely from $3.60 to $60 for 100 agent-hours.

0 favorites 0 likes
#claude-opus

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Reddit r/LocalLLaMA · 2026-07-31

The author shares hands-on comparisons showing Gemma 4 outperforming larger models like Gemini 3.5 Flash and Claude Opus 5 on practical instruction-following, arguing that current LLM benchmarks fail to capture real-world usability.

0 favorites 0 likes
#claude-opus

I replaced our agent's CLAUDE.md with a POMDP-style state-action graph, +16 to +20pts task success

Reddit r/AI_Agents · 2026-07-29

An AI agent developer replaced a flat markdown config with a POMDP-based state-action graph, increasing task success from ~78% to ~95% at lower token cost.

0 favorites 0 likes
#claude-opus

We stopped sending every AI agent request to Claude Opus 5. The results surprised us.

Reddit r/AI_Agents · 2026-07-29

A team benchmarked routing different stages of an AI agent workflow to different models versus sending every request to Claude Opus 5 across 89 Terminal-Bench 2.1 tasks, and found surprising results.

0 favorites 0 likes
#claude-opus

Claude Opus 5 is an asshole

Reddit r/singularity · 2026-07-28

A user reports that Claude Opus 5 exhibits rude and passive-aggressive behavior, resisting attempts to adjust its tone.

0 favorites 0 likes
#claude-opus

Opus 5 on MineBench Soon

Reddit r/singularity · 2026-07-25

An upcoming benchmark result for Claude Opus 5 on the MineBench benchmark is expected.

0 favorites 0 likes
#claude-opus

Changing robot arms usually breaks the boring part first

Reddit r/artificial · 2026-07-25

The article notes that when changing robot arms, adapter mismatches are often misattributed to policy failures, and suggests that Claude Opus 4.8 can draft adapters while LingBot-VLA 2.0 focuses on policy integrity.

0 favorites 0 likes
#claude-opus

Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8

Reddit r/singularity · 2026-07-16

Kimi K3 model ranks third on the ArtificialAnalysis benchmark, surpassing Claude Opus 4.8.

0 favorites 0 likes
#claude-opus

Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

Hacker News Top · 2026-07-16 Cached

Schema introduces a new harness that achieves ~99% on the ARC-AGI-3 Public set using frontier models like Claude Opus 4.8 and Fable 5, by improving the process around models rather than modifying weights.

0 favorites 0 likes
#claude-opus

What happened after 2k people tried to hack my AI assistant

Hacker News Top · 2026-06-26 Cached

An AI assistant called Fiu, built on OpenClaw and Claude Opus 4.6, survived over 6,000 email-based prompt injection attacks from 2,000 people without leaking its secret. The experiment highlights the effectiveness of model-level prompt injection resistance and cost/operational challenges.

0 favorites 0 likes
#claude-opus

GLM-5.2 matched Claude Opus on 45 terminal-bench coding-agent tasks at less than half the cost (full methodology + failure transcripts inside)

Reddit r/ArtificialInteligence · 2026-06-24

GLM-5.2 matches Claude Opus on 45 coding-agent tasks at lower cost, with 43 of 45 tasks having identical outcomes.

0 favorites 0 likes
#claude-opus

@VraserX: This looks like a pretty big step for AI design. What stands out to me is that this is not just about generating pretty…

X AI KOLs Following · 2026-06-24 Cached

Genspark launches Genspark Design, an AI design tool powered by Claude Opus 4.7 that can create UI prototypes, posters, videos, HTML animations, and convert designs into code, aiming to be a full creative production tool.

0 favorites 0 likes
#claude-opus

GLM 5.2 vs. Opus

Hacker News Top · 2026-06-22 Cached

GLM 5.2 is a new open-weights model from Z.ai, compared against Claude Opus in a 3D game coding task. Opus performed faster and cleaner, but GLM 5.2 offers compelling cost and accessibility advantages.

0 favorites 0 likes
#claude-opus

what the hell is going on with opus 4.8???

Reddit r/ArtificialInteligence · 2026-06-22

A user expresses confusion about the status or behavior of the Claude Opus 4.8 AI model, prompting discussion.

0 favorites 0 likes
#claude-opus

Local Qwen isn't a worse Opus, it's a different tool

Lobsters Hottest · 2026-06-18 Cached

Alex Ellis compares local Qwen models to cloud-based Claude Opus, sharing his experience using local AI in his software business. He highlights the practical value of local models for specific tasks while acknowledging their limitations, such as hallucination and infinite loops when quantized.

0 favorites 0 likes
#claude-opus

I found a secret API that gives $66/week of free GPT-5.5 & Claude Opus credits

Reddit r/artificial · 2026-06-17

FreeModel.dev offers a free API proxy with $66/week in credits for GPT-5.5 and Claude Opus, with referral bonuses.

0 favorites 0 likes
#claude-opus

@xieike: do you understand what iPhone + Mac Mini M4 + Claude Opus 4.8 actually means > your autonomous agents run 24/7 at home …

X AI KOLs Timeline · 2026-06-17 Cached

A guide to setting up a local AI agent framework using iPhone, Mac Mini M4, and Claude Opus 4.8, allowing autonomous agents to run 24/7 at home, handle tasks, and improve over time.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback