model-performance

Tag

Cards List
#model-performance

Even Fable 5 is losing money in Andon Market (fully AI-operated retail store in San Francisco)

Reddit r/singularity · 2026-08-16

A real-world experiment in San Francisco shows that all Claude AI models, including Fable 5, are operating at a loss in the AI-run Andon Market, with thousands of dollars lost over varying periods.

0 favorites 0 likes
#model-performance

Big Pickle on SWE Atlas – Codebase QnA

Hacker News Top · 2026-08-16 Cached

Big-pickle, a free stealth AI model, achieved a 50.8% resolve rate on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold, outperforming other models in its class.

0 favorites 0 likes
#model-performance

@FinanceYF5: Kimi K3, which was all over the feed just days ago, has now been overtaken — Claude Opus 5 (Max reasoning) storms into Frontend Code Arena and Text Arena, taking first place in both. The default high-reasoning version is also impressive, ranking 3rd in Code Arena behind Kimi…

X AI KOLs Following · 2026-07-28 Cached

Claude Opus 5 (Max reasoning) surpasses Kimi K3 in Frontend Code Arena and Text Arena, claiming first place in both, while the default high-reasoning version also performs well.

0 favorites 0 likes
#model-performance

Opus 5's effort dial is not monotonic. Above "high", coding scores go down, and Anthropic's own migration guide says so.

Reddit r/artificial · 2026-07-25

Anthropic's Opus 5 shows non-monotonic performance on coding tasks; the 'high' effort setting outperforms 'max' due to unnecessary refactors. The model also has a 6% higher hallucination rate than Opus 4.8, and safety classifiers may silently fall back to the older model.

0 favorites 0 likes
#model-performance

@elonmusk: Grok 4.5 is excellent for real-world work

X AI KOLs Timeline · 2026-07-24 Cached

Elon Musk endorses Grok 4.5 for real-world tasks, citing a Ramp test where Grok achieved the highest perfect-extraction rate on 150k business invoices.

0 favorites 0 likes
#model-performance

Opus 5 claims second place on simple bench

Reddit r/singularity · 2026-07-24

Opus 5 achieved second place on the Simple Bench benchmark, highlighting its competitive performance among AI models.

0 favorites 0 likes
#model-performance

Absurd claim: the distilled model outperforms the originals

Reddit r/LocalLLaMA · 2026-07-23

A claim suggests that a distilled AI model can outperform its larger original model, which is counterintuitive.

0 favorites 0 likes
#model-performance

Does Kimi K3 change the distillation debate?

Reddit r/singularity · 2026-07-19

Kimi K3's recent ranking challenges the notion that Chinese AI models rely heavily on distillation from US models, as it was released soon after Fable 5 and GPT-5.6. The article suggests Kimi K3 reflects substantial Chinese innovation.

0 favorites 0 likes
#model-performance

@LiorOnAI: The mood of 20+ million developers now depends on how well Anthropic and OpenAI’s models perform that day.

X AI KOLs Timeline · 2026-07-15

A tweet observing that the mood of over 20 million developers now depends on the daily performance of Anthropic and OpenAI's AI models.

0 favorites 0 likes
#model-performance

@elonmusk: Grok Build

X AI KOLs Timeline · 2026-07-10 Cached

Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.

0 favorites 0 likes
#model-performance

NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise?

Reddit r/LocalLLaMA · 2026-07-09

NVIDIA's Puzzle-75B-A9B model achieves 132 tokens per second using NVFP4 quantization on three RTX 3090 GPUs, raising discussion about the lack of competition in this model size category.

0 favorites 0 likes
#model-performance

@danshipper: This is wrong, it’s the same model But it does fall back to Opud 4.8 slightly more, so the benchmarks are measuring a m…

X AI KOLs Following · 2026-07-03 Cached

Dan Shipper argues that the Fable 5 model is not nerfed but falls back to Opus 4.8 more often, causing mixed benchmark results, contrary to claims of severe degradation.

0 favorites 0 likes
#model-performance

@jun_song: GPT-5.6 seems very disappointing. Nothing better than GLM-5.2

X AI KOLs Following · 2026-06-23 Cached

A user expresses disappointment with GPT-5.6, claiming it is not better than GLM-5.2.

0 favorites 0 likes
#model-performance

@haider1: GLM 5.2 feels like the opus 4.5 moment for open-weight models what genuinely impressed me was during long, multi-step a…

X AI KOLs Following · 2026-06-17 Cached

GLM 5.2 marks a significant milestone for open-weight models, demonstrating strong context retention across long multi-step tasks and more reliable tool calling.

0 favorites 0 likes
#model-performance

humanity's last exam current benchmarks thoughts?

Reddit r/singularity · 2026-06-15

Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.

0 favorites 0 likes
#model-performance

Opus 4.8 Thinking keeps deteroriating on Hard Prompts English in LMArena (again)

Reddit r/singularity · 2026-06-07

Opus 4.8 Thinking continues to deteriorate on the Hard Prompts English benchmark on LMArena, scoring 23 points lower than Opus 4.6 Thinking, which retains the top spot.

0 favorites 0 likes
#model-performance

Performance When Offloading Large Models to System RAM?

Reddit r/LocalLLaMA · 2026-05-24

Discusses performance trade-offs of offloading large AI model weights from GPU VRAM to system RAM, comparing different GPU configurations like RTX 5090 vs RTX6000 for models like DeepSeek V4 Pro.

0 favorites 0 likes
#model-performance

@swyx: very belated but in retrospect i think @sama's mythical "build a business that gets better when models get better" is b…

X AI KOLs Following · 2026-05-20 Cached

swyx reflects on Sam Altman's idea of building businesses that improve as AI models improve, linking it to the emerging concept of Agent Labs, and notes a clear correlation with revenue spikes in Q4 2025.

0 favorites 0 likes
#model-performance

Gemini 3.5 Flash Benchmarks

Reddit r/singularity · 2026-05-19

Benchmark results for the Gemini 3.5 Flash model are discussed, likely showcasing its performance across various AI tasks.

0 favorites 0 likes
#model-performance

Consider running a bigger quant if possible

Reddit r/LocalLLaMA · 2026-04-22

A user reports that switching from a highly-compressed IQ4_XS quant to the larger IQ4_NL_XL quant of Qwen 3.6 dramatically improves agentic-coding accuracy, despite lower tok/s, urging others to favor bigger quants when VRAM allows.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback