model-benchmark

Tag

Cards List
#model-benchmark

Are AI labs pelicanmaxxing?

Simon Willison's Blog · 2026-07-22 Cached

Dylan Castillo conducted a rigorous investigation to determine if AI labs have been secretly training models to draw pelicans riding bicycles. Testing multiple models with various animal-vehicle combinations, he found no evidence of 'pelicanmaxxing'.

0 favorites 0 likes
#model-benchmark

Good example of the gap between Fable and Sol

Reddit r/singularity · 2026-07-15 Cached

The article compares the performance of OpenAI GPT-5.6 Soul and Anthropic Claude Fable 5 in physical 3D printed part replication and autonomous magazine production. Soul slightly outperforms in speed and design precision, but both require significant human intervention in complex real-world tasks, exposing the limitations of current AI in real-world manufacturing tasks.

0 favorites 0 likes
#model-benchmark

Qwen 3.6 27B - VLLM Performance Benchmark Results (BF16, FP8, NVFP4)

Reddit r/LocalLLaMA · 2026-07-05

A detailed benchmark of Qwen 3.6 27B using VLLM across BF16, FP8, and NVFP4 quantizations, showing NVFP4 fastest for token generation but FP8 best for prompt processing, with practical advice on choosing the right quantization for coding tasks.

0 favorites 0 likes
#model-benchmark

@MMMusol: Gemini 3.1 Pro, GPT 5.5, Deepseek V4, and the latest Claude Fable 5 performed the same test, as shown in the video. Compare for yourself~ The prompt is as follows: Create an HTML file to render a high-speed, aggressive fighter jet at full afterburner...

X AI KOLs Timeline · 2026-06-10 Cached

Multiple AI models (Gemini 3.1 Pro, GPT 5.5, Deepseek V4, Claude Fable 5) were asked to generate the same fighter jet HTML animation. The video shows a comparison of each model's output.

0 favorites 0 likes
#model-benchmark

Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, largely because it is one of the most cautious models.

Reddit r/singularity · 2026-05-21

Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, measuring how often models change judgment to side with the user. The benchmark reveals that some models are sycophantic while others are decisive or cautious.

0 favorites 0 likes
#model-benchmark

Gave GPT-4o and Claude the exact same double pendulum prompt. They picked opposite angle conventions within seconds.

Reddit r/ArtificialInteligence · 2026-05-16

An experiment feeding GPT-4o, Claude 3.5 Sonnet, and other models the same double pendulum prompt reveals they pick opposite angle conventions, causing immediate visible mismatch in a shared renderer. The convention split, non-random across model families, suggests a bias in training data distribution for classical mechanics problems.

0 favorites 0 likes
← Back to home

Submit Feedback