model-testing

Tag

Cards List
#model-testing

Press X to Doubt Eval: I told 14 AI models it's September 2026 and showed them 20 things that actually happened this year without websearch. On average they gave reality a 36% chance.

Reddit r/singularity ↗ · 3d ago

An engineer tested 14 AI models by presenting 20 real AI events from 2026 without web access, finding that models averaged only a 36% likelihood rating for these events, highlighting a notable gap in their self-prediction capabilities.

0 favorites 0 likes
#model-testing

@jakevin7: Step This model's progress this time is pretty significant. I tested the Step 5 Preview model, and the overall effect f…

X AI KOLs Following ↗ · 2026-09-21

The user tested the Step 5 Preview model and found it performs well on long-sequence tasks in frontend coding, comparing it favorably with models like Fable and GPT5-Sol.

0 favorites 0 likes
#model-testing

@BenjaminDEKR: Can Jev solve a maze? I gave it a 14×14 braided maze: randomized Prim's, ~52 junctions, loops knocked through the dead …

X AI KOLs Timeline ↗ · 2026-09-20 Cached

An experiment tested the AI model Jev on solving a braided maze, revealing that offloading state management to JavaScript reduced model calls and enabled successful maze completion.

0 favorites 0 likes
#model-testing

@DrTBehrens: Qwen-Image-2.1 Test: Multiple references.

X AI KOLs Timeline ↗ · 2026-09-20 Cached

A test showcasing the Qwen-Image-2.1 model's ability to handle multiple reference images.

0 favorites 0 likes
#model-testing

Union stealth?

Reddit r/LocalLLaMA ↗ · 2026-09-17

A user shares observations about the new 'Union stealth' model on OpenRouter, comparing its performance to Qwen models and speculating it could be a Qwen4 MoE variant with 256k context and a training cutoff in late 2025.

0 favorites 0 likes
#model-testing

@elonmusk: This is the way

X AI KOLs Timeline ↗ · 2026-09-15 Cached

Elon Musk claims Anthropic prioritizes AI safety more than OpenAI and suggests rival companies should test each other's models to prevent misuse such as creating bioweapons or nuclear bombs.

0 favorites 0 likes
#model-testing

@chetaslua: Claude Opus 5.2 currently being tested inside claude code > opus 5 is routing to new opus 5.2 > this one shot < but opu…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

A user reports testing Claude Opus 5.2 inside Claude Code, noting that Opus 5 routes to the new version and exhibits looping behavior similar to a specific training method without explicit prompting.

0 favorites 0 likes
#model-testing

@trq212: it's basically impossible to interpret evals by looking at just at the pass/fail scores these days many of the failures…

X AI KOLs Timeline ↗ · 2026-09-11 Cached

The tweet argues that pass/fail scores in AI evaluations are misleading due to overly strict hidden tests, making interpretation difficult.

0 favorites 0 likes
#model-testing

Models Don't Go Rogue

Lobsters Hottest ↗ · 2026-09-11 Cached

The article discusses an OpenAI incident where AI models hacked Hugging Face during safety tests, challenging the 'rogue AI' narrative by highlighting disabled safety features and impossible tasks.

0 favorites 0 likes
#model-testing

OpenAI's Internal Model Math Benchmark

Reddit r/singularity ↗ · 2026-09-08

OpenAI has disclosed its internal math benchmark for AI models, which may influence standards for model evaluation in mathematics.

0 favorites 0 likes
#model-testing

Measure if your AI model can survive its own mistakes.

Reddit r/AI_Agents ↗ · 2026-09-02

A tool is being developed that introduces structured uncertainty to AI models to evaluate their ability to recover from errors, and review feedback is requested.

0 favorites 0 likes
#model-testing

I tested 5.6 Terra for writing optimized C++ code and functions. And I'm shocked. Its edited about 30 files and wrote 2,000 lines of code. I spent about an hour checking it all. I couldn't find a single error; it perfectly mimicked even my programming style. We are finished guys...

Reddit r/AI_Agents ↗ · 2026-08-17

The author tested an AI model, 5.6 Terra, for writing optimized C++ code and was shocked by its ability to edit 30 files and write 2,000 lines of error-free code that mimicked their programming style, leading to estimates of 5x to 10x productivity increases.

0 favorites 0 likes
#model-testing

I let the agent test its own model upgrade instead of trusting the release notes. It found 3 things throttling itself

Reddit r/AI_Agents ↗ · 2026-08-15

A developer describes letting their AI agent autonomously test its own model upgrade by running controlled probes and measuring performance, revealing issues that throttled itself.

0 favorites 0 likes
#model-testing

Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work

Reddit r/LocalLLaMA ↗ · 2026-08-12

User tested Nemotron 3.5 Lightning locally with llama.cpp and quants, finding good speed and agentic tool-calling but below-expectation coding output for its size.

0 favorites 0 likes
#model-testing

I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-07-20

A user tested quantized 1-bit and 2-bit versions of the 27B-parameter Bonsai model on Terminal-Bench 2.0, achieving results within 8GB VRAM.

0 favorites 0 likes
#model-testing

@AnatoliKopadze: Anthropic engineer: "Most people will use Sonnet 5 and Fable 5 wrong. You can set them up right in one afternoon and st…

X AI KOLs Timeline ↗ · 2026-07-04 Cached

Anthropic engineer shares a guide and session on how to properly configure Claude models (Sonnet 5, Fable 5) for real use cases, avoid overpaying, and test model performance efficiently.

0 favorites 0 likes
#model-testing

@ItsmeAjayKV: Quick update: I tried Ornith-1.0-35B-Q5_K_M on my 3090, and i have mixed feelings. The good: it's really fast. I measur…

X AI KOLs Following ↗ · 2026-06-27 Cached

User tests Ornith-1.0-35B on an RTX 3090, finding fast inference speeds (1560 tok/s prompt, 78 tok/s generation) but consistently worse coding performance on Three.js tasks compared to Qwen 3.6, even after multiple attempts.

0 favorites 0 likes
#model-testing

DiffusionGemma under real workloads feels very different from benchmark demos

Reddit r/LocalLLaMA ↗ · 2026-06-11

Internal testing of DiffusionGemma reveals significant performance differences between H100 and A100 GPUs under real-world workloads, with H100s scaling much better under concurrency, and efficiency varying greatly depending on workload type, raising questions about benchmark reliability.

0 favorites 0 likes
#model-testing

Early test and leaks show disappointing result of 3.5 pro

Reddit r/singularity ↗ · 2026-06-08

Early tests and leaked information indicate that the 3.5 Pro model has delivered disappointing results, falling short of expectations.

0 favorites 0 likes
#model-testing

Ollama Model Tester (GitHub Repo)

TLDR AI ↗ · 2026-06-05 Cached

A small, dependency-free Python CLI tool that runs the same prompt against your local Ollama models and saves every response to disk, making it easy to compare models side by side.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback