ai-evaluation

Tag

Cards List
#ai-evaluation

@enginenerdx: same model, three medals: sonnet 5 scored 14/42 on the IMO in a web ui, 21 in claude code, 35 in a structured multi-age…

X AI KOLs Timeline · 2026-07-25 Cached

A comparison shows that LLM performance on the 2026 IMO varies dramatically depending on the evaluation harness, with structured multi-agent setups achieving far higher scores than simple web UI, indicating that current gains are absorbed at the frontier by better orchestration.

0 favorites 0 likes
#ai-evaluation

ARC AGI 3 could be gamed if Opus is a loop and not a pure model

Reddit r/singularity · 2026-07-24

Discusses a potential vulnerability in the ARC AGI 3 benchmark where the Opus model could be gamed if it functions as a loop rather than a pure model.

0 favorites 0 likes
#ai-evaluation

@no_stp_on_snek: Ran this on Laguna S 2.1 in Poolside's own agent (pool), pointed at a local instance on a DGX Spark, using the prompt l…

X AI KOLs Timeline · 2026-07-24 Cached

A comparison of two AI coding agents building a Mario game: Laguna S 2.1 in Poolside's agent took 62 minutes with self-correction and passed tests, while a previous Qwen model took hours and needed human help; highlights oracle discipline and native harness advantages.

0 favorites 0 likes
#ai-evaluation

tested whether AI models can recognize their own writing in a blind lineup. grok went 0 for 9. it wrote something, then a minute later insisted someone else wrote it

Reddit r/artificial · 2026-07-22 Cached

A self-awareness exam given to Claude, Gemini, and Grok found Claude and Gemini nearly aced it, while Grok scored 0 on self-recognition, failing to identify its own writing. The study measures six dimensions of functional self-knowledge.

0 favorites 0 likes
#ai-evaluation

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

arXiv cs.AI · 2026-07-22 Cached

Introduces SciHazard, a benchmark for measuring scientific safety risks in LLMs with a decomposed harm scoring framework, and evaluates 31 frontier models, finding deep research agents pose higher risks.

0 favorites 0 likes
#ai-evaluation

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

arXiv cs.LG · 2026-07-21 Cached

BACON proposes a four-stage pipeline that combines budgeted human labels with multiple AI judge outputs to produce calibrated item-level surrogate predictions, supporting both population-level estimation and individual-level scoring with improved accuracy and reduced bias.

0 favorites 0 likes
#ai-evaluation

@rohanpaul_ai: Self-improving AI is only as real as the signal it was tested on to see if it worked. Sorting 1,250 papers reveals a si…

X AI KOLs Timeline · 2026-07-20 Cached

Analysis of 1,250 papers on recursive self-improvement in AI reveals that the evaluator signal is the critical bottleneck. Models improve reliably only with strong, trustable signals like proof checkers, while weak signals cause loops to collapse or reinforce errors.

0 favorites 0 likes
#ai-evaluation

Agents Last Exam will be saturated by next February at the latest.

Reddit r/singularity · 2026-07-20

The article predicts that the AI benchmark 'Agents Last Exam' will reach saturation (models maxing out performance) by February next year.

0 favorites 0 likes
#ai-evaluation

Agent failures should become evals, not just traces

Reddit r/AI_Agents · 2026-07-20

Advocates for treating agent failures as evaluation benchmarks rather than just trace logs, emphasizing the need for systematic testing of AI agent behaviors.

0 favorites 0 likes
#ai-evaluation

Fable 5 and GPT-5.6 Lead the Singularity Gate. Benchmark for testing whether AI can predict paradigm-breaking discoveries after model cutoff

Reddit r/singularity · 2026-07-16

The Singularity Gate benchmark tests whether frontier AI models can predict paradigm-breaking scientific discoveries made after their training cutoff. Claude Fable 5 leads but has a low response rate due to refusals, while GPT-5.6 Sol shows strong performance without refusals at a lower price point.

0 favorites 0 likes
#ai-evaluation

@sayashk: I'm defending my PhD next Tuesday! The talk is titled "The Missing Science of AI Evaluation" and is based on my faculty…

X AI KOLs Following · 2026-07-15 Cached

Sayash Kapoor announces his upcoming PhD defense on 'The Missing Science of AI Evaluation,' discussing AI-based science, agent evaluation, and open model risks, livestreamed and open to all.

0 favorites 0 likes
#ai-evaluation

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

arXiv cs.AI · 2026-07-15 Cached

This paper introduces NameRank, a method to measure how well LLMs recognize individuals and artifacts from their parametric knowledge, finding that models recognize named projects and methods far better than credentials or contributors.

0 favorites 0 likes
#ai-evaluation

Good Benchmarks

arXiv cs.AI · 2026-07-15 Cached

This paper from July 2026 defines principles for designing effective benchmark tasks for AI agents, drawing on experience from the Terminal Bench project. Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons, with an emphasis on real-world relevance and outcome-based verification.

0 favorites 0 likes
#ai-evaluation

@HamelHusain: New Blog Post: Do Automated Evals Work? There has been a rise of tools that look through your traces with AI and identi…

X AI KOLs Timeline · 2026-07-14 Cached

A blog post from Parlance Labs tests automated AI evaluation tools (Braintrust Loop, Arize Alyx, LangSmith Engine) on real production data, finding they catch 87% of issues humans flag but miss domain-specific failures and add noise, recommending iterative human-in-the-loop use.

0 favorites 0 likes
#ai-evaluation

GPT-5.5 with tools now surpasses the 10-year-old level on the BabyVision benchmark

Reddit r/singularity · 2026-07-11

GPT-5.5, when equipped with tools, has surpassed the performance of a 10-year-old child on the BabyVision benchmark, marking a notable advancement in AI visual reasoning.

0 favorites 0 likes
#ai-evaluation

Psychological Competence as a Missing Dimension in AI Evaluation

arXiv cs.AI · 2026-07-10 Cached

This paper introduces psychological competence as a missing dimension in AI evaluation, proposing a conceptual framework to assess how AI systems support user cognition, emotional interpretation, and decision-making in human-facing roles.

0 favorites 0 likes
#ai-evaluation

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

arXiv cs.AI · 2026-07-10 Cached

This paper audits whether self-consistency and cross-model agreement are reliable indicators of correctness in LLMs, finding that agreement is a weak, regime-dependent proxy and that frontier models exhibit overconfidence.

0 favorites 0 likes
#ai-evaluation

Measuring Intelligence Beyond Human Scale

arXiv cs.AI · 2026-07-09 Cached

This paper proposes a new paradigm for measuring intelligence beyond human capability using adversarial psychometric rating systems, where models generate challenges to separate other systems, enabling evaluation that scales with AI capabilities.

0 favorites 0 likes
#ai-evaluation

We made Grok 4.5, GPT-5.5, and Claude build the same apps

Reddit r/singularity · 2026-07-09 Cached

This article benchmarks Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 by having each model build three interactive apps (3D Rubik's Cube, particle gravity sandbox, Breakout game) from a single prompt, comparing their one-shot coding capabilities.

0 favorites 0 likes
#ai-evaluation

@OpenAI: As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchmarks help the field u…

X AI KOLs · 2026-07-08 Cached

OpenAI emphasizes the need for more rigorous and trustworthy evaluations for coding AI models to better measure real progress.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback