Tag
An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.
A user shares updated benchmark results for DeepSeek V4 Flash on SlopCodeBench using local quants (antirez imatrix quant) with the pi harness, showing improved performance over previous runs but still slower than the hosted API.
Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.
Maka is an open-source Agent Harness. Through mechanisms such as log-as-runtime, context pruning, and thinking feedback, it cuts the cost of the same DeepSeek task to 1/8 of OpenCode, while achieving a higher pass rate on Terminal-Bench at lower cost.
Grok Imagine Image 2.0 (Low) from xAI jumped to #2 in the Text-to-Image Arena, beating its own older quality model and showing significant improvement.
A local experiment comparing Qwen 35B-A3B MoE and Qwen 27B dense on coding-maintenance tasks, finding the MoE model ~3.9× faster with a smaller quality gap than expected.
Meta's Muse Spark 1.2 reaches #4 in the Text Arena, moving the cost-quality Pareto frontier upward with a ~91% price cut while sacrificing only 9 points relative to the top model.
Ahmad Osman shares performance numbers from running DeepSeek V4 Flash 0731 on an NVIDIA DGX Station.
New research from Tencent introduces Recursive Synthetic Terminal Tasks (RST), a method that progressively generates harder, verifiable training tasks for AI agents. Starting from 639 tasks, it produced 37,484 verified tasks across 15 rounds, and reinforcement learning with these tasks improved Qwen3.5-27B from 22.7% to 32.0% on Terminal-Bench Hard.
Introducing VGI-Bench, a multimodal benchmark that probes 12 distinct visual and audio-visual skills with 550 human-curated questions to expose failures in today's video benchmarks.
DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.
Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.
Jeff Geerling reviews the Dell XPS 13 with Intel Core 5 320, highlighting its impressive power efficiency and Linux compatibility, and expresses optimism about Intel's low-end processor direction.
A study comparing 10 coding agent harnesses on SWE-bench Pro shows harness choice dramatically affects pass@1 and cost, revealing vendor harnesses are overfitted to large models and do not transfer to smaller ones.
A user shares a head-to-head comparison of Seedance 2.5 versus Seedance 2.0 across four cinematic genres, noting significant differences in output quality.
LabyrinthBench is a local, judge-free LLM benchmark that deterministically measures context recall under interference for multi-step agentic tasks. Initial experiments show that wiping context and re-injecting gate answers helped 7 of 9 models but backfired on 2, highlighting the complexity of context management strategies.
ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.
This paper proposes a unified definition of uncertainty as pointwise posterior risk and introduces a theory-backed benchmark using semi-synthetic datasets to directly compute oracle epistemic and aleatoric uncertainty, enabling fine-grained evaluation beyond proxy tasks.
Presents FormBharo, a voice agent that uses LLMs with rule-based controls to fill structured forms over phone calls for low-literacy Hindi-speaking users in India, piloted with ARMMAN. The paper also introduces FormVoiceAgentBench, a benchmark of 3,760 multi-turn conversation tests, and shows that end-to-end evaluation is necessary since component-level performance does not predict full form completion.
EpiBench is a new closed-book, sequence-based benchmark for evaluating how well LLMs understand epitopes across five antibody-drug-discovery tasks, finding that current models capture partial signals but struggle with antibody-specific reasoning.