CogScale: Scalable Benchmark for Sequence Processing
Summary
CogScale is a benchmark of 14 scalable synthetic tasks designed to isolate and evaluate cognitive and memory abilities in sequence processing models. It provides a lightweight framework for rapid architectural validation and includes evaluations of seven architectures under strict parameter budgets.
Similar Articles
CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
CogGym is a scalable framework for comparing human and AI cognition using cognitive experiments, revealing that larger language models better mimic human reasoning but still lag behind formal benchmarks.
Scaling Test-Time Compute for Agentic Coding
A test-time scaling framework for agentic coding that compresses rollout trajectories into structured summaries and uses recursive voting/PDR to boost Claude-4.5-Opus to 77.6% on SWE-Bench Verified.
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
CogArena introduces a procedurally generated 13-paradigm benchmark to evaluate whether LLMs exhibit separable cognitive abilities or a single general competence, finding only weak support for stable five-dimensional profiles across 55 models.
CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility
CoMedBench is a reproducible benchmark evaluating synthetic medical data generators across 37 dataset-task pairs, showing that synthetic training data preserves most downstream signal on tabular tasks but temporal ICU tasks remain generator-sensitive.
@swyx: Finally! the first eval ship from cog!!!!!!!!!! To contextualize: @METR_Evals cap out at ~16 hours. Cog has private ent…
Cognition released the first evaluation suite for Devin, offering up to 100-hour enterprise evals with a financial guarantee. The dataset includes real-world Java/TypeScript/Python/C# tasks from 126 enterprise users, aiming to measure engineering productivity more accurately than existing benchmarks.