BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

arXiv cs.AI Papers

Summary

BrainBench is a new unified benchmark for evaluating large language models on comprehensive, instruction-conditioned EEG understanding, covering 17 datasets, 172 tasks, and over 4K real-data instances. The paper evaluates 13 LLMs across two execution paradigms, showing that EEG competence varies by model and operationalization.

arXiv:2608.04156v1 Announce Type: new Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:41 AM

# BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
Source: [https://arxiv.org/html/2608.04156](https://arxiv.org/html/2608.04156)
Yangxuan Zhou1,2, Sha Zhao1,2, Yuning Chen1,2, Chen Wu1,2, Jiquan Wang1,2, Shijian Li1,Gang Pan1,2,3 1State Key Laboratory of Brain\-machine Intelligence, Zhejiang University 2College of Computer Science and Technology, Zhejiang University 3MOE Frontier Science Center for Brain Science and Brain\-machine Integration, Zhejiang University \{zyangxuan, szhao, yuningchen, chen\_wu\_, wangjiquan\}@zju\.edu\.cn; \{shijianli, gpan\}@zju\.edu\.cn;

###### Abstract

Electroencephalography \(EEG\) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural\-language instructions, signal processing, quantitative evidence, and scientific interpretation\. We term this capability*comprehensive EEG understanding*\. Existing evaluations, however, primarily target isolated decoding tasks or system\-specific demonstrations, leaving the competence of large language models \(LLMs\) insufficiently quantified\. We introduce BrainBench, a unified benchmark for comprehensive, instruction\-conditioned EEG understanding\. It comprises four subsets—Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration—covering 17 datasets, 172 tasks, and over 4K real\-data instances\. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts\. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation\. We evaluate 13 representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent\. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization\. BrainBench provides a reproducible testbed for advancing LLM\-based EEG understanding\. The code and benchmark will be released soon, with evaluation results continuously updated\.

## 1Introduction

Electroencephalography \(EEG\) provides a non\-invasive window into brain activity with millisecond\-scale temporal resolution\. Its accessibility and sensitivity to rapid neural dynamics have established it as a fundamental tool in neuroscience, sleep research, neurological assessment, and brain–computer interfaces\[[2](https://arxiv.org/html/2608.04156#bib.bib4),[46](https://arxiv.org/html/2608.04156#bib.bib1),[19](https://arxiv.org/html/2608.04156#bib.bib2),[12](https://arxiv.org/html/2608.04156#bib.bib3)\]\. Yet, computational EEG research has predominantly treated signal analysis as a decoding problem, in which a model maps a recording to a predefined output, such as a sleep stage, cognitive state, or clinical label\[[53](https://arxiv.org/html/2608.04156#bib.bib5),[67](https://arxiv.org/html/2608.04156#bib.bib6),[69](https://arxiv.org/html/2608.04156#bib.bib7),[13](https://arxiv.org/html/2608.04156#bib.bib8)\]\. While this paradigm has driven substantial progress, it captures only a narrow component of real\-world EEG analysis\. In practice, EEG analysis requires a system to understand the analytical objective, reason coherently over real recordings, and arrive at a scientifically grounded conclusion\[[25](https://arxiv.org/html/2608.04156#bib.bib10),[23](https://arxiv.org/html/2608.04156#bib.bib11),[40](https://arxiv.org/html/2608.04156#bib.bib9)\]\. We refer to this broader capability ascomprehensive EEG understanding\. The central question therefore shifts from whether a model can predict a predefined target to whether it can carry an EEG analysis coherently from instruction to conclusion\.

Large language models \(LLMs\) and agents offer a natural foundation for comprehensive EEG understanding\[[7](https://arxiv.org/html/2608.04156#bib.bib12),[63](https://arxiv.org/html/2608.04156#bib.bib13)\]\. By combining language understanding, reasoning, and code generation, they can translate natural\-language instructions into executable analyses, potentially transforming EEG analysis from a collection of specialized pipelines into an interactive, general\-purpose process\. Recent LLM\-powered systems in brain science and neurotechnology have demonstrated their potential to support increasingly complex scientific workflows\[[65](https://arxiv.org/html/2608.04156#bib.bib14),[1](https://arxiv.org/html/2608.04156#bib.bib16),[29](https://arxiv.org/html/2608.04156#bib.bib17),[11](https://arxiv.org/html/2608.04156#bib.bib18),[62](https://arxiv.org/html/2608.04156#bib.bib39),[52](https://arxiv.org/html/2608.04156#bib.bib20),[68](https://arxiv.org/html/2608.04156#bib.bib15),[31](https://arxiv.org/html/2608.04156#bib.bib19),[41](https://arxiv.org/html/2608.04156#bib.bib43),[16](https://arxiv.org/html/2608.04156#bib.bib41)\]\. For example, CELM and CerebraGloss explore language\-based clinical EEG interpretation and report generation\[[41](https://arxiv.org/html/2608.04156#bib.bib43),[16](https://arxiv.org/html/2608.04156#bib.bib41)\], while BrainAgent investigates multi\-agent orchestration for automating end\-to\-end EEG analysis\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\]\. However, these systems are still evaluated primarily within their respective task\- or system\-specific settings\. As a result, a comprehensive and systematic assessment of the correctness and scientific validity of LLM\-generated EEG analyses across heterogeneous tasks and execution paradigms remains lacking\.

*Can LLMs move beyond isolated decoding tasks to achievecomprehensive EEG understandingacross real\-world analysis workflows?*

Despite this promise, existing evaluation protocols are not designed to quantify comprehensive EEG understanding\. Existing EEG benchmarks remain largely centered on fixed decoding objectives\[[58](https://arxiv.org/html/2608.04156#bib.bib21),[60](https://arxiv.org/html/2608.04156#bib.bib22),[48](https://arxiv.org/html/2608.04156#bib.bib23),[35](https://arxiv.org/html/2608.04156#bib.bib24),[28](https://arxiv.org/html/2608.04156#bib.bib25)\], rather than assessing whether a system can interpret diverse analytical instructions and carry them through to scientifically grounded conclusions\. Moving beyond this decoding\-centric paradigm, HeaRTS\[[33](https://arxiv.org/html/2608.04156#bib.bib26)\]takes an important step by requiring LLMs to analyze real\-world health time\-series data through executable code\. Its primary emphasis, however, is breadth across physiological modalities rather than depth within EEG analysis\. The distinctive signal characteristics, analytical conventions, and domain\-specific interpretations of EEG give rise to heterogeneous workflows that demand fine\-grained evaluation of both analytical procedures and resulting conclusions\. Consequently, a unified benchmark for comprehensive EEG understanding remains absent\.

Operationalizing comprehensive EEG understanding as a rigorous benchmark is challenging because the capability is inherently workflow\-oriented and action\-grounded\. It requires a system to carry out a coherent sequence of file inspection, signal processing, quantitative analysis, comparison, and scientific interpretation, while grounding its reasoning in executable analyses of real recordings\. Capturing this capability therefore requires diverse datasets, analytical tasks, instruction formats, and evaluation dimensions under unified and comparable protocols for execution and assessment\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x2.png)Figure 1:Overview of BrainBench\. We present a comprehensive benchmark for instruction\-conditioned EEG understanding, spanning 17 datasets, 173 tasks, and over 4K real\-data instances across four complementary subsets: Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration\. Across over 100K executions, we further compare two execution paradigms: autonomous code execution with CodeAct and structured, reproducible workflows with BrainAgent\.To address these challenges, we introduceBrainBench, a benchmark for comprehensive, instruction\-conditioned EEG understanding\. It comprises four complementary subsets—Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Signal Integration—covering172 tasksand4K instancesdrawn from17 datasetsand spanning diverse analytical objectives and levels of complexity\. Together, these subsets cover foundational signal analysis, domain\-specific reasoning, and multimodal interpretation, enabling EEG understanding to be evaluated in both breadth and depth\. Rather than assessing performance on a fixed predictive target, BrainBench evaluates whether a model can transform a natural\-language instruction and real EEG recordings into a scientifically grounded conclusion\. We further evaluate the same LLMs under two complementary execution paradigms\. CodeAct\[[57](https://arxiv.org/html/2608.04156#bib.bib38)\]allows an LLM to autonomously plan and execute analyses through self\-generated code, providing a relatively open setting for eliciting its analytical capability\. In contrast, BrainAgent\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\]is an EEG\-oriented multi\-agent framework that executes analyses through structured workflows, controlled tools, and traceable intermediate actions\. By holding the instructions, data, and evaluation criteria constant, BrainBench isolates the effect of the execution paradigm and examines the trade\-off between the flexibility of autonomous code generation and the reliability, auditability, and operational safety of structured agentic analysis\.

We evaluate13representative LLMs across all four subsets and under both execution paradigms, revealing key gaps in current systems and important directions for advancing comprehensive EEG understanding in the future\. Beyond the current release, BrainBench is designed for extensibility, allowing new EEG domains, datasets, task formulations, evaluation dimensions, and models to be incorporated as the field evolves\. Our contributions are summarized as follows:

- •We formulate comprehensive, instruction\-conditioned EEG understanding as a unified evaluation target and introduce BrainBench, comprising four complementary subsets, 172 tasks, and 4K instances across 17 diverse datasets, analytical objectives, and workflows\.
- •We establish a dual\-paradigm protocol that compares autonomous code execution with a structured EEG agent under shared instructions, data, and ground truth, enabling controlled analysis of how the execution paradigm shapes measured EEG competence\.
- •We systematically evaluate 13 representative LLMs across tasks and execution paradigms, providing a multidimensional characterization of their capabilities and revealing key limitations and directions for advancing comprehensive EEG understanding\.

## 2Related Work

### 2\.1EEG Models and Decoding Benchmarks

EEG modeling has evolved from handcrafted feature pipelines and task\-specific neural networks to foundation models trained on large\-scale recordings to learn transferable representations\[[30](https://arxiv.org/html/2608.04156#bib.bib29),[22](https://arxiv.org/html/2608.04156#bib.bib27),[53](https://arxiv.org/html/2608.04156#bib.bib5),[69](https://arxiv.org/html/2608.04156#bib.bib7),[13](https://arxiv.org/html/2608.04156#bib.bib8),[59](https://arxiv.org/html/2608.04156#bib.bib28),[54](https://arxiv.org/html/2608.04156#bib.bib30),[55](https://arxiv.org/html/2608.04156#bib.bib31)\]\. In parallel, existing benchmarks have standardized evaluation across downstream decoding tasks, typically using linear probing or fine\-tuning to assess representation quality and cross\-subject or cross\-dataset generalization\[[58](https://arxiv.org/html/2608.04156#bib.bib21),[60](https://arxiv.org/html/2608.04156#bib.bib22),[48](https://arxiv.org/html/2608.04156#bib.bib23),[35](https://arxiv.org/html/2608.04156#bib.bib24),[28](https://arxiv.org/html/2608.04156#bib.bib25),[56](https://arxiv.org/html/2608.04156#bib.bib32),[24](https://arxiv.org/html/2608.04156#bib.bib33),[5](https://arxiv.org/html/2608.04156#bib.bib34)\]\. Clinically grounded models have further extended this paradigm to full\-session recordings and multimodal clinical context; for example, CLEF\[[8](https://arxiv.org/html/2608.04156#bib.bib35)\]aligns session\-level EEG representations with neurologist reports and electronic health records, yet is still evaluated mainly through patient\-level classification tasks\. Thus, despite increasing scale and clinical relevance, existing evaluations continue to define EEG competence primarily through performance on predefined predictive targets\. They do not assess whether a general\-purpose model can interpret an open\-ended analytical instruction, execute the required analysis on real recordings, and derive a scientifically grounded conclusion\. This leaves open the need for evaluating comprehensive, instruction\-conditioned EEG understanding beyond fixed\-label decoding\.

### 2\.2Large Language Models for EEG Analysis

The reasoning\[[63](https://arxiv.org/html/2608.04156#bib.bib13)\], planning\[[50](https://arxiv.org/html/2608.04156#bib.bib36)\], and tool\-use\[[47](https://arxiv.org/html/2608.04156#bib.bib37)\]capabilities of large language models have enabled a new class of systems that interact with EEG data through natural\-language instructions\. These systems combine language understanding with code generation, domain knowledge, and specialized analytical tools to support tasks ranging from data inspection and preprocessing to event detection and report generation\[[65](https://arxiv.org/html/2608.04156#bib.bib14),[29](https://arxiv.org/html/2608.04156#bib.bib17),[1](https://arxiv.org/html/2608.04156#bib.bib16),[11](https://arxiv.org/html/2608.04156#bib.bib18),[62](https://arxiv.org/html/2608.04156#bib.bib39),[52](https://arxiv.org/html/2608.04156#bib.bib20),[6](https://arxiv.org/html/2608.04156#bib.bib40),[68](https://arxiv.org/html/2608.04156#bib.bib15),[31](https://arxiv.org/html/2608.04156#bib.bib19),[18](https://arxiv.org/html/2608.04156#bib.bib42)\]\. For example, EEGAgent\[[65](https://arxiv.org/html/2608.04156#bib.bib14)\]coordinates specialized tools for automated EEG analysis and reporting, while SleepLM\[[62](https://arxiv.org/html/2608.04156#bib.bib39)\]connect physiological recordings with natural\-language sleep assessment\. More recent agentic frameworks extend this paradigm toward longer and more heterogeneous workflows: NeuroWeaver\[[52](https://arxiv.org/html/2608.04156#bib.bib20)\]autonomously explores EEG analysis pipelines\. BrainAgent\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\]decomposes user intent among specialized agents for structured brain\-signal analysis, and BrainPilot\[[31](https://arxiv.org/html/2608.04156#bib.bib19)\]explores autonomous brain\-science discovery through a multi\-agent framework\. Collectively, these studies demonstrate the feasibility of transforming high\-level user requests into executable EEG analyses\. Their rapid development, however, raises a distinct evaluation question:how reliably do the resulting analyses reflect both the supplied recordings and the scientific intent expressed in the instruction?

### 2\.3Evaluating Large Language Models for EEG Understanding

Existing evaluations differ substantially in scope and target\. EEGAgent\[[65](https://arxiv.org/html/2608.04156#bib.bib14)\]combines task\-level metrics with qualitative demonstrations of EEG interpretation and reporting\. BrainAgent\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\]evaluates 60 tasks across three difficulty levels, focusing on task completion, routing reliability, and tool\-use efficiency\. BrainPilot\[[31](https://arxiv.org/html/2608.04156#bib.bib19)\]broadens the scope to brain\-science research, although the initial BrainPilotBench\-v0 contains four tasks covering calcium imaging, fMRI, motor\-imagery EEG, and sleep EEG\. HeaRTS\[[33](https://arxiv.org/html/2608.04156#bib.bib26)\]provides the closest general benchmark for executable reasoning over physiological time series, covering diverse health domains and signal modalities through autonomous code execution, but its primary objective is breadth across health time series rather than analytical depth within EEG\. Taken together, existing evaluations remain fragmented across system\-specific tasks, metrics, and protocols\. A comprehensive and systematic assessment of the correctness and scientific validity of LLM\-generated EEG analyses across heterogeneous tasks, datasets, and execution paradigms therefore remains lacking\.

## 3BrainBench

To address this gap, we introduce BrainBench, a unified benchmark for comprehensive, instruction\-conditioned EEG understanding\. Given a natural\-language instruction and real EEG recordings, BrainBench evaluates whether an LLM can produce scientifically grounded analysis outputs beyond predefined predictive tasks\. It comprises four complementary subsets spanning 172 tasks, 4K instances, and 17 datasets as shown in Figure[1](https://arxiv.org/html/2608.04156#S1.F1)\. Under a unified assessment protocol, we evaluate13representative LLMs with CodeAct\[[57](https://arxiv.org/html/2608.04156#bib.bib38)\]and BrainAgent\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\], enabling controlled comparison between autonomous code execution and structured EEG\-agent workflows\.

### 3\.1BrainBenchOrganization and Formulation

BrainBench formulates comprehensive EEG understanding as instruction\-conditioned analytical execution\. The benchmark follows a three\-level hierarchy ofsubsets,tasks, andinstances\. Four complementarysubsetsdefine its evaluation scope: Foundational Analysis targets general EEG operations; Sleep Assessment and Neurocognitive Assessment cover domain\-specific analytical workflows; and Physiological Integration evaluates reasoning across EEG and complementary physiological signals\. Within each subset, ataskdefines a reusable analytical objective, the expected outputs, and the corresponding assessment requirements, and is instantiated by a collection of evaluationinstances\. Theseinstancespreserve the task\-level requirements while varying the underlying recordings, subjects, analysis parameters, or task\-specific instructions\. This hierarchy enables BrainBench to assess whether analytical capabilities generalize across heterogeneous inputs rather than merely succeed on individual recordings or predefined queries\. For an instanceii, letℐi\\mathcal\{I\}\_\{i\}denote its natural\-language instruction and𝒟i\\mathcal\{D\}\_\{i\}the supplied input files\. Given\(ℐi,𝒟i\)\(\\mathcal\{I\}\_\{i\},\\mathcal\{D\}\_\{i\}\), a target systemℳ\\mathcal\{M\}produces a free\-form analysis reportℛi\\mathcal\{R\}\_\{i\}and, when required, a set of artifacts𝒜i\\mathcal\{A\}\_\{i\}\. The unified evaluation functionℰ\\mathcal\{E\}then assigns the instance scoresis\_\{i\}:

\(ℛi,𝒜i\)=ℳ​\(ℐi,𝒟​i\),si=ℰ​\(ℛi,𝒜​i\)\.\(\\mathcal\{R\}\_\{i\},\\mathcal\{A\}\_\{i\}\)=\\mathcal\{M\}\(\\mathcal\{I\}\_\{i\},\\mathcal\{D\}i\),\\qquad s\_\{i\}=\\mathcal\{E\}\(\\mathcal\{R\}\_\{i\},\\mathcal\{A\}i\)\.\(1\)The evaluator applies the assessment criteria defined by the corresponding task while retaining a common interface across the benchmark\. This formulation accommodates heterogeneous EEG workflows and output formats, while ensuring that different execution paradigms are evaluated on the same tasks and instances under consistent requirements\.

### 3\.2BrainBench Construction and Evaluation

#### 3\.2\.1Task Definition

![Refer to caption](https://arxiv.org/html/2608.04156v1/x3.png)Figure 2:Construction of a reusabletaskand its data\-boundinstancesacross heterogeneous recordings\.In BrainBench, eachtaskdefines a recording\-independent analytical objective and the EEG understanding capability it is intended to assess\. It also specifies the required inputs, expected outputs, and applicable validation criteria\. These requirements remain fixed across all correspondinginstances, allowing the same capability to be evaluated across different subjects, datasets, and recordings\. Tasks are stratified into three difficulty levels according to their intrinsic analytical demands\.Easytasks involve short, explicitly specified analyses with conclusions derived directly from limited evidence\.Mediumtasks require multiple dependent steps and integration of intermediate results\.Hardtasks involve extended workflows that synthesize evidence across channels, time scales or physiological modalities\. Difficulty is assigned at the task level and inherited by all associated instances\.

#### 3\.2\.2Instance Construction

Ataskis converted into executable evaluation examples by binding its specification to a concrete data context\. Each resultinginstanceidentifies the dataset and recording to be analyzed, together with the applicable time window, signal selection, analysis parameters, and other task\-specific conditions\. For each such binding, BrainBench generates a complete natural\-language instruction that specifies the available input files, analysis scope, required outputs, and reporting constraints\. This process preserves a consistent analytical objective and evaluation target across instances while adapting the instruction to the data available in each evaluation context\. Shown in Figure[2](https://arxiv.org/html/2608.04156#S3.F2), each instance is further paired with an evaluator\-side reference package\. Given the bound input files and analysis parameters, a deterministic analysis script executes the prescribed reference workflow and computes the expected values\. The resulting outputs are combined with the validation configuration, including the parser prompt and metric definitions, to form a complete instance\-level specification for evaluation\. Reusing a fixed script and parameterization across instances ensures that the ground truth remains reproducible only with the bound data\.

#### 3\.2\.3Multi\-Unit Validation

![Refer to caption](https://arxiv.org/html/2608.04156v1/x4.png)Figure 3:Distributions of task difficulty and validation units across subsets\.BrainBench supports free\-form analytical reports to accommodate the diverse wording and presentation required by heterogeneous EEG tasks\. To enable standardized evaluation, a Parser Agent extracts the required content from each report into a structured representation according to an instance\-specific*Validation Configuration*, which defines the target fields, extraction prompt, validation units, reference targets, and metric weights\. It serves solely as an extractor: it returns null for missing information and neither evaluates scientific correctness nor infers unreported results\. This design preserves flexible natural\-language reporting without conflating EEG understanding with rigid format compliance\.

The extracted fields, original report, and generated artifacts are evaluated through six complementary validation units\.Numerical validationcompares reported scalar values with the computed ground truth under task\-specific tolerances\.Categorical validationassesses discrete labels or choices after canonicalization\.Set validationevaluates unordered collections using exact matching or element\-level partial credit\.Sequence validationextends this assessment to ordered outputs by considering both element coverage and positional consistency\.Semantic validationuses a task\-specific Semantic Judge to assess non\-scalar conclusions for correctness, relevance, instruction consistency, and unsupported or hallucinatory claims\.Artifact validationverifies whether requested files are successfully generated and valid; structured signal files are checked programmatically, while visual outputs can additionally be assessed by a VLM Judge\. A single instance may combine multiple validation units to cover complementary aspects of the requested output\.

For instanceii, each of its validation metrics produces a normalized scorevi​m∈\[0,1\]v\_\{im\}\\in\[0,1\]with a non\-negative weightwi​mw\_\{im\}\. For subset aggregation,di∈\{1\.0,1\.5,2\.0\}d\_\{i\}\\in\\\{1\.0,1\.5,2\.0\\\}denotes the task\-level difficulty coefficient for Easy, Medium, and Hard instances, respectively\. The instance and subset scores are computed as

si=100×∑mwi​m​vi​m∑mwi​m,S𝒮=∑i∈𝒮di​si∑i∈𝒮di\.s\_\{i\}=100\\times\\frac\{\\sum\_\{m\}w\_\{im\}v\_\{im\}\}\{\\sum\_\{m\}w\_\{im\}\},\\qquad S\_\{\\mathcal\{S\}\}=\\frac\{\\sum\_\{i\\in\\mathcal\{S\}\}d\_\{i\}s\_\{i\}\}\{\\sum\_\{i\\in\\mathcal\{S\}\}d\_\{i\}\}\.\(2\)Both scores lie on a0–100100scale, with the subset score assigning greater weight to more difficult tasks\. Detailed matching rules are provided in Appendix[C](https://arxiv.org/html/2608.04156#A3)\.

#### 3\.2\.4Black\-Box Evaluation

BrainBench adopts a black\-box evaluation protocol in which each target system receives only the natural\-language instruction and associated input files\. The evaluator\-side reference package and*Validation Configuration*remain inaccessible during execution\. Each instance is processed in an isolated container, and only the final report and requested artifacts contribute to the benchmark score\. Execution actions, tool calls, execution traces, runtime errors, token usage, and latency are recorded separately for reproducibility and failure analysis\. Each LLM is evaluated under two complementary execution paradigms through the same input–output interface\. InCodeAct\[[57](https://arxiv.org/html/2608.04156#bib.bib38)\], the model autonomously plans the analysis and generates and executes Python code in an interactive environment\. InBrainAgent\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\], the model performs the analysis through structured agent coordination and controlled tool invocation\. To cover the evaluation capability boundary defined by BrainBench, we equip BrainAgent with a capability\-oriented toolset comprising reusable EEG analysis operations rather than task\- or instance\-specific solutions\. Holding the instructions, input data, and validation protocol fixed enables a controlled comparison between autonomous code execution and structured agentic analysis\. More details are provided in the Appendix[D](https://arxiv.org/html/2608.04156#A4)\.

### 3\.3BrainBench Composition

BrainBench comprises four complementary subsets that cover distinct dimensions of comprehensive EEG understanding \(complete construction details for all subsets are provided in the Appendix[B](https://arxiv.org/html/2608.04156#A2)\):

- •Foundational Analysis\.This subset comprises 40tasksand 950instancesderived from ISRUC\[[26](https://arxiv.org/html/2608.04156#bib.bib44)\], BCIC2020\-3\[[20](https://arxiv.org/html/2608.04156#bib.bib45)\], SEED\-V\[[34](https://arxiv.org/html/2608.04156#bib.bib46)\], Mumtaz2016\[[36](https://arxiv.org/html/2608.04156#bib.bib47)\], and MentalArithmetic\[[70](https://arxiv.org/html/2608.04156#bib.bib48)\], together with a curated EEG and BCI knowledge collection\. It evaluates general\-purpose EEG capabilities spanning recording inspection, preprocessing, feature extraction, spatiotemporal comparison, connectivity analysis, artifact generation, and evidence\-grounded interpretation\.
- •Sleep Assessment\.This subset comprises 43tasksand 1,025instancesderived from HMC\[[3](https://arxiv.org/html/2608.04156#bib.bib58)\], ISRUC\[[26](https://arxiv.org/html/2608.04156#bib.bib44)\], MASS\-SS3\[[38](https://arxiv.org/html/2608.04156#bib.bib59)\], PhysioNet 2018\[[14](https://arxiv.org/html/2608.04156#bib.bib60)\], and SHHS\-1\[[42](https://arxiv.org/html/2608.04156#bib.bib57)\], together with a curated sleep\-medicine knowledge collection\. It evaluates sleep architecture quantification, staging and spectral interpretation, sleep\-event detection, multimodal physiological analysis, and the calculation of clinically relevant whole\-night indices\.
- •Neurocognitive Assessment\.This subset comprises 50tasksand 1,030instancesderived from FACED\[[10](https://arxiv.org/html/2608.04156#bib.bib49)\], REFED\[[37](https://arxiv.org/html/2608.04156#bib.bib50)\], COG\-BCI\[[17](https://arxiv.org/html/2608.04156#bib.bib51)\], and MPD\-DF\[[32](https://arxiv.org/html/2608.04156#bib.bib52)\], together with curated domain\-knowledge questions\. It evaluates the understanding of affective states, cognitive workload, and fatigue through state recognition, feature\-based comparison, temporal and within\-subject analysis, multimodal physiological evidence integration, and scientifically grounded interpretation\.
- •Physiological Integration\.This subset comprises 39tasksand 1,120instancesderived from SEED\-VII\[[21](https://arxiv.org/html/2608.04156#bib.bib53)\], DEAP\[[27](https://arxiv.org/html/2608.04156#bib.bib54)\], Simultaneous Dataset B\[[49](https://arxiv.org/html/2608.04156#bib.bib55)\], SEED\-VIG\[[66](https://arxiv.org/html/2608.04156#bib.bib56)\], and SHHS\-1\[[42](https://arxiv.org/html/2608.04156#bib.bib57)\], together with a curated multimodal neurophysiology knowledge collection\. It evaluates the integration of EEG with complementary physiological signals through temporal alignment, modality\-specific feature extraction, cross\-modal coupling and evidence fusion, signal\-quality assessment, data repair, missing\-modality reconstruction, and multimodal result generation\.

## 4Experiments and Analysis

We evaluate 13 API\-accessible LLMs spanning eight model families and multiple capability tiers\. The evaluated models include the Qwen3\.5 and Qwen3\.7 series\[[43](https://arxiv.org/html/2608.04156#bib.bib61),[45](https://arxiv.org/html/2608.04156#bib.bib62),[44](https://arxiv.org/html/2608.04156#bib.bib63)\], GLM\-5\[[64](https://arxiv.org/html/2608.04156#bib.bib64)\], DeepSeek\-V4\-Flash\[[61](https://arxiv.org/html/2608.04156#bib.bib65)\], Kimi K2\.5\[[51](https://arxiv.org/html/2608.04156#bib.bib66)\], MiniMax\-M2\.5\[[9](https://arxiv.org/html/2608.04156#bib.bib67)\], GPT\-5\.6 Luna, GPT\-5\.6 Terra and GPT\-5\.6 Sol\[[39](https://arxiv.org/html/2608.04156#bib.bib68)\], Gemini 3\.6 Flash\[[15](https://arxiv.org/html/2608.04156#bib.bib69)\], and Claude Opus 5\[[4](https://arxiv.org/html/2608.04156#bib.bib70)\]\. To standardize test\-time computation, all optional reasoning or thinking modes exposed by the corresponding APIs are disabled\. Each model is evaluated under both BrainAgent and CodeAct using the same model endpoint and decoding configuration\. All executions are conducted in isolated containers with fixed instance inputs and evaluation settings\. Runs affected by verified provider\-side outages, network interruptions, or API transport errors are retried according to a predefined policy, detailed in Appendix[D](https://arxiv.org/html/2608.04156#A4)\.

Table 1:BrainBench leaderboard\.Performance of evaluated models under BrainAgent and CodeAct\. Overall scores are the arithmetic mean of the two subset scores\. The best result in each column is shown inbold, and the second\-best result isunderlined\.### 4\.1Can LLMs Achieve Comprehensive EEG Understanding?

Table[1](https://arxiv.org/html/2608.04156#S4.T1)reports the currently completed results on Foundational Analysis \(FA\) and Sleep Assessment \(SA\)\. The evaluated models demonstrate substantial but uneven EEG understanding, with performance varying across both model families and execution paradigms\. Averaged across models and subsets, BrainAgent achieves an overall score of 68\.78, compared with 62\.83 for CodeAct\. BrainAgent outperforms CodeAct for 10 of the 13 models on FA and 12 of the 13 models on SA, indicating that structured, domain\-oriented execution benefits most evaluated LLMs\. The advantage is not universal, however\. Gemini 3\.6 Flash achieves the highest BrainAgent overall score of 75\.17, whereas Claude Opus 5 attains the highest CodeAct score of 79\.25\. GPT\-5\.6 Sol ranks second under both paradigms, while several models exhibit substantial changes in relative performance between BrainAgent and CodeAct\. These results show that comprehensive EEG understanding is not determined by the underlying LLM alone; it emerges from the interaction between model capability, analytical task, and execution paradigm\. Moreover, the best overall score remains below 80, leaving considerable room for improvement toward reliable EEG understanding across heterogeneous workflows\.

Finding 1:Current LLMs demonstrate substantial but incomplete EEG understanding\. Structured agentic execution improves most models, yet the measured capability remains jointly determined by the LLM and its execution paradigm\.

### 4\.2How Does Task Difficulty Shape EEG Understanding?

Table[2](https://arxiv.org/html/2608.04156#S4.T2)stratifies performance by task difficulty\. Averaged across models, FA performance under BrainAgent decreases from 82\.6 on Easy tasks to 68\.9 on Medium tasks and 64\.7 on Hard tasks; the corresponding CodeAct scores decrease from 66\.0 to 64\.8 and 62\.3\. The decline is more pronounced on SA, where BrainAgent falls from 91\.0 to 72\.8 and 35\.8, while CodeAct decreases from 86\.0 to 59\.6 and 32\.5\. These results expose a clear capability boundary: current systems perform well on short and explicitly specified analyses but remain limited on extended workflows requiring multi\-step evidence integration\. The effect of the execution paradigm also changes with difficulty\. On FA, the average BrainAgent advantage decreases from 16\.6 points on Easy tasks to 4\.1 on Medium tasks and 2\.4 on Hard tasks\. On SA, the advantage is 5\.0 points on Easy tasks, peaks at 13\.2 points on Medium tasks, and falls to 3\.3 points on Hard tasks\. Structured execution therefore provides substantial support when analytical procedures can be effectively organized through domain\-oriented workflows, but its advantage diminishes on the most demanding tasks\. This pattern suggests that workflow structure can improve execution reliability but cannot fully compensate for limitations in long\-horizon reasoning and scientific evidence integration\.

Finding 2:Performance declines systematically with task difficulty, while the advantage of structured agentic execution narrows on the hardest tasks\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x5.png)Figure 4:Validation\-unit\-specific effects of execution paradigms\.Within each subset, paired markers show model\-wise unweighted mean validation scores under BA and CA, diamonds indicate across\-model means, and heatmaps report the correspondingΔ\\Deltain percentage points\.\)Table 2:Difficulty\-stratified performance under BrainAgent \(BA\) and CodeAct \(CA\)\.Δ=BA−CA\\Delta=\\mathrm\{BA\}\-\\mathrm\{CA\}; blue and orange shading denotes positive and negativeΔ\\Delta, respectively\.
### 4\.3Validation\-Unit\-Specific Effects of Execution Paradigms

Figure[4](https://arxiv.org/html/2608.04156#S4.F4)decomposes LLM\-based EEG understanding across the six validation units for FA and SA subsets\. Paired scores compare BrainAgent and CodeAct for the same model, while the mode\-gap heatmaps show the direction and magnitude of their differences, thereby separating execution\-paradigm effects from variations in underlying LLM capability\. The most consistent cross\-subset advantage of BrainAgent occurs in numerical and artifact validation\. Across model families, its structured tools generally improve the accuracy of quantitative results and the reliability of requested files relative to autonomous code generation\. The remaining units exhibit more heterogeneous mode gaps: categorical, set, sequence, and semantic performance varies with the model and subset, without a uniform advantage for either paradigm\. Structured execution therefore changes specific components of EEG analysis rather than producing an indiscriminate improvement across all output types\. Further details are provided in Appendix[E](https://arxiv.org/html/2608.04156#A5)\.

Finding 3:Structured agentic workflows strengthen execution\-grounded EEG understanding, enabling more accurate quantitative analysis and more reliable artifact production\.

### 4\.4Cross\-Instance Stability within Reusable Tasks

We assess cross\-instance consistency within reusable tasks by quantifying within\-task variability as the mean pairwise absolute difference \(MPAD\) among normalized instance scores \(the formal definition is detailed in Appendix[E](https://arxiv.org/html/2608.04156#A5)\), with lower values indicating greater stability\. Figure[5](https://arxiv.org/html/2608.04156#S4.F5)presents task\-level MPAD and its model\-level aggregation for FA and SA subsets\. Across both subsets, BrainAgent generally exhibits lower within\-task dispersion than CodeAct, with the clearest and most consistent separation observed in SA\. Although exceptions remain for individual model–task pairs, the aggregate shift toward lower MPAD suggests that capability\-oriented tools and controlled workflows reduce sensitivity to variations in recordings and analytical context\. This stability complements mean performance: it does not necessarily imply greater correctness, but indicates that a measured capability transfers more consistently across instances sharing the same analytical objective\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x6.png)Figure 5:Cross\-instance stability within reusable tasks\.For each model–task pair, the heatmap shows the mean pairwise absolute difference among within\-task normalized instance scores\. Model\-level panels report the median across tasks with task\-bootstrap 95% confidence intervals\.Finding 4:Structured agentic workflows improve the cross\-context reliability of EEG understanding, enabling consistent generalization across recordings with shared objectives\.

## 5Conclusion

We introduced BrainBench, a unified benchmark that moves EEG evaluation beyond predefined decoding targets toward comprehensive, instruction\-conditioned understanding\. Across four complementary subsets, 172 tasks, 4K real\-data instances, and 17 datasets, BrainBench assesses whether LLMs can transform analytical instructions and EEG recordings into scientifically grounded conclusions\. Evaluations of 13 representative LLMs reveal substantial differences across models, analytical capabilities, and task complexity, highlighting the remaining gap between isolated task success and comprehensive EEG understanding\. Our comparison further shows that structured agentic workflows and autonomous LLM coding are complementary rather than interchangeable: the former improves reliability through domain\-oriented execution, whereas the latter preserves analytical flexibility but is more sensitive to model capability and workflow complexity\. These findings establish BrainBench as a reproducible and extensible foundation for developing EEG\-oriented systems that better balance analytical autonomy, domain structure, and operational reliability\.

## References

- \[1\]A\. Abdou, M\. Ivanov, S\. Shaya, A\. Rueda, F\. G\. Nezhad, I\. Demchenko, M\. A\. Kamaleddin, P\. A\. Frewen, B\. T\. Dunkley, B\. Brady,et al\.\(2026\)EEG\-ai: an agentic system for ai\-assisted semi\-automated eeg preprocessing and artifact removal\.Journal of Neuroscience Methods432,pp\. 110759\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[2\]D\. Aeschbach and A\. A\. Borbely\(1993\)All\-night dynamics of the human sleep eeg\.Journal of sleep research2\(2\),pp\. 70–81\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[3\]D\. Alvarez\-Estevez and R\. M\. Rijsman\(2020\)Inter\-database validation of a deep learning approach for automatic sleep scoring\.arXiv preprint arXiv:2009\.10365\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.9.5.1),[2nd item](https://arxiv.org/html/2608.04156#S3.I1.i2.p1.1)\.
- \[4\]Anthropic\(2026\-07\)Claude Opus 5 system card\.Note:System cardExternal Links:[Link](https://www.anthropic.com/system-cards)Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[5\]H\. Banville, S\. d’Ascoli, S\. Dahan, J\. Rapin, M\. Careil, Y\. Benchetrit, J\. Lévy, S\. Panchavati, A\. Ratouchniak, E\. Cascardi,et al\.\(2026\)NeuralBench: a unifying framework to benchmark neuroai models\.arXiv preprint arXiv:2605\.08495\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[6\]D\. Baradari, N\. Kosmyna, O\. Petrov, R\. Kaplun, and P\. Maes\(2025\)NeuroChat: a neuroadaptive ai chatbot for customizing learning experiences\.InProceedings of the 7th ACM Conference on Conversational User Interfaces,pp\. 1–21\.Cited by:[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[7\]T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1)\.
- \[8\]P\. Cao, A\. Mirzazadeh, J\. W\. Lee, A\. Videnovic, and D\. Katabi\(2026\)CLEF: eeg foundation model for learning clinical semantics\.arXiv preprint arXiv:2605\.10817\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[9\]A\. Chen, A\. Li, B\. Zhou, B\. Gong, B\. Jiang, B\. Dan, C\. Yu, C\. Wang, C\. Ma, C\. Zhong,et al\.\(2026\)The minimax\-m2 series: mini activations unleashing max real\-world intelligence\.arXiv preprint arXiv:2605\.26494\.Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[10\]J\. Chen, X\. Wang, C\. Huang, X\. Hu, X\. Shen, and D\. Zhang\(2023\)A large finer\-grained affective computing eeg dataset\.Scientific Data10\(1\),pp\. 740\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.11.7.1),[3rd item](https://arxiv.org/html/2608.04156#S3.I1.i3.p1.1)\.
- \[11\]Y\. Chen, X\. Zhang, Y\. Zhang, Y\. Li, H\. Zou, C\. Miao, W\. Zhang, S\. X\. Liu, and P\. S\. Yu\(2026\)Embracing trustworthy brain\-agent collaboration as paradigm extension for intelligent assistive technologies\.Advances in Neural Information Processing Systems38\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[12\]F\. L\. da Silva\(2013\)EEG and meg: relevance to neuroscience\.Neuron80\(5\),pp\. 1112–1128\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[13\]Y\. El Ouahidi, J\. Lys, P\. Thölke, N\. Farrugia, B\. Pasdeloup, V\. Gripon, K\. Jerbi, and G\. Lioi\(2026\)REVE: a foundation model for eeg\-adapting to any setup with large\-scale pretraining on 25,000 subjects\.Advances in Neural Information Processing Systems38,pp\. 22541–22577\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[14\]M\. M\. Ghassemi, B\. E\. Moody, L\. H\. Lehman, C\. Song, Q\. Li, H\. Sun, R\. G\. Mark, M\. B\. Westover, and G\. D\. Clifford\(2018\)You snooze, you win: the physionet/computing in cardiology challenge 2018\.In2018 Computing in Cardiology Conference \(CinC\),Vol\.45,pp\. 1–4\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.2.2.2.2.2.2.2.2.2.2.2.2),[2nd item](https://arxiv.org/html/2608.04156#S3.I1.i2.p1.1)\.
- \[15\]Google DeepMind\(2026\-07\)Gemini 3\.6 Flash model card\.Note:Model cardExternal Links:[Link](https://deepmind.google/models/model-cards/gemini-3-6-flash/)Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[16\]W\. Gu, L\. Tianming, Q\. Zhang, M\. Ye, X\. Shen, W\. Chen, Y\. Li, Y\. Zhang, J\. Hong, B\. Lu,et al\.CerebraGloss: instruction\-tuning a large vision\-language model for fine\-grained clinical eeg interpretation\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1)\.
- \[17\]M\. F\. Hinss, E\. S\. Jahanpour, B\. Somon, L\. Pluchon, F\. Dehais, and R\. N\. Roy\(2023\)Open multi\-session and multi\-task eeg cognitive dataset for passive brain\-computer interface applications\.Scientific Data10\(1\),pp\. 85\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.13.9.1),[3rd item](https://arxiv.org/html/2608.04156#S3.I1.i3.p1.1)\.
- \[18\]J\. Hong, W\. Wang, and L\. Najafizadeh\(2024\)ChatBCI: a p300 speller bci leveraging large language models for improved sentence composition in realistic scenarios\.arXiv preprint arXiv:2411\.15395\.Cited by:[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[19\]J\. Jeong\(2004\)EEG dynamics in patients with alzheimer’s disease\.Clinical neurophysiology115\(7\),pp\. 1490–1505\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[20\]J\. Jeong, J\. Cho, Y\. Lee, S\. Lee, G\. Shin, Y\. Kweon, J\. d\. R\. Millán, K\. Müller, and S\. Lee\(2022\)2020 international brain–computer interface competition: a review\.Frontiers in human neuroscience16,pp\. 898300\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.5.1.1),[1st item](https://arxiv.org/html/2608.04156#S3.I1.i1.p1.1)\.
- \[21\]W\. Jiang, X\. Liu, W\. Zheng, and B\. Lu\(2024\)SEED\-vii: a multimodal dataset of six basic emotions with continuous labels for emotion recognition\.IEEE Transactions on Affective Computing16\(2\),pp\. 969–985\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.15.11.1),[4th item](https://arxiv.org/html/2608.04156#S3.I1.i4.p1.1)\.
- \[22\]W\. Jiang, L\. Zhao, and B\. Lu\(2024\)Large brain model for learning generic representations with tremendous eeg data in bci\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 16405–16426\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[23\]N\. Kane, J\. Acharya, S\. Beniczky, L\. Caboclo, S\. Finnigan, P\. W\. Kaplan, H\. Shibasaki, R\. Pressler, and M\. J\. Van Putten\(2017\)A revised glossary of terms most commonly used by clinical electroencephalographers and updated proposal for the report format of the eeg findings\. revision 2017\.Clinical neurophysiology practice2,pp\. 170–185\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[24\]A\. Kastrati, J\. Bürki, J\. Lauer, C\. Xuan, R\. Iaquinto, and R\. Wattenhofer\(2025\)EEG\-bench: a benchmark for eeg foundation models in clinical applications\.arXiv preprint arXiv:2512\.08959\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[25\]A\. Keil, S\. Debener, G\. Gratton, M\. Junghöfer, E\. S\. Kappenman, S\. J\. Luck, P\. Luu, G\. A\. Miller, and C\. M\. Yee\(2014\)Committee report: publication guidelines and recommendations for studies using electroencephalography and magnetoencephalography\.Psychophysiology51\(1\),pp\. 1–21\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[26\]S\. Khalighi, T\. Sousa, J\. M\. Santos, and U\. Nunes\(2016\)ISRUC\-sleep: a comprehensive public dataset for sleep researchers\.Computer methods and programs in biomedicine124,pp\. 180–192\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.1.1.1.1.1.1.1.1.1.1.1.2),[1st item](https://arxiv.org/html/2608.04156#S3.I1.i1.p1.1),[2nd item](https://arxiv.org/html/2608.04156#S3.I1.i2.p1.1)\.
- \[27\]S\. Koelstra, C\. Muhl, M\. Soleymani, J\. Lee, A\. Yazdani, T\. Ebrahimi, T\. Pun, A\. Nijholt, and I\. Patras\(2011\)Deap: a database for emotion analysis; using physiological signals\.IEEE transactions on affective computing3\(1\),pp\. 18–31\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.16.12.1),[4th item](https://arxiv.org/html/2608.04156#S3.I1.i4.p1.1)\.
- \[28\]K\. Kontras, T\. Osselaer, S\. G\. Mouslech, A\. Karaiskou, G\. Gagliardi, T\. Strypsteen, M\. H\. Badiei, A\. Rani, M\. Vanmarcke, M\. Bhagubai,et al\.\(2026\)NeuroAtlas: benchmarking foundation models for clinical eeg and brain\-computer interfaces\.arXiv preprint arXiv:2605\.14698\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[29\]N\. Kosmyna and E\. Hauptmann\(2026\)NeuroSkill \(tm\): proactive real\-time agentic system capable of modeling human state of mind\.arXiv preprint arXiv:2603\.03212\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[30\]V\. J\. Lawhern, A\. J\. Solon, N\. R\. Waytowich, S\. M\. Gordon, C\. P\. Hung, and B\. J\. Lance\(2016\)EEGNet: a compact convolutional network for eeg\-based brain\-computer interfaces\.arXiv preprint arXiv:1611\.08024\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[31\]H\. Li, T\. Gao, J\. Li, Y\. Fan, R\. Shi, W\. Wang, T\. Zhao, Z\. Wu, X\. Jiang, Q\. Zhang,et al\.\(2026\)BrainPilot: automating brain discovery with agentic research\.arXiv preprint arXiv:2607\.15079\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2608.04156#S2.SS3.p1.1)\.
- \[32\]J\. Li, C\. Fu, J\. Tang, L\. Zhou, H\. Chen, W\. Zhou, C\. Chen, and J\. Luo\(2026\)Multimodal phenotyping dataset of driving fatigue\.Scientific Data\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.14.10.1),[3rd item](https://arxiv.org/html/2608.04156#S3.I1.i3.p1.1)\.
- \[33\]S\. Li, S\. Xiao, M\. Joshi, A\. Metwally, D\. McDuff, W\. Wang, and Y\. Yang\(2026\)Hearts: benchmarking llm reasoning on health time series\.arXiv preprint arXiv:2603\.06638\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.04156#S2.SS3.p1.1)\.
- \[34\]W\. Liu, J\. Qiu, W\. Zheng, and B\. Lu\(2021\)Comparing recognition performance and robustness of multimodal deep learning models for multimodal emotion recognition\.IEEE Transactions on Cognitive and Developmental Systems14\(2\),pp\. 715–729\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.6.2.1),[1st item](https://arxiv.org/html/2608.04156#S3.I1.i1.p1.1)\.
- \[35\]Z\. Lu, Z\. Li, X\. Shen, K\. Lou, Y\. Xin, X\. Chen, S\. Wang, X\. Chen, J\. Fan, C\. Huang,et al\.\(2026\)OmniEEG\-bench: a standardized evaluation benchmark for eeg foundation models\.arXiv preprint arXiv:2606\.00815\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[36\]W\. Mumtaz, L\. Xia, S\. S\. A\. Ali, M\. A\. M\. Yasin, M\. Hussain, and A\. S\. Malik\(2017\)Electroencephalogram \(eeg\)\-based computer\-aided technique to diagnose major depressive disorder \(mdd\)\.Biomedical Signal Processing and Control31,pp\. 108–115\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.7.3.1),[1st item](https://arxiv.org/html/2608.04156#S3.I1.i1.p1.1)\.
- \[37\]X\. Ning, J\. Wang, Z\. Feng, T\. Xin, S\. Zhang, S\. Zhang, Z\. Lian, Y\. Ding, Y\. Lin, and Z\. Jia\(2026\)REFED: a subject real\-time dynamic labeled eeg\-fnirs synchronized recorded emotion dataset\.Advances in Neural Information Processing Systems38\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.12.8.1),[3rd item](https://arxiv.org/html/2608.04156#S3.I1.i3.p1.1)\.
- \[38\]C\. O’reilly, N\. Gosselin, J\. Carrier, and T\. Nielsen\(2014\)Montreal archive of sleep studies: an open\-access resource for instrument benchmarking and exploratory research\.Journal of sleep research23\(6\),pp\. 628–635\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.10.6.1),[2nd item](https://arxiv.org/html/2608.04156#S3.I1.i2.p1.1)\.
- \[39\]OpenAI\(2026\-07\)GPT\-5\.6 system card\.Note:System cardExternal Links:[Link](https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf)Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[40\]C\. Pernet, M\. I\. Garrido, A\. Gramfort, N\. Maurits, C\. M\. Michel, E\. Pang, R\. Salmelin, J\. M\. Schoffelen, P\. A\. Valdes\-Sosa, and A\. Puce\(2020\)Issues and recommendations from the ohbm cobidas meeg committee for reproducible eeg and meg research\.Nature neuroscience23\(12\),pp\. 1473–1483\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[41\]J\. Pradeepkumar, Z\. Chen, and J\. Sun\(2026\)Neural signals generate clinical notes in the wild\.arXiv preprint arXiv:2601\.22197\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1)\.
- \[42\]S\. F\. Quan, B\. V\. Howard, C\. Iber, J\. P\. Kiley, F\. J\. Nieto, G\. T\. O’Connor, D\. M\. Rapoport, S\. Redline, J\. Robbins, J\. M\. Samet,et al\.\(1997\)The sleep heart health study: design, rationale, and methods\.Sleep20\(12\),pp\. 1077–1085\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.3.2),[2nd item](https://arxiv.org/html/2608.04156#S3.I1.i2.p1.1),[4th item](https://arxiv.org/html/2608.04156#S3.I1.i4.p1.1)\.
- \[43\]Qwen Team\(2026\-02\)Qwen3\.5: towards native multimodal agents\.Note:Technical reportExternal Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[44\]Qwen Team\(2026\-05\)Qwen3\.7\-Plus: multimodal agent intelligence\.Note:Technical reportExternal Links:[Link](https://qwen.ai/blog?id=qwen3.7-plus)Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[45\]Qwen Team\(2026\-05\)Qwen3\.7: the agent frontier\.Note:Technical reportExternal Links:[Link](https://qwen.ai/blog?id=qwen3.7)Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[46\]G\. Schalk, D\. J\. McFarland, T\. Hinterberger, N\. Birbaumer, and J\. R\. Wolpaw\(2004\)BCI2000: a general\-purpose brain\-computer interface \(bci\) system\.IEEE Transactions on biomedical engineering51\(6\),pp\. 1034–1043\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[47\]T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. Scialom\(2023\)Toolformer: language models can teach themselves to use tools\.Advances in Neural Information Processing Systems36,pp\. 68539–68551\.Cited by:[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[48\]F\. Shen, E\. Yang, J\. Li, J\. Hong, X\. Pan, Z\. Yuan, M\. Li, and Y\. Yang\(2026\)Brain4FMs: a benchmark of foundation models for electrical brain signal\.arXiv preprint arXiv:2602\.11558\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[49\]J\. Shin, A\. Von Lühmann, D\. Kim, J\. Mehnert, H\. Hwang, and K\. Müller\(2018\)Simultaneous acquisition of eeg and nirs during cognitive tasks for an open access dataset\.Scientific data5\(1\),pp\. 180003\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.17.13.1),[4th item](https://arxiv.org/html/2608.04156#S3.I1.i4.p1.1)\.
- \[50\]N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. Yao\(2023\)Reflexion: language agents with verbal reinforcement learning\.Advances in Neural Information Processing Systems36,pp\. 8634–8652\.Cited by:[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[51\]K\. Team, T\. Bai, Y\. Bai, Y\. Bao, S\. Cai, Y\. Cao, Y\. Charles, H\. Che, C\. Chen, G\. Chen,et al\.\(2026\)Kimi k2\. 5: visual agentic intelligence\.arXiv preprint arXiv:2602\.02276\.Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[52\]G\. Wang, S\. Yang, J\. Ding, and F\. Liu\(2026\)Neuroweaver: an autonomous evolutionary agent for exploring the programmatic space of eeg analysis pipelines\.arXiv preprint arXiv:2602\.13473\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[53\]J\. Wang, S\. Zhao, Z\. Luo, Y\. Zhou, H\. Jiang, S\. Li, T\. Li, and G\. Pan\(2025\)Cbramod: a criss\-cross brain foundation model for eeg decoding\.InInternational conference on learning representations,Vol\.2025,pp\. 75310–75346\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[54\]J\. Wang, S\. Zhao, Z\. Luo, Y\. Zhou, S\. Li, and G\. Pan\(2025\)Eegmamba: an eeg foundation model with mamba\.Neural Networks,pp\. 107816\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[55\]J\. Wang, S\. Zhao, Y\. Zhou, Y\. Kang, S\. Li, and G\. Pan\(2026\)DeeperBrain: a neuro\-grounded eeg foundation model towards universal bci\.arXiv preprint arXiv:2601\.06134\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[56\]X\. Wang, Y\. Yang, and D\. Coyle\(2026\)EEG\-fm\-audit: a systematic evaluation and analysis pipeline for eeg foundation models\.arXiv preprint arXiv:2605\.26910\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[57\]X\. Wang, Y\. Chen, L\. Yuan, Y\. Zhang, Y\. Li, H\. Peng, and H\. Ji\(2024\)Executable code actions elicit better llm agents\.InForty\-first International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p5.1),[§3\.2\.4](https://arxiv.org/html/2608.04156#S3.SS2.SSS4.p1.1),[§3](https://arxiv.org/html/2608.04156#S3.p1.1)\.
- \[58\]J\. Wu, Z\. Ren, J\. Wang, P\. Zhu, Y\. Song, M\. Liu, Q\. Zheng, L\. Bai, W\. Ouyang, and C\. Song\(2025\)Adabrain\-bench: benchmarking brain foundation models for brain\-computer interface applications\.arXiv preprint arXiv:2507\.09882\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[59\]Q\. Xiao, Z\. Cui, C\. Zhang, S\. Chen, W\. Wu, A\. Thwaites, A\. Woolgar, B\. Zhou, and C\. Zhang\(2026\)Brainomni: a brain foundation model for unified eeg and meg signals\.Advances in Neural Information Processing Systems38,pp\. 41179–41212\.Cited by:[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[60\]W\. Xiong, J\. Li, J\. Li, and K\. Zhu\(2025\)Eeg\-fm\-bench: a comprehensive benchmark for the systematic evaluation of eeg foundation models\.arXiv preprint arXiv:2508\.17742\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[61\]A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.\(2026\)Deepseek\-v4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[62\]Z\. Xu, Z\. Shuai, E\. Mozaffari, R\. S\. Aysola, R\. Kumar, and Y\. Yang\(2026\)Sleeplm: natural\-language intelligence for human sleep\.InForty\-third International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[63\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\(2022\)React: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1)\.
- \[64\]A\. Zeng, X\. Lv, Z\. Hou, Z\. Du, Q\. Zheng, B\. Chen, D\. Yin, C\. Ge, C\. Huang, C\. Xie,et al\.\(2026\)Glm\-5: from vibe coding to agentic engineering\.arXiv preprint arXiv:2602\.15763\.Cited by:[§4](https://arxiv.org/html/2608.04156#S4.p1.1)\.
- \[65\]S\. Zhao, M\. Peng, H\. Jiang, T\. Li, and S\. Li\(2026\)EEG agent: a unified framework for automated eeg analysis using large language models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 18063–18071\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2608.04156#S2.SS3.p1.1)\.
- \[66\]W\. Zheng and B\. Lu\(2016\)A multimodal approach to estimating vigilance using eeg and forehead eog\.arXiv preprint arXiv:1611\.08492\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.18.14.1),[4th item](https://arxiv.org/html/2608.04156#S3.I1.i4.p1.1)\.
- \[67\]Y\. Zhou, S\. Zhao, J\. Wang, H\. Jiang, S\. Li, B\. Luo, T\. Li, and G\. Pan\(2025\)Personalized sleep staging leveraging source\-free unsupervised domain adaptation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 14529–14537\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1)\.
- \[68\]Y\. Zhou, S\. Zhao, J\. Wang, S\. Li, and G\. Pan\(2026\)BrainAgent: a large language model\-driven multi\-agent framework for autonomous brain signal understanding\.arXiv preprint arXiv:2606\.25400\.Cited by:[§D\.2\.1](https://arxiv.org/html/2608.04156#A4.SS2.SSS1.p1.1),[§1](https://arxiv.org/html/2608.04156#S1.p2.1),[§1](https://arxiv.org/html/2608.04156#S1.p5.1),[§2\.2](https://arxiv.org/html/2608.04156#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2608.04156#S2.SS3.p1.1),[§3\.2\.4](https://arxiv.org/html/2608.04156#S3.SS2.SSS4.p1.1),[§3](https://arxiv.org/html/2608.04156#S3.p1.1)\.
- \[69\]Y\. Zhou, J\. Wu, Z\. Ren, Z\. Yao, W\. Lu, K\. Peng, Q\. Zheng, C\. Song, W\. Ouyang, and C\. Gou\(2026\)Csbrain: a cross\-scale spatiotemporal brain foundation model for eeg decoding\.Advances in Neural Information Processing Systems38,pp\. 87150–87195\.Cited by:[§1](https://arxiv.org/html/2608.04156#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.04156#S2.SS1.p1.1)\.
- \[70\]I\. Zyma, S\. Tukaev, I\. Seleznov, K\. Kiyono, A\. Popov, M\. Chernykh, and O\. Shpenkov\(2019\)Electroencephalograms during mental arithmetic task performance\.Data4\(1\),pp\. 14\.Cited by:[Table 3](https://arxiv.org/html/2608.04156#A1.T3.3.3.3.3.3.3.3.3.3.3.8.4.1),[1st item](https://arxiv.org/html/2608.04156#S3.I1.i1.p1.1)\.

## Appendix ADataset Statistics

BrainBench incorporates 17 unique datasets spanning research\-oriented EEG recordings, clinical polysomnography, neurocognitive experiments, and multimodal neurophysiological measurements\. For each dataset,BrainBench includes recordings from the first five subjects in the source\-defined ordering\. The sole exception is MDP\-DF, for which subjects 1–4 and 6are used because the fifth subject was excluded due to data\-quality issues\. Table[3](https://arxiv.org/html/2608.04156#A1.T3)summarizes the subset assignment, signal modalities and their native sampling rates, number of EEG channels, and accessibility of each dataset\. We do not apply additional filtering, resampling, re\-referencing, or artifact removal to the source\-distributed signals\. File\-format harmonization and standardized naming are applied where necessary to support benchmark construction and data management\. For datasets containing signals acquired at different sampling rates, each modality retains its native temporal resolution\.

Table 3:Dataset composition and acquisition characteristics of BrainBench, including subset assignment, retained signal modalities and native sampling rates, EEG channel counts, and data accessibility\. FA, SA, NA, and PI denote Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration, respectively\.Openindicates access without case\-by\-case approval, whereasRestrictedindicates access requiring an application, institutional affiliation, ethical approval, or a data\-use agreement\.
## Appendix BTask Design Details

This appendix provides a task\-level inventory of BrainBench\. Each task is distilled from representative EEG analysis capabilities and practical analytical workflows, and each entry specifies the globally unique task identifier, difficulty level, number of associated instances, assessment objective, and validation configuration\.All task contents were cross\-reviewed by three domain experts, who additionally verified the consistency among the instruction, ground\-truth script, and validation configuration\.E, M, and H denote Easy, Medium, and Hard, respectively\. Validation units are abbreviated as follows: Num\. for numerical validation, Cat\. for categorical validation, Set for unordered\-set validation, Seq\. for ordered\-sequence validation, Sem\. for semantic validation, and File for artifact validation\. The notation×n\\times nindicates the number of validation units of a given type, while “\+” denotes the joint use of multiple validation\-unit types within a task\.\.

### B\.1Foundational Analysis

As shown in Table \.LABEL:tab:fa\_task\_inventory, the Foundational Analysis subset comprises 40tasksand 950instancesconstructed from ISRUC, BCIC2020\-3, SEED\-V, Mumtaz2016, and MentalArithmetic\. It evaluates recording inspection, spectral and nonlinear feature extraction, preprocessing and artifact generation, channel\- and region\-level comparison, connectivity analysis, robustness to invalid requests, and foundational knowledge\.

Table 4:Task\-level inventory of the Foundational Analysis subset\.Task IDDiff\.\# Inst\.Assessment contentValidation unit\(s\)FA\-01E25Estimate mean channel\-wise alpha relative power after standardized EEG filtering and PSD integration\.Num\.×1\\times 1FA\-02E25Quantify prefrontal alpha/beta and theta/beta power ratios using available frontal channels\.Num\.×2\\times 2FA\-03E25Measure short\-window occipital signal complexity using sample entropy and aggregate it across channels\.Num\.×1\\times 1FA\-04M25Execute a multistep central\-channel preprocessing pipeline and compute mean Hjorth mobility and complexity\.Num\.×2\\times 2FA\-05M25Compute and compare regional spectral edge frequency \(SEF95\) over parietal/central and occipital channels\.Num\.×2\\times 2FA\-06M20Quantify low\-frequency baseline drift and classify its severity without removing the target phenomenon\.Cat\.×1\\times 1FA\-07M5Identify the EEG electrode\-placement system and detect channels outside the corresponding standard montage\.Cat\.×1\\times 1\+ Set×1\\times 1FA\-08E5Determine whether the available channel montage supports left–right spectral asymmetry analysis\.Cat\.×1\\times 1FA\-09M25Rank the three EEG channel pairs with the strongest broadband Pearson correlations\.Seq\.×1\\times 1FA\-10M25Characterize alpha\-band inter\-channel synchronization and identify the strongest correlated pair\.Num\.×1\\times 1\+ Cat\.×1\\times 1FA\-11E25Identify the dominant and secondary EEG frequency bands within a specified recording segment\.Cat\.×2\\times 2FA\-12M25Rank channels by delta\-band variance after filtering and resampling\.Seq\.×1\\times 1FA\-13M25Rank channels independently by signal kurtosis and skewness after standardized preprocessing\.Seq\.×2\\times 2FA\-14E25Extract, filter, resample, and export a frontal EEG segment as an EDF artifact\.File×1\\times 1FA\-15E25Select genuine EEG channels, apply average referencing, and export a fixed\-duration EDF artifact\.File×1\\times 1FA\-16M25Select the channel with maximal alpha\-relative energy, isolate its alpha component, and export it as an array\.File×1\\times 1FA\-17M25Apply line\-noise suppression and generate a whole\-recording EEG PSD visualization\.File×1\\times 1FA\-18M25Generate a band\-limited PSD visualization for a specified pair of EEG channels\.File×1\\times 1FA\-19M25Compute and visualize inter\-channel correlation for a specified multichannel segment\.File×1\\times 1FA\-20E5Handle an unavailable EEG file safely without fabricating a brain\-state interpretation\.Sem\.×1\\times 1FA\-21E15Recognize that a requested channel is absent and respond without inventing channel\-level analysis\.Sem\.×1\\times 1FA\-22E25Detect that a requested time window lies outside the recording and avoid unsupported analysis\.Sem\.×1\\times 1FA\-23M25Verify a supplied dominant\-band claim against the signal and resist an incorrect premise\.Cat\.×1\\times 1\+ Sem\.×1\\times 1FA\-24M25Detect when preprocessing removes the frequency content required by the requested downstream analysis\.Sem\.×1\\times 1FA\-25M25Quantify notch\-filter attenuation and determine whether suppression succeeds, including mismatched\-frequency controls\.Num\.×1\\times 1\+ Cat\.×1\\times 1FA\-26H25Compare frontal and occipital alpha\-relative power and identify the region with stronger activity\.Num\.×1\\times 1\+ Cat\.×1\\times 1FA\-27H25Compare regional band\-power dominance across two time windows and interpret the spatial\-state transition\.Cat\.×2\\times 2\+ Sem\.×1\\times 1FA\-28H25Rank brain regions by alpha\-relative power and assess whether alpha activity is posterior dominant\.Seq\.×1\\times 1\+ Cat\.×1\\times 1\+ Sem\.×1\\times 1FA\-29H25Compute alpha\-band phase\-locking connectivity, global mean PLV, and the strongest channel pairs\.Num\.×1\\times 1\+ Set×1\\times 1FA\-30H25Compare global phase locking across four canonical bands and explain the dominant synchronization band\.Num\.×4\\times 4\+ Sem\.×1\\times 1FA\-31H25Compare frontal alpha/beta ratios across consecutive windows and infer the direction of attention\-state change\.Num\.×2\\times 2\+ Sem\.×1\\times 1FA\-32H25Contrast global and occipital alpha rankings to assess whether global aggregation masks regional structure\.Cat\.×2\\times 2\+ Sem\.×1\\times 1FA\-33M25Remove variance\-outlier channels and return channels whose alpha\-relative power exceeds the retained\-channel mean\.Set×1\\times 1FA\-34H25Estimate occipital alpha peak frequency and derive personalized alpha\-band relative power\.Num\.×2\\times 2FA\-35M25Execute standardized preprocessing and locate maximum global\-field\-power peaks in three windows\.Num\.×3\\times 3FA\-36H25Compare left\- and right\-hemisphere wPLI in alpha and beta bands and assess cross\-band consistency\.Num\.×2\\times 2\+ Sem\.×1\\times 1FA\-37M25Identify channels that repeatedly appear among the highest relative\-power channels across multiple bands\.Set×1\\times 1FA\-38H25Build an alpha\-band PLV network and rank channels by weighted connectivity degree\.Seq\.×1\\times 1FA\-39H25Rank channels by alpha power and infer the dominant anatomical region from the leading channels\.Seq\.×1\\times 1\+ Set×1\\times 1FA\-40E50Answer a standalone foundational EEG and BCI knowledge question without requiring a signal file\.Cat\.×1\\times 1
### B\.2Sleep Assessment

As shown in Table\.LABEL:tab:sa\_task\_inventory, the Sleep Assessment subset comprises 43tasksand 1,025instancesconstructed from HMC, ISRUC, MASS\-SS3, PhysioNet 2018, and SHHS\-1\. It evaluates sleep architecture, staging, spectral and temporal analysis, PSG artifact generation, arousal and respiratory\-event analysis, oxygenation, and sleep\-medicine knowledge\.

Table 5:Task\-level inventory of the Sleep Assessment subset\.Task IDDiff\.\# Inst\.Assessment contentValidation unit\(s\)SA\-01E25Calculate sleep onset latency from epoch\-level sleep\-stage labels\.Num\.×1\\times 1SA\-02E15Inventory all PSG channels and distinguish genuine EEG channels from auxiliary sensors\.Num\.×1\\times 1\+ Set×1\\times 1SA\-03E25Derive total sleep time, time in bed, and sleep efficiency from sleep\-stage labels\.Num\.×3\\times 3SA\-04E25Quantify wake after sleep onset and count post\-onset awakening bouts\.Num\.×2\\times 2SA\-05E25Compute NREM duration, REM proportion, and REM latency from whole\-night staging\.Num\.×3\\times 3SA\-06E25Identify the dominant and secondary sleep stages over the full recording\.Cat\.×2\\times 2SA\-07M25Quantify the light\-to\-deep sleep ratio and interpret its implication for sleep architecture\.Num\.×1\\times 1\+ Sem\.×1\\times 1SA\-08M25Find the three longest uninterrupted sleep episodes and interpret whole\-night continuity\.Num\.×3\\times 3\+ Sem\.×1\\times 1SA\-09E25Count all sleep\-stage transitions and deep\-sleep\-to\-wake interruptions\.Num\.×2\\times 2SA\-10E25Count sustained N3 bouts and calculate their mean duration\.Num\.×2\\times 2SA\-11M25Locate the longest nocturnal wake interruption and classify its continuity impact\.Num\.×2\\times 2\+ Cat\.×1\\times 1SA\-12M25Quantify first\-half N3 and second\-half REM concentration and interpret the overnight pattern\.Num\.×2\\times 2\+ Sem\.×1\\times 1SA\-13M25Calculate a whole\-night sleep fragmentation index from awakenings and non\-wake stage shifts\.Num\.×1\\times 1SA\-14E25Calculate the whole\-sequence percentage distribution of W, N1, N2, N3, and REM\.Num\.×5\\times 5SA\-15H25Screen the sleep\-stage sequence for a predefined panel of twelve sleep\-structure abnormalities\.Cat\.×12\\times 12SA\-16E25Determine the dominant EEG band and sleep stage in a specified recording segment\.Cat\.×2\\times 2SA\-17M25Infer the sleep stages of ten consecutive epochs from the sleep recording\.Seq\.×1\\times 1SA\-18E25Generate and save a whole\-night sleep hypnogram\.File×1\\times 1SA\-19E25Execute a sleep\-staging preprocessing pipeline and export fixed\-length EEG epochs as an array\.File×1\\times 1SA\-20E25Generate and save a sleep EEG spectrogram from the available EEG channels\.File×1\\times 1SA\-21M25Compare delta\-relative power across three windows and identify the most slow\-wave\-rich segment\.Num\.×3\\times 3\+ Cat\.×1\\times 1SA\-22M25Compare event\-window and whole\-night chin\-EMG activity to assess whether the segment is REM\-like\.Num\.×1\\times 1\+ Sem\.×1\\times 1SA\-23M25Compare first\- and second\-half delta\-relative energy and assess consistency with canonical sleep architecture\.Num\.×2\\times 2\+ Sem\.×1\\times 1SA\-24M25Perform whole\-recording automatic sleep staging and summarize sleep efficiency, N3, and REM proportions\.Num\.×3\\times 3SA\-25M25Test a supplied band\-and\-stage claim against the signal and reject unsupported conclusions\.Sem\.×1\\times 1SA\-26M25Rank three segments by EOG activity and determine whether the strongest segment represents REM sleep\.Seq\.×1\\times 1\+ Cat\.×1\\times 1SA\-27M25Rank three segments by EMG activity and identify whether and where REM sleep is present\.Seq\.×1\\times 1\+ Sem\.×1\\times 1SA\-28M25Estimate whole\-night mean, minimum, and maximum heart rate from the ECG channel\.Num\.×3\\times 3SA\-29H40Detect arousal independently in five specified sleep segments\.Cat\.×5\\times 5SA\-30H40Classify five specified respiratory segments as apnea, hypopnea, or no target event\.Cat\.×5\\times 5SA\-31H10Calculate the whole\-night arousal index using detected arousals and total sleep time\.Num\.×1\\times 1SA\-32H15Calculate the whole\-night apnea–hypopnea index from respiratory events and total sleep time\.Num\.×1\\times 1SA\-33H15Quantify respiratory disturbance or detect respiratory\-effort\-related arousals in selected segments\.Num\.×1\\times 1; or Cat\.×5\\times 5SA\-34M9Identify the longest apnea event and report its duration, onset, and sleep stage\.Num\.×2\\times 2\+ Cat\.×1\\times 1SA\-35M46Distinguish central, obstructive, and mixed apnea subtypes in two specified segments\.Cat\.×2\\times 2SA\-36H13Calculate stage\-specific apnea–hypopnea indices for REM and NREM sleep\.Num\.×2\\times 2SA\-37H40Reconstruct respiratory\-event and arousal chronology and assess its likely impact on sleep continuity\.Sem\.×1\\times 1SA\-38H10Quantify N3\-specific arousal burden and transitions and assess deep\-sleep disruption\.Num\.×2\\times 2\+ Sem\.×1\\times 1SA\-39H15Detect respiratory\-event clusters in fixed windows and determine their dominant sleep stage\.Num\.×1\\times 1\+ Cat\.×1\\times 1SA\-40M14Calculate sleep\-period oxygen desaturation indices using 3% and 4% thresholds\.Num\.×2\\times 2SA\-41M14Derive sleep/wake mean oxygen saturation and minimum sleep oxygen saturation after signal cleaning\.Num\.×3\\times 3SA\-42M14Quantify cumulative sleep time below 90% and 80% oxygen saturation\.Num\.×2\\times 2SA\-43E40Answer a standalone sleep\-medicine and polysomnography knowledge question\.Cat\.×1\\times 1
### B\.3Neurocognitive Assessment

The Neurocognitive Assessment subset comprises 50tasksand 1,030instancesconstructed from FACED, REFED, COG\-BCI, and MPD\-DF\. It includes 20 affective\-state tasks, 15 cognitive\-load tasks, and 15 fatigue and vigilance tasks, covering state recognition, feature\-based and temporal comparison, multimodal evidence integration, claim verification, artifact generation, and domain knowledge\. Tasks NA\-20, NA\-35, and NA\-50 are knowledge\-only tasks; the remaining tasks are grounded in neurophysiological recordings or associated behavioral annotations shown in Table\.LABEL:tab:na\_task\_inventory\.

Table 6:Task\-level inventory of the Neurocognitive Assessment subset\.Task IDDiff\.\# Inst\.Assessment contentValidation unit\(s\)NA\-01E30Determine the emotional polarity of a specified EEG segment as positive, neutral, or negative\.Cat\.×1\\times 1NA\-02E20Classify the arousal or valence level of a specified EEG segment as high or low\.Cat\.×1\\times 1NA\-03E30Compute frontal alpha asymmetry from F3/F4 alpha power in a specified EEG segment\.Num\.×1\\times 1NA\-04E10Determine whether a statement about emotional polarity is supported by SAM arousal and valence ratings\.Sem\.×1\\times 1NA\-05M15Correlate channel\-band relative\-power features with continuous SAM ratings and identify the strongest positively correlated pairs\.Set×1\\times 1NA\-06M20Compare two trials using beta/alpha relative\-power patterns, identify discriminative channels, determine the stronger target\-emotion trial, and explain the evidence\.Set×1\\times 1\+ Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-07M20Compare positive and negative trials, identify the three channels with the largest alpha\-power differences, and determine their dominant brain region\.Set×1\\times 1\+ Cat\.×1\\times 1NA\-08M20Compute frontal alpha asymmetry across homologous pairs, determine the primary valence direction, and explain the evidence\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-09M20Determine whether an emotional state changes within a continuous EEG segment and explain the temporal evidence\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-10M20Fit temporal trends of emotion\-related EEG features and assess whether they agree with the known trial label\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-11H20Compare positive and negative trials using theta\-band coherence, identify the three channel pairs with the largest negative\-trial increase, and explain the result\.Set×1\\times 1\+ Sem\.×1\\times 1NA\-12M20Compare an emotional trial with a neutral baseline and assess whether the target emotion exhibits the specified EEG pattern more strongly\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-13H20Rank three trials by emotion\-related EEG state or intensity and identify the emotion category of the highest\-ranked trial\.Seq\.×1\\times 1\+ Cat\.×1\\times 1NA\-14H20Compare positive and negative trial groups, identify the most stable discriminative channel set and its dominant region, and explain the evidence\.Set×1\\times 1\+ Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-15M20Evaluate a statement about alpha\-band hemispheric asymmetry using left–right channel\-power evidence\.Cat\.×1\\times 1NA\-16M20Compare a target trial with a neutral baseline and determine whether a supplied emotion\-EEG statement is supported\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-17H20Correlate multiband relative power with arousal or valence ratings, generate a heatmap, and identify the two strongest positively correlated channel\-band pairs\.File×1\\times 1\+ Set×1\\times 1NA\-18M20Compare sample entropy between two trials and determine which trial exhibits greater signal complexity\.Num\.×1\\times 1\+ Cat\.×1\\times 1NA\-19H20Determine which of three trials best matches a specified emotional state or maximizes the target EEG index\.Cat\.×1\\times 1NA\-20E40Answer a standalone EEG and BCI emotion\-recognition knowledge question\.Cat\.×1\\times 1NA\-21E20Classify a continuous N\-back EEG segment as low or high cognitive workload\.Cat\.×1\\times 1NA\-22M20Compare low\- and high\-workload EEG using frontal theta power and interpret whether the observed direction agrees with cognitive\-load physiology\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-23M15Compare behavioral performance across two N\-back conditions and determine which condition imposes greater behavioral load\.Num\.×1\\times 1\+ Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-24H20Rank three EEG segments by frontal\-theta/parietal\-alpha ratio and explain its relation to cognitive workload\.Seq\.×1\\times 1\+ Sem\.×1\\times 1NA\-25M15Compare two within\-condition segments using frontal\-theta/parietal\-alpha ratio and determine whether either is higher under a specified equivalence margin\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-26M15Compare beta/\(alpha\+theta\) engagement between two segments over centroparietal channels\.Cat\.×1\\times 1NA\-27H15Correlate theta, alpha, and beta power sequences with binary group labels and rank the two strongest frequency\-band associations\.Seq\.×1\\times 1\+ Sem\.×1\\times 1NA\-28M20Compare target\-minus\-nontarget ERP amplitudes between workload groups and interpret the difference\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-29M20Compute regional relative theta during a high\-load block, identify the dominant region, and interpret its working\-memory relevance\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-30H20Compare P300 amplitude and latency between two stimulus\-locked segments and infer their relative cognitive load\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-31H20Integrate behavioral errors and reaction time with EEG theta/alpha ratio to assess agreement between behavioral and neural workload evidence\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-32H20Evaluate a compound claim about cognitive load and engagement using theta/alpha and beta/\(alpha\+theta\) ratios\.Cat\.×1\\times 1\+ Sem\.×1\\times 1NA\-33M15Determine whether a cognitive\-task or workload\-state change occurs within a continuous EEG window\.Cat\.×1\\times 1NA\-34H15Compute target and nontarget P300/P3b waveforms at Pz and export their joint visualization\.File×1\\times 1NA\-35E40Answer a standalone EEG and BCI cognitive\-load knowledge question\.Cat\.×1\\times 1NA\-36E15Compute blink rate, mean blink duration, and slow\-eye\-movement power from an EOG window\.Num\.×1\\times 1NA\-37M20Compare two EOG windows for fatigue\-related long eye closure, slow eye movements, or blink behavior\.Cat\.×1\\times 1NA\-38M20Compare band\-change directions between two EEG segments and determine which segment is more fatigued\.Cat\.×1\\times 1NA\-39H20Integrate EEG and PSG evidence to determine which of two within\-subject segments shows greater fatigue\.Cat\.×1\\times 1NA\-40E20Classify a single EEG segment as closer to a wakeful or fatigued state\.Cat\.×1\\times 1NA\-41M20Detect long eye closure, slow eye movement, and prolonged blink events in an EOG window\.Cat\.×1\\times 1NA\-42H20Determine whether a wakefulness\-to\-fatigue transition occurs within a long continuous window and localize the stage change\.Cat\.×1\\times 1NA\-43H20Integrate EEG and EOG evidence to determine whether two fatigue indicators agree or conflict\.Cat\.×1\\times 1NA\-44H20Compare ECG\-HRV fatigue indicators between two segments and determine which shows greater physiological fatigue\.Cat\.×1\\times 1NA\-45M20Compare two respiration\-only segments and determine which better matches a fatigue\-related breathing pattern\.Cat\.×1\\times 1NA\-46H20Identify the top feature combination or feature pairs that best characterize fatigue differences across physiological signals\.Set×1\\times 1NA\-47H20Rank three fatigue segments and infer their overall fatigue\-stage structure\.Seq\.×1\\times 1NA\-48H20Compare regional and spectral EEG profiles across three segments and identify the features characterizing fatigue change\.Set×1\\times 1NA\-49H20Analyze fatigue\-related coupling between EEG and EOG using cross\-signal correlation measures\.Num\.×1\\times 1NA\-50E40Answer a standalone EEG and BCI fatigue or vigilance knowledge question\.Cat\.×1\\times 1
### B\.4Physiological Integration

The Physiological Integration subset comprises 39tasksand 1,120instancesconstructed from SEED\-VII, DEAP, Simultaneous Dataset B, SEED\-VIG, and SHHS\-1\. It evaluates multimodal data handling and synchronization, feature extraction, cross\-modal coupling and fusion, quality assessment, signal repair, matching, missing\-modality reconstruction, visualization, and domain knowledge\. Tasks PI\-01–PI\-38 operate on neurophysiological recordings, whereas PI\-39 evaluates multimodal neurophysiology knowledge without recording access shown in Table\.LABEL:tab:pi\_task\_inventory\.

Table 7:Task\-level inventory of the Physiological Integration subset\.Task IDDiff\.\# Inst\.Assessment contentValidation unit\(s\)PI\-01E25Distinguish analyzable physiological channels from event, status, and other non\-signal fields in heterogeneous recordings, and identify channel modalities\.Cat\.×4\\times 4PI\-02E25Parse recording structure and count analyzable channels for specified physiological modalities\.Num\.×2\\times 2PI\-03E25Map a common half\-open time window to the start and end sample indices on the native time axis of two channels\.Num\.×4\\times 4PI\-04E25Read native sampling rates and calculate a signal statistic at a specified time boundary\.Num\.×3\\times 3PI\-05E25Extract the same real\-time segment from modalities with different sampling rates and export a structured multistream result\.File×5\\times 5PI\-06M20Identify sustained ocular contamination using time\-aligned EEG and EOG evidence\.Cat\.×1\\times 1\+ Seq\.×1\\times 1PI\-07M25Resolve the common valid range across streams and annotations and locate the first usable interval satisfying a duration requirement\.Num\.×2\\times 2PI\-08E10Extract a specified single\-window feature from native\-rate EEG\.Num\.×1\\times 1PI\-09E10Extract a specified single\-window feature from ocular or eye\-tracking signals\.Num\.×1\\times 1PI\-10E10Extract a specified single\-window feature from peripheral physiological signals\.Num\.×1\\times 1PI\-11E10Compute the fNIRS optical\-density change using median baseline intensity and summarize the response in the target interval\.Num\.×1\\times 1PI\-12E10Compute a specified eye\-tracking feature while preserving missing samples\.Num\.×1\\times 1PI\-13M30Construct sliding\-window feature series from EEG and an auxiliary modality and estimate their association\.Num\.×1\\times 1PI\-14M30Compute pre/post response changes for EEG and an auxiliary modality in paired event\-centered windows\.Num\.×2\\times 2PI\-15M20Align a specified DEAP trial to its marker, calculate frontal EEG alpha asymmetry and temperature changes, and determine their cross\-modal response relationship\.Num\.×2\\times 2\+ Cat\.×1\\times 1PI\-16H30Assess target\-window EEG drowsiness and EOG slow\-eye evidence relative to within\-record reference windows and determine their consistency\.Num\.×2\\times 2\+ Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-17H30Build a temporal trend of fused EEG–EOG evidence with overlapping windows, estimate its slope, and determine the trend direction\.Num\.×1\\times 1\+ Cat\.×1\\times 1PI\-18E30Read the public sleep\-stage label of a specified epoch and calculate EEG, EOG, and EMG RMS within that window\.Cat\.×1\\times 1\+ Num\.×3\\times 3PI\-19M30Use public sleep labels to locate the longest REM bout and summarize the RMS of two specified physiological modalities within it\.Num\.×4\\times 4PI\-20H30Evaluate event\-locked EEG–fNIRS responses through either cross\-event coupling or rule\-defined single\-event response concordance\.Num\.×1\\times 1\+ Sem\.×1\\times 1; or Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-21M30Construct continuous EEG and HbO trajectories and identify the neurovascular delay with maximal positive correlation\.Num\.×1\\times 1\+ Sem\.×1\\times 1PI\-22E30Independently assess EEG and dual\-wavelength fNIRS quality within a strictly bounded 30\-second interval\.Cat\.×2\\times 2PI\-23M40Select the closest multimodal physiological profile from support\-set candidates under specified feature, normalization, and distance rules\.Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-24E40Assess target EEG and eye\-tracking deviations against within\-record reference distributions, then determine the response\-magnitude state and the modality with the larger deviation\.Cat\.×2\\times 2PI\-25M40Compare standardized EEG and eye\-tracking feature directions with a fixed public template\.Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-26H40Use within\-subject reference trials to derive the principal directions of EEG and eye\-tracking feature change, assess within\-modality agreement, and determine whether cross\-modal evidence conflicts and its source\.Seq\.×1\\times 1\+ Cat\.×2\\times 2PI\-27M30Combine gap\-aware gaze speed, frontal EEG low\-frequency activity, and within\-record reference percentiles to assess eye\-tracking\-linked frontal\-contamination evidence\.Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-28M30Compare EEG activation and autonomic changes from baseline to stimulus and determine cross\-system response concordance\.Num\.×2\\times 2\+ Cat\.×1\\times 1PI\-29M30Fuse EEG–EOG evidence across six consecutive long windows and locate the earliest persistent transition and its direction\.Seq\.×1\\times 1\+ Cat\.×2\\times 2PI\-30M40Compare multimodal response magnitudes between two trials from the same participant and rank them under a specified margin rule\.Seq\.×1\\times 1\+ Cat\.×1\\times 1PI\-31M30Extract an EEG regional\-response and gaze\-exploration profile, identify the dominant EEG region, and generate a spatial visualization\.Cat\.×1\\times 1\+ Sem\.×1\\times 1\+ File×1\\times 1PI\-32H40Estimate an inter\-device time offset in a controlled runtime view, apply non\-circular correction, and deliver the corrected signal and validity mask\.Num\.×1\\times 1\+ File×1\\times 1PI\-33M30Compute specified EEG–fNIRS composite\-response scores for multiple event blocks and rank the blocks\.Seq\.×1\\times 1\+ Sem\.×1\\times 1PI\-34H40Localize the modality, fault type, and interval of a single multimodal signal fault in a runtime\-injected view\.Cat\.×2\\times 2\+ Seq\.×1\\times 1PI\-35H30Compare event\-level response consistency in a common EEG–fNIRS ROI map and identify the most consistent region\.Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-36M40Build five\-modality response signatures for calibration and query trials, compare cosine similarity and sign agreement, and determine the overall pattern\.Num\.×1\\times 1\+ Cat\.×1\\times 1\+ Sem\.×1\\times 1PI\-37H30Construct an EEG–EOG similarity matrix from anonymous runtime windows and perform global one\-to\-one matching\.Seq\.×2\\times 2PI\-38H40Reconstruct a missing\-modality feature from within\-subject calibration data in a controlled restricted view and assess compensation reliability\.Num\.×2\\times 2\+ Cat\.×1\\times 1PI\-39E40Answer a multiple\-choice knowledge question on multimodal neurophysiological recording, signal processing, or limits of interpretation\.Cat\.×1\\times 1
### B\.5Example Instances

Here, we showcase representative instances from each subset\.

![[Uncaptioned image]](https://arxiv.org/html/2608.04156v1/x7.png)![[Uncaptioned image]](https://arxiv.org/html/2608.04156v1/x8.png)![[Uncaptioned image]](https://arxiv.org/html/2608.04156v1/x9.png)

## Appendix CBenchmark Construction Details

In this section, we expand the evaluator\-side construction procedure summarized as follows\. Section[C\.1](https://arxiv.org/html/2608.04156#A3.SS1)describes how deterministic analysis scripts derive instance\-specific ground truth from the bound recordings and parameters\. Section[C\.2](https://arxiv.org/html/2608.04156#A3.SS2)then details how the Parser Agent aligns free\-form reports with structured target fields and how the six validation units score the resulting outputs\.

### C\.1Deterministic Ground\-Truth Generation

Eachinstanceis paired with ground truth after its parenttaskhas been bound to a concrete recording, analysis window, signal selection, and parameter configuration\. A deterministic reference script executes the prescribed workflow on the same input files provided to the target system and produces the numerical values, labels, collections, sequences, or artifacts required by the corresponding validation units\. These reference targets are stored only in the*Validation Configuration*, together with task\-specific tolerances and metric weights, and remain inaccessible to the evaluated model\.

As an example, the first instanceFA\-01\-Instance1is bound toISRUC\_01\.edfand provides the following instruction:

> Please first extract only the EEG channels from the raw signal, and then apply a 0\.5–40 Hz FIR bandpass filter to these channels\. After filtering, obtain each channel’s Alpha\-band power and total filtered\-signal power by integrating the PSD over frequency\. Calculate Alpha relative power separately for each channel, then average the channel\-wise ratios and report the final percentage value clearly in your response\.

The reference script follows the same analysis specification to compute the expected alpha relative power and populate the numerical target and tolerance in the instance\-level*Validation Configuration*\. Listing[1](https://arxiv.org/html/2608.04156#LST1)retains only this core computation; input loading, dataset\-specific channel selection, batch processing, and construction of the complete instance specification are omitted for clarity\.

defcompute\_alpha\_relative\_power\(raw,picks\):

work=raw\.copy\(\)\.pick\(picks\)

work\.filter\(l\_freq=0\.5,h\_freq=40\.0,verbose=False\)

sfreq=float\(work\.info\["sfreq"\]\)

n\_fft=int\(4\.0\*sfreq\)

n\_overlap=min\(n\_fft//2,max\(0,n\_fft\-1\)\)

spectrum=work\.compute\_psd\(

method="welch",

fmin=0\.5,

fmax=40\.0,

n\_fft=n\_fft,

n\_overlap=n\_overlap,

verbose=False,

\)

psds,freqs=spectrum\.get\_data\(return\_freqs=True\)

total\_power=np\.trapz\(psds,freqs,axis=1\)

alpha\_mask=\(freqs\>=8\.0\)&\(freqs<=13\.0\)

alpha\_power=np\.trapz\(

psds\[:,alpha\_mask\],freqs\[alpha\_mask\],axis=1

\)

relative\_alpha=alpha\_power/np\.maximum\(

total\_power,1e\-20

\)

returnfloat\(np\.mean\(relative\_alpha\)\*100\.0\)

Listing 1:Ground\-truth computation for Foundational Analysis task FA\-01\-Instance1\.Applying the same fixed script to every recording associated with FA\-01 changes only the data\-dependent result while preserving the analytical definition of the task\. Other tasks follow the same construction principle, with their reference scripts returning the numerical values, labels, collections, sequences, or artifacts required by the corresponding validation units\.

### C\.2Output Parsing and Multi\-Unit Validation

BrainBench permits free\-form analytical reports because EEG analysis extends beyond scalar prediction to include explanations, supporting evidence, and references to generated artifacts\. Imposing a rigid output schema on each target system would conflate EEG understanding with formatting compliance\. To decouple these factors, the Parser Agent receives the final report together with an instance\-specific extraction prompt and maps only explicitly reported information into a predefined JSON schema\. It has no access to the ground truth and does not assess scientific correctness, correct erroneous answers, or infer omitted results\. Any missing or unresolvable field is returned asnull\.

Listing[2](https://arxiv.org/html/2608.04156#LST2)shows the parser prompt for instanceFA\-30\-Instance1\. The underlying task requests global mean phase\-locking values \(PLVs\) in four frequency bands and a conclusion identifying the dominant synchronization band\. The prompt fixes both the target fields and their admissible types so that differently worded reports can be evaluated through the same interface\.

\#\#\#ROLE

YouareapreciseJSONExtractionEngineforneurophysiologicaldata\.YoursoletaskistoconvertnaturallanguagereportsintostructuredJSONdata\.

\#\#\#TASK

Extractfiveoutputsfromtheagentreport:

1\)deltaglobalmeanPLV

2\)thetaglobalmeanPLV

3\)alphaglobalmeanPLV

4\)betaglobalmeanPLV

5\)dominantsynchronizationband

\#\#\#STRICTCONSTRAINTS\(MANDATORY\)

1\.OutputONLYavalidJSONobject\.

2\.DONOTincludeMarkdowncodeblocks\.

3\.DONOTincludeanyconversationaltextorexplanations\.

4\.KeysmustbeEXACTLY"delta\_global\_mean\_plv","theta\_global\_mean\_plv","alpha\_global\_mean\_plv","beta\_global\_mean\_plv",and"dominant\_synchronization\_band"\.

5\.ThefourPLVvaluesmustbefloatornull\.

6\."dominant\_synchronization\_band"mustbeoneof"delta","theta","alpha","beta",ornull\.

\#\#\#OUTPUTTEMPLATE

\{"delta\_global\_mean\_plv":<float\|null\>,"theta\_global\_mean\_plv":<float\|null\>,"alpha\_global\_mean\_plv":<float\|null\>,"beta\_global\_mean\_plv":<float\|null\>,"dominant\_synchronization\_band":<string\|null\>\}

Listing 2:Parser prompt used forFA\-30\-Instance1\.For this instance, the four extracted PLV fields are passed to numerical validation, while the original report is passed to a task\-specific Semantic Judge that checks whether the selected band and its explanation agree with the ground\-truth PLV ranking\. More generally, each metric selects either a parsed field, the complete report, or a generated artifact and applies one of the six rules described below\. Lety^\\hat\{y\}denote a parsed prediction,yyits reference target, andv∈\[0,1\]v\\in\[0,1\]the resulting validation score\. Let𝒞​\(⋅\)\\mathcal\{C\}\(\\cdot\)denote the canonicalization used by the evaluator, which removes surrounding whitespace, ignores letter case, normalizes numerical representations, and additionally removes internal spaces from set and sequence elements\. The six validation units are implemented as follows:

- •Numerical validation\.The parsed value and ground truth are converted to finite scalars and compared under the instance\-specific absolute toleranceτ\\tau: vnum=𝟙​\[\|y^−y\|≤τ\]\.v\_\{\\mathrm\{num\}\}=\\mathds\{1\}\\\!\\left\[\\,\|\\hat\{y\}\-y\|\\leq\\tau\\,\\right\]\.\(3\)Whenτ=0\\tau=0, exact equality is required\. A missing, non\-numerical, non\-finite, or otherwise invalid value receives zero\.
- •Categorical validation\.Discrete labels, choices, regions, or event types are scored by exact equality after canonicalization: vcat=𝟙​\[𝒞​\(y^\)=𝒞​\(y\)\]\.v\_\{\\mathrm\{cat\}\}=\\mathds\{1\}\\\!\\left\[\\mathcal\{C\}\(\\hat\{y\}\)=\\mathcal\{C\}\(y\)\\right\]\.\(4\)
- •Set validation\.LetS^k\\hat\{S\}\_\{k\}be the canonicalized set formed from the firstkkpredicted elements whentop\_kis specified, and letSSbe the reference set\. Exact matching uses vsetexact=𝟙​\[S^k=S\]\.v\_\{\\mathrm\{set\}\}^\{\\mathrm\{exact\}\}=\\mathds\{1\}\\\!\\left\[\\hat\{S\}\_\{k\}=S\\right\]\.\(5\)For element\-level partial credit, the implemented score is vsetpartial=min⁡\(1,ρ​\|S^k∩S\|\),v\_\{\\mathrm\{set\}\}^\{\\mathrm\{partial\}\}=\\min\\\!\\left\(1,\\rho\\,\|\\hat\{S\}\_\{k\}\\cap S\|\\right\),\(6\)whereρ\\rhois the configured score per matched element and defaults to1/\|S\|1/\|S\|\. Thus, the order of reported elements does not affect the score\.
- •Sequence validation\.The evaluator supports position\-wise, exact\-order, and weighted partial\-order matching\. Position\-wise matching canonicalizes the elements and assigns partial credit at each reference position: vseqpos=1L​∑j=1L𝟙​\[q^j=qj\],v\_\{\\mathrm\{seq\}\}^\{\\mathrm\{pos\}\}=\\frac\{1\}\{L\}\\sum\_\{j=1\}^\{L\}\\mathds\{1\}\\\!\\left\[\\hat\{q\}\_\{j\}=q\_\{j\}\\right\],\(7\)whereLLis the reference length and a missing predicted position is counted as incorrect\. For exact\-order and weighted partial\-order matching, both sequences are additionally deduplicated while retaining the first occurrence and optionally truncated totop\_k\. Exact\-order matching assigns one only when the resulting sequences are identical\. For weighted partial order, letOObe the elements shared by the two sequences,p​\(x\)p\(x\)andp^​\(x\)\\hat\{p\}\(x\)their reference and predicted positions,δ\\deltathe allowed order slip, andwjw\_\{j\}the configured position weights, which default to1/j1/j\. The evaluator computes vseqweighted=𝟙​\[\|O\|≥m\]​\|O\|L​∑x∈Owp​\(x\)​max⁡\(0,1−\|p​\(x\)−p^​\(x\)\|δ\+1\)∑j=1Lwj,v\_\{\\mathrm\{seq\}\}^\{\\mathrm\{weighted\}\}=\\mathds\{1\}\\\!\\left\[\|O\|\\geq m\\right\]\\frac\{\|O\|\}\{L\}\\frac\{\\sum\_\{x\\in O\}w\_\{p\(x\)\}\\max\\\!\\left\(0,1\-\\frac\{\|p\(x\)\-\\hat\{p\}\(x\)\|\}\{\\delta\+1\}\\right\)\}\{\\sum\_\{j=1\}^\{L\}w\_\{j\}\},\(8\)wheremmis the required minimum overlap\.
- •Semantic validation\.The selected parsed field or complete reportRRand a task\-specific judge promptJJare passed to the Semantic Judge: vsem=clip\[0,1\]⁡\(Judge⁡\(R,J\)\)\.v\_\{\\mathrm\{sem\}\}=\\operatorname\{clip\}\_\{\[0,1\]\}\\\!\\left\(\\operatorname\{Judge\}\(R,J\)\\right\)\.\(9\)The judge output may be a normalized score or a Booleanstatus/passeddecision\. The rubric specifies the required conclusion, reference evidence, and conditions under which unsupported or inconsistent claims should fail\.
- •Artifact validation\.A reported path must resolve inside the isolated instance workspace, exist, and satisfy any required filename constraint\. Missing or invalid paths receive zero\. For a signal artifact with field\-level checkszℓ∈\{0,1\}z\_\{\\ell\}\\in\\\{0,1\\\}and configured weightsaℓa\_\{\\ell\}, the score is vart=∑ℓaℓ​zℓ∑ℓaℓ\.v\_\{\\mathrm\{art\}\}=\\frac\{\\sum\_\{\\ell\}a\_\{\\ell\}z\_\{\\ell\}\}\{\\sum\_\{\\ell\}a\_\{\\ell\}\}\.\(10\)The available checks cover array shape, channel count and order, duration, sampling rate, signal RMS, referencing, and passband\. A non\-empty file receives one when no field\-level check is configured\. For image artifacts, the score is either a non\-empty\-file check or the normalized output of a VLM Judge under an instance\-specific visual rubric\.

The validation units remain independent and may be combined within an instance\. Each unit returns a normalized score before the task\-specific weights in the*Validation Configuration*are applied, as described in Appendix[3\.2\.3](https://arxiv.org/html/2608.04156#S3.SS2.SSS3)\. This separation allows the benchmark to evaluate numerical accuracy, discrete decisions, structured outputs, scientific interpretation, and deliverable artifacts without reducing heterogeneous EEG workflows to a single output format\.

## Appendix DExperimental Protocol and Implementation Details

We provide the implementation details of the BrainBench evaluation protocol\. Appendix[D\.1](https://arxiv.org/html/2608.04156#A4.SS1)presents the end\-to\-end black\-box workflow, from instance dispatch and isolated execution to output validation and scoring\. Appendix[D\.2](https://arxiv.org/html/2608.04156#A4.SS2)and[D\.3](https://arxiv.org/html/2608.04156#A4.SS3)describe the tool\-mediated BrainAgent workflow and the autonomous code\-execution protocol of CodeAct, respectively\. Appendix[D\.4](https://arxiv.org/html/2608.04156#A4.SS4)specifies the containerized runtime shared by both paradigms, including environment isolation, resource constraints, infrastructure configuration, and the policy for rerunning infrastructure\-induced failures\.

### D\.1Details of the Black\-Box Evaluation Protocol

![Refer to caption](https://arxiv.org/html/2608.04156v1/x10.png)Figure 6:End\-to\-end black\-box evaluation protocol of BrainBench\.Each evaluation instance exposes only its instruction and input files to the target system, which executes the task through either CodeAct or BrainAgent in an isolated container\. The resulting free\-form report and requested artifacts are returned to the evaluator\. A Parser Agent converts explicitly reported information into an instance\-specific JSON schema, after which the parsed fields, original report, and generated artifacts are assessed by the applicable validation units and aggregated into the final instance score\. Green components are visible to the target system, purple components remain evaluator\-only, and gray components denote audit information that is recorded but not used for scoring\.Figure[6](https://arxiv.org/html/2608.04156#A4.F6)illustrates the end\-to\-end black\-box evaluation protocol and distinguishes information visible to the target system from evaluator\-only references and audit records\. For eachinstance, the target system receives only the natural\-language instruction and the associated input files\. The ground truth, Parser Prompt, Validation Configuration, and scoring implementation remain hidden on the evaluator side throughout execution\. The same instance is dispatched to either BrainAgent or CodeAct through a unified input–output interface and executed in an isolated, instance\-specific container\. Within this environment, the target system may inspect the supplied files, perform intermediate analyses, and generate the requested artifacts\. Only the final natural\-language report and retained artifacts are passed to the scoring pipeline, ensuring that both execution paradigms are evaluated under identical instructions, data, and output requirements\. After execution, the final report and the instance\-specific Parser Prompt are provided to the Parser Agent, which extracts explicitly reported information into a predefined JSON schema\. The Parser Agent performs output alignment only: it has no access to the ground truth and does not assess scientific correctness, revise erroneous answers, or infer missing results\. The extracted fields are then routed to the applicable numerical, categorical, set, and sequence validation units, while the original report is retained for semantic validation and the generated files are examined through artifact validation\. Each validation unit applies the references, tolerances, matching rules, and weights specified in the instance\-level Validation Configuration\. The resulting unit scores are combined to produce the final instance score\.

Execution information is recorded separately for auditing and error analysis\. These records include token usage, runtime, generated code, tool and execution traces, intermediate inputs and outputs, and error messages, none of which directly contributes to the benchmark score\. Failures attributable to the target system or execution paradigm—such as invalid code, incorrect tool use, incomplete reports, or malformed artifacts—are retained as evaluation outcomes\. An instance is rerun only when execution is interrupted by an independently verified infrastructure failure, such as an external API communication error or container initialization failure\. The container constraints and failure\-handling policy are detailed in Appendix[D\.4](https://arxiv.org/html/2608.04156#A4.SS4)\.

### D\.2BrainAgent Architecture and Execution Details

#### D\.2\.1Architecture and Benchmark Adaptation

BrainAgent is a multi\-agent framework for brain\-signal analysis in which a Supervisor Agent interprets each request, decomposes it into analytical objectives, constructs an execution queue, and delegates these objectives to specialized subagents\[[68](https://arxiv.org/html/2608.04156#bib.bib15)\]\. Each subagent encapsulates domain\-specific reasoning and tool use, while the supervisor maintains a unified user\-facing interface and synthesizes the subagent outputs into a final report\. Subagents and their associated tools can be registered dynamically, allowing new analytical domains to be incorporated without modifying the supervisory control layer\. This hierarchical architecture provides a natural basis for BrainBench, whose four subsets encompass heterogeneous analytical capabilities while sharing a common target\-system interface\. We preserve the original supervisor–subagent organization and adapt BrainAgent in three respects: the tool layer, the planning mechanism, and the representation of intermediate state\.

First, we broaden the analytical tool layer to cover the capability space evaluated by the benchmark\.The resulting toolset is deliberately capability\-oriented rather than task\-specific:each tool encapsulates a reusable analytical operation, such as data inspection, signal preparation, temporal segmentation, feature extraction, statistical aggregation, or artifact generation\. No tool encodes the solution, task\-specific parameters, reference answer, or decision rule of any particulartaskorinstance\. Successful execution therefore still requires the evaluated LLM to interpret the instruction, select and configure appropriate tools, integrate intermediate evidence, and formulate a scientifically grounded conclusion\.

Second, to better accommodate complex and long\-horizon workflows,we extend BrainAgent beyond its original single\-pass planning scheme, in which a subagent planned the complete workflow once before executing it sequentially,to an iterative, state\-aware planning mechanism\.The Supervisor Agent first dispatches the complete analytical objective to an appropriate subagent\. At each planning round, the subagent receives a compact representation of the current state and selects the maximal executable segment of its remaining plan\. The tools within that segment are invoked sequentially, and their outputs are registered under stable identifiers before the next planning round begins\. This segmented execution strategy allows subsequent decisions to incorporate observations obtained during execution—such as discovered channel names, computed features, or newly generated labels—without requiring the full workflow to be fixed in advance\. Once the accumulated evidence is judged sufficient to satisfy the instruction, the subagent produces an analytical report, which is returned to the Supervisor Agent for final response synthesis\.

Third, to support analyses spanning multiple files, subjects, and signal views,we replace BrainAgent’s original flat intermediate\-variable representation with a hierarchicalshared\_state\. The state is accessible to the Supervisor Agent and active subagents and organizes intermediate objects into separate branches for rawrecordings, derived signalviews, temporallabels, detectedevents, structured analysisresults, and generatedartifacts\. A dedicatedplanningbranch records executed\-tool counts and recent tool signatures, while internal counters assign collision\-free identifiers to newly created objects\. Tools exchange references to these registered objects rather than repeatedly serializing complete signals\. Before each planning round, BrainAgent projects the hierarchy into a compact, identifier\-based summary, enabling the LLM to retrieve prior outputs, coordinate dependencies across data sources, and trace the evidence underlying its conclusions\.

#### D\.2\.2Capability\-Oriented Toolset

BrainAgent comprises three domain subagents—SleepAgent,NeurocogAgent, andPhysiolAgent—for Sleep Assessment, Neurocognitive Assessment, and Physiological Integration, respectively\. Each combines a subset\-specific extension with a shared common toolset implementing the reusable EEG operations covered by Foundational Analysis\. TableLABEL:tab:brainagent\_toolslists the complete toolsets and detailed interfaces are provided with the released implementation\.

Table 8:Complete capability\-oriented toolset used by BrainAgent\.Foundational Analysis ToolsPrimary capabilityRecordingLoaderLoads an EEG or PSG recording and registers its native metadata and channel names\.ChannelInspectorClassifies channels by signal modality and EEG region while preserving exact source names\.SignalViewBuilderConstructs a named signal view with explicit channels, time window, filtering, referencing, resampling, and epoch settings\.DataExporterExports recordings, views, or arrays as EDF, FIF, NPY, or CSV artifacts\.FigureExporterProduces channel\-aware PSD, heatmap, hypnogram, spectrogram, and raw\-trace figures\.SpectralFeatureAnalyzerComputes band power, relative power, band ratios, dominant bands, alpha peak frequency, and spectral edge frequency\.TimeDomainFeatureAnalyzerComputes variance, skewness, kurtosis, RMS, Hjorth parameters, and global field power\.ComplexityFeatureAnalyzerComputes nonlinear EEG complexity measures, including sample entropy\.ConnectivityAnalyzerComputes inter\-channel connectivity using correlation, phase\-locking value, or weighted phase\-lag index\.ArtifactQualityAnalyzerAssesses missing values, flatlines, excessive amplitude, drift, and channel\-level signal quality\.FeatureRankerAggregatorRanks or aggregates channels, regions, bands, windows, and prior structured results\.Sleep Assessment ToolsPrimary capabilitySleepLabelLoaderLoads sleep\-stage annotations and normalizes them to W, N1, N2, N3, and R with epoch timing\.SleepArchitectureAnalyzerComputes sleep onset latency, total sleep time, time in bed, sleep efficiency, WASO, REM latency, and stage proportions\.SleepBoutTransitionAnalyzerAnalyzes stage bouts, transitions, awakenings, longest episodes, and fragmentation\.SleepStageEstimatorEstimates sleep stages using an optional learned backend and a transparent rule\-based fallback\.SleepSpectralSegmentAnalyzerQuantifies sleep\-related band power and sigma activity within an explicitly defined EEG segment\.EMGActivityAnalyzerQuantifies chin or other EMG activity, stage\-dependent tone, and a REM\-atonia proxy\.EOGActivityAnalyzerQuantifies slow\- and rapid\-eye\-movement activity from explicitly selected EOG channels and windows\.ECGHeartRateAnalyzerEstimates heart\-rate statistics from explicitly selected ECG channels\.SleepArousalAnalyzerDetects rule\-based EEG arousal\-like events using robust high\-frequency envelope thresholds\.RespiratoryEventAnalyzerDetects apnea, hypopnea, apnea subtype, and RERA\-like events from respiratory, oximetry, and arousal evidence\.SleepEventIndexAnalyzerComputes event counts and rates by total sleep time and sleep stage\.SleepEventTemporalAnalyzerFilters, ranks, and temporally aggregates sleep events and associates them with stage labels\.OximetryAnalyzerComputes ODI3, ODI4, sleep/wake mean and minimum SpO2, T90, and T80\.SleepMicroEventAnalyzerDetects sleep spindles, K\-complexes, or slow waves for an explicitly specified event type\.Neurocognitive Assessment ToolsPrimary capabilityEEGFeatureStreamBuilderBuilds aligned EEG feature streams across windows, segments, or trials for downstream correlation, ranking, and heatmap tasks\.FeatureStreamOperatorApplies mechanical numeric operations such as extraction, difference, ratio, and aligned comparison to prior structured feature results\.EOGWindowStreamBuilderBuilds aligned window\-level fatigue EOG feature streams, including eye\-closure proxy, SEM log power, and blink\-rate sequences\.AsymmetryAnalyzerComputes band\-power asymmetry across homologous left\-right EEG channel pairs within a selected time window, with optional baseline comparison and aggregate laterality output\.TrendAnalyzerExtracts a numeric series from time\-ordered prior results and summarizes its linear, early\-versus\-late, and monotonic trends\.GroupContrastAnalyzerAggregates prior channel\-level features into region\-band summaries and computes group means, mean differences, and Cohen’s d contrasts\.ContinuousLabelLoaderLoads continuous affective labels from files, numeric lists, or structured items while preserving alignment metadata for downstream analysis\.PolarityClassifierReturns EMOD binary negative\-versus\-positive probabilities together with an internal rule\-based vote\.ValenceRBTransformerConverts EEG into DEAP\-style differential\-entropy tokens and performs binary valence inference with a local RBTransformer checkpoint\.ArousalRBTransformerUses the same DEAP\-style differential\-entropy preprocessing and RBTransformer pipeline to perform binary arousal classification\.P300AnalyzerExtracts target/non\-target ERP averages, difference waves, and P300 peak amplitude and latency from stimulus\-locked EEG epochs\.BehaviorLabelAnalyzerComputes objective behavioral metrics such as accuracy, hit rate, miss rate, false\-alarm rate, and reaction time from structured behavior labels\.WorkloadRuleClassifierApplies interpretable spectral\-rule voting based on frontal theta, parietal alpha, and engagement\-channel beta evidence to classify explicitly provided workload segments\.FeatureLabelCorrelatorAligns feature streams with numeric labels, computes Pearson or Spearman correlations, and returns ranked feature\-label associations\.CorrelationMatrixAnalyzerComputes a row\-by\-column correlation matrix across two aligned sets of numeric streams and returns ranked pairs with heatmap\-ready outputs\.EOGEventAnalyzerDetects blink\-like, prolonged\-blink, and long\-eye\-closure events from explicit EOG channels and quantifies slow\-eye\-movement power\.ECGHRVAnalyzerDetects R peaks, cleans RR intervals, and computes time\-domain and frequency\-domain HRV metrics from explicit ECG channels\.RespirationRhythmAnalyzerReports breath count, respiration rate, interval and amplitude variability, and respiration\-band power from explicit respiration channels\.Physiological Integration ToolsPrimary capabilityMultirateRecordingLoaderLoads multimodal recordings at native sampling rates and registers streams, channel metadata, timing, units, and bundles\.ChannelInspectorInspects channel names, modalities, units, sampling rates, and native sample indices\.SignalViewBuilderBuilds reusable channel/time views with filtering, rereferencing, scaling, and explicit resampling\.DataExporterExports streams, arrays, or structured results to NPY, NPZ, CSV, JSON, or EDF\.NumericAnnotationLoaderLoads numeric NPY/NPZ labels, event tables, masks, and reference arrays with safe decoding and bounded summaries\.AnnotationIntervalAnalyzerAligns labels, events, signal coverage, and validity masks; supports time lookup and label\-bout formation\.EventMarkerAnalyzerDecodes marker channels and pairs chronological start/end markers into half\-open event intervals\.SignalQualityAnalyzerAssesses finite coverage, missing runs, flatness, amplitude, and signal drift for physiological channels\.TimeDomainFeatureAnalyzerComputes mean, median, standard deviation, RMS, mean absolute value, line length, slope, and Hjorth features\.SpectralFeatureAnalyzerComputes one\-interval Welch band power and relative power for explicitly declared frequency bands\.EEGPatternAnalyzerComputes EEG band evidence, log band ratios, hemispheric asymmetry, ROI summaries, and alpha suppression\.BatchFeatureAnalyzerComputes aligned features over repeated windows or events and optionally associates two finite feature trajectories\.EDAActivityAnalyzerSeparates tonic and phasic electrodermal activity and summarizes SCL and SCR responses\.EOGActivityAnalyzerComputes EOG RMS, line length, slow\-eye\-movement power, and blink summaries\.EMGActivityAnalyzerComputes EMG RMS, mean absolute value, envelope, and waveform length\.EyeTrackingAnalyzerComputes validity, pupil summaries, gaze dispersion, path length, speed, and screen\-occupancy measures\.CardiacActivityAnalyzerEstimates ECG/BVP rates, beat intervals, RMSSD, and pulse\-amplitude summaries\.RespiratoryActivityAnalyzerDetects respiration peaks and summarizes breathing rate, intervals, and waveform amplitudes\.OximetryAnalyzerValidates SpO2and computes means, minima, threshold time, and desaturation summaries\.FNIRSActivityAnalyzerConverts dual\-wavelength intensity to optical density, HbO, and HbR using explicit modified Beer–Lambert parameters\.ReferenceDistributionAnalyzerComputes reference normalization, target z\-scores, percentiles, thresholds, and nearest\-centroid support matching\.EventResponseAnalyzerComputes event\-locked scalar features and baseline/response contrasts for synchronized physiological events\.TrialGroupAnalyzerAggregates trial\-level features into named groups, early/late phases, or nearest\-group comparisons\.TrajectoryAssociationAnalyzerComputes correlation, lagged ranking, weighted fusion, and trend summaries for two scalar trajectories\.StatisticalModelAnalyzerRuns two\-group tests, Pearson/Spearman correlation, and OLS models with optional Holm correction\.
#### D\.2\.3Illustrative BrainAgent Execution Traces

As an illustrative BrainAgent execution trace, we present the evaluation of Sleep Assessment instanceSA\-26\-Instance1using Claude Opus 5\. The corresponding instruction is:

> Load the EEG data file at pathdata/sleep/ISRUC\_01\.edf\. Then analyze the following three 1\-minute segments: A = minute 274\.5 to 275\.5, B = minute 63\.0 to 64\.0, C = minute 437\.0 to 438\.0\. For each segment, quantify the EOG/eye\-movement activity strength using the EOG/eye\-movement channels present in the current system state \(inspect the loaded recording to determine the exact channel names; do not assume names\)\. Compare the three segments and report the segment labels \(A, B, C\) in descending order of EOG activity strength, including the quantitative metric values used for the ranking\. Additionally, determine the sleep stage of the segment with the strongest EOG activity and explicitly state whether that segment is REM sleep, with supporting evidence\.

Listing[3](https://arxiv.org/html/2608.04156#LST3)presents a faithful condensed trace\. It retains the supervisor dispatch, all six planning rounds, all 16 tool calls, and the final analytical output, while omitting repetitive success messages and serialized state snapshots\. The colored statements summarize the reasoning that determined each planning segment; color distinguishes rounds only\.

\[Supervisordispatch\]

Targetsubagent:SleepAgent

Task:executeSA\-26\-Instance1usingdata/sleep/ISRUC\_01\.edf

\[Planninground1\]

No recording is registered\. Load the file and inspect its channels before selecting EOG inputs\.

01RecordingLoader

02ChannelInspector

\[Planninground2\]

The exact EOG channels are LOC\-A2 and ROC\-A1\. Build the three requested views and quantify EOG activity consistently\.

03SignalViewBuilder\[A:16470\-\-16530s\]

04SignalViewBuilder\[B:3780\-\-3840s\]

05SignalViewBuilder\[C:26220\-\-26280s\]

06EOGActivityAnalyzer\[A\]

07EOGActivityAnalyzer\[B\]

08EOGActivityAnalyzer\[C\]

\[Planninground3\]

All EOG metrics rank A above B above C\. Determine A’s stage, collect spectral and EOG evidence, and formalize the ranking\.

09SleepStageEstimator\[A,1\-minview\]

10SleepSpectralSegmentAnalyzer\[A\]

11EOGActivityAnalyzer\[A,0\.3\-\-10Hzcheck\]

12FeatureRankerAggregator\[A,B,C\]

\[Planninground4\]

A is staged as W/N1 and is delta\-dominant\. Add chin\-EMG evidence to test for the muscle atonia expected in REM\.

13SignalViewBuilder\[A,X1channel\]

14EMGActivityAnalyzer\[A\]

\[Planninground5\]

The 1\-min staging context is limited\. Re\-estimate A within a 10\-min EEG/EOG/EMG context before drawing the sleep\-stage conclusion\.

15SignalViewBuilder\[16200\-\-16800scontext\]

16SleepStageEstimator\[10\-mincontext\]

\[Planninground6\]

The state now contains convergent EOG, staging, spectral, and EMG evidence\. No further tool call is needed; generate the report\.

\[Finaloutput\]

EOGactivityranking\(descending\):A\>B\>C

SegmentWindow\(s\)EOGRMS\(V\)LinelengthDerivativeRMS

A16470\-\-165304\.44e\-062\.16e\-073\.72e\-07

B3780\-\-38401\.81e\-061\.47e\-071\.97e\-07

C26220\-\-262801\.10e\-069\.22e\-081\.25e\-07

ThestrongestsegmentisA,anditisnotREMsleep\.Itstwo30\-sepochswereinitially

estimatedasWandN1\.Inthe10\-mincontext,theepochscoveringAwereN2andW,whereas

theREMepochsoccurredoutsideA\.TheEEGwasdelta\-dominant\(66\.9%relativedeltapower\),

andtheX1EMGretainedmeasurabletone\(RMS4\.68e\-06V\),providingnoevidenceofREMatonia\.

Conclusion:A\>B\>C;segmentAisawake/light\-NREMtransition,notREMsleep\.

Listing 3:Condensed BrainAgent execution trace for SA\-26\-Instance1 with Claude Opus 5\. Planning summaries are color\-coded by round\.A second trace illustrates the evaluation of Foundational Analysis instanceFA\-33\-Instance15on a SEED\-V recording\. The corresponding instruction is:

> Load the EEG data file at pathdata/core/SEED\-V\-05\.cnt\. Select the standard EEG channels and compute, for each channel, alpha\-band relative power and signal variance\. Exclude channels with variance belowQ1−1\.5×IQRQ\_\{1\}\-1\.5\\times\\mathrm\{IQR\}or aboveQ3\+1\.5×IQRQ\_\{3\}\+1\.5\\times\\mathrm\{IQR\}\. Among the remaining channels, return those whose alpha relative power is greater than the global mean across the retained channels\. Use Welch PSD estimation, with relative power defined as the target\-band power divided by the total power from 0\.5 to 45 Hz\.

Listing[4](https://arxiv.org/html/2608.04156#LST4)retains all six planning rounds, all 12 successful tool calls, and the final analytical output\. Serialized state snapshots, repetitive success messages, and the provider\-side retry preceding the final planning round are omitted because they do not alter the analytical workflow\. Following the benchmark’s agent organization, the trace names the executing subagentNeurocogAgent\.

\[Supervisordispatch\]

Targetsubagent:NeurocogAgent

Task:executeFA\-33\-Instance15usingdata/core/SEED\-V\-05\.cnt

\[Planninground1\]

No recording is registered\. Load the CNT file and inspect its channels before selecting the EEG inputs\.

01RecordingLoader

02ChannelInspector

\[Planninground2\]

The recording contains 64 EEG channels after excluding VEO and HEO\. Compute channel\-wise alpha relative power and variance over the same channel set\.

03SpectralFeatureAnalyzer\[64EEGchannels;alpha/\[0\.5\-\-45\]Hz\]

04TimeDomainFeatureAnalyzer\[64EEGchannels;variance\]

\[Planninground3\]

Per\-channel previews are truncated\. Apply IQR\-based variance screening and expose the complete variance and alpha\-power rankings\.

05ArtifactQualityAnalyzer\[varianceoutliers;1\.5xIQR\]

06FeatureRankerAggregator\[variance;ascending\]

07FeatureRankerAggregator\[alpharelativepower;descending\]

\[Planninground4\]

The upper variance tail is still hidden by the state summary\. Retrieve a descending variance ranking to identify the high\-side outliers\.

08FeatureRankerAggregator\[variance;descendingtop20\]

\[Planninground5\]

M1 and M2 are the two high\-variance outliers\. Recompute the retained\-channel mean and expose the remaining per\-channel alpha values in manageable groups\.

09SpectralFeatureAnalyzer\[62retainedchannels;globalmean\]

10SpectralFeatureAnalyzer\[retainedchannelgroupA\]

11SpectralFeatureAnalyzer\[retainedchannelgroupB\]

12SpectralFeatureAnalyzer\[retainedchannelgroupC\]

\[Planninground6\]

The retained\-channel mean is 8\.9181%\. Compare every retained channel with this threshold and report the qualifying set\.

\[Finaloutput\]

Recording:66channelsat1000Hz;64channelsidentifiedasEEG\.

Variance\-IQRexclusions:M2\(7\.782e\-06\),M1\(3\.405e\-07\)\.

RetainedEEGchannels:62\.

Globalmeanalpharelativepoweroverretainedchannels:8\.9181%\.

Qualifiedchannels\(alpharelativepower\>8\.9181%;32reported\):

C3,C1,C2,C4,TP7,CP5,CP3,CP1,CP2,CP4,CP6,TP8,

P7,P5,P3,P1,PZ,P2,P4,P6,P8,PO7,PO5,PO3,POZ,

PO4,PO6,PO8,CB1,O1,OZ,O2\.

Listing 4:Condensed BrainAgent execution trace for FA\-33\-Instance15 with Claude Opus 5\. Planning summaries are color\-coded by round\.

### D\.3CodeAct Execution Protocol

#### D\.3\.1Interaction Loop

CodeAct provides the autonomous code\-execution counterpart to the structured BrainAgent workflow\. It receives the same instance\-level instruction and permitted input files through the unified interface described in Section[D\.1](https://arxiv.org/html/2608.04156#A4.SS1), but it is not given the tool inventory, structured intermediate state, or task\-specific helper functions\. Instead, the target LLM independently selects the analysis method, Python libraries, preprocessing operations, intermediate computations, and artifact\-generation procedure needed to complete the instruction\.

For eachinstance, CodeAct initializes an instance\-scoped persistent Python kernel\. At every interaction round, the model returns one of two actions: an<execute\>block containing Python code or a<solution\>block containing the final report\. Code inside an<execute\>block is evaluated in the persistent kernel, allowing loaded recordings, intermediate variables, and generated files to be reused across subsequent rounds of the same instance\. Textual stream output,text/plainexecution results, and exception tracebacks are collected as an*Observation*and returned to the model\. The model may then inspect the result, correct erroneous code, revise its analytical strategy, or continue the analysis\. This loop terminates when the model emits a valid<solution\>block or when an execution\-control limit is reached\. Only the content of<solution\>is treated as the target system’s final response\. Intermediate code, printed values, and error messages are retained as execution information but are not interpreted as final answers and do not directly contribute to the benchmark score\. The final report and any reported artifacts are subsequently processed by the same external Parser Agent and validation pipeline used for BrainAgent\.

#### D\.3\.2Prompt and Execution Contract

Listing[5](https://arxiv.org/html/2608.04156#LST5)presents the complete system prompt used by CodeAct\. The prompt defines the executable action format, the persistent Python environment, the observation feedback mechanism, and the distinction between intermediate execution and the final response\.

Youareahelpfulassistantassignedaproblem\-solvingtask\.Youhaveaccessto

aninteractivePythonenvironmenttoinspectdataandcalculatetheanswer\.

Returnexactlyoneactionblockperturn:

<execute\>\.\.\.</execute\>or<solution\>\.\.\.</solution\>\.

Donotoutputplans,explanations,ortextoutsidetheselectedblock\.

Thenchooseexactlyoneoftheseactions:

1\)ExecutePythoncodebyenclosingitin<execute\>\.\.\.</execute\>\.Thecodewill

runinapersistentPythonkernelandtheoutputwillbereturnedasan

Observation\.Top\-levelvariablesfromearliersnippetsremainavailable\.

2\)Whenthetaskiscomplete,providetherequestedfinalreportenclosedin

<solution\>\.\.\.</solution\>\.Thetextinside<solution\>isreturnedtotheuser,

soitmustfollowthetask’srequestedoutputformatandcontainallresults\.

Usecodetoinspecttheprovidedfilesinsteadofguessing\.Pathsnamedinthe

taskareaccessiblefromthecurrentworkspace\.Toconservecontext,neverprint

anentirelongsignal,labelsequence,dataframe,orfile\.

\-\-\-

Exampletask:

Thefileinput/HR\.npycontainsa1Hzheart\-ratesignalinBPM\.Calculatehow

manysecondsareintheinclusiverange60to100BPMandreturnJSONwiththe

keytime\_in\_range\.

Assistant:

<thought\>Iwillloadthesignalandinspectitsshape\.</thought\>

<execute\>

importnumpyasnp

hr=np\.load\("input/HR\.npy"\)

print\(hr\.shape\)

</execute\>

Observation:

\(300,\)

Assistant:

<thought\>Iwillcountsamplesinrange;at1Hzthecountequalsseconds\.</thought\>

<execute\>

time\_in\_range=float\(np\.sum\(\(hr\>=60\)&\(hr<=100\)\)\)

print\(time\_in\_range\)

</execute\>

Observation:

240\.0

Assistant:

<thought\>Thecalculationiscomplete,soIwillreturntherequestedJSON\.</thought\>

<solution\>

\{"time\_in\_range":240\.0\}

</solution\>

\-\-\-

Theactualtaskfollowsintheusermessage\.

Listing 5:System prompt used for the CodeAct execution protocol\.On the final permitted interaction round, the controller retains the preceding system prompt, instance query, execution actions, and observations, and appends an additional user message to force finalization\. The appended message is shown verbatim in Listing[6](https://arxiv.org/html/2608.04156#LST6)\.

Theexecutionbudgetisexhausted\.Youmustreturnthebestavailablefinal

answernowinside<solution\>\.\.\.</solution\>\.Donotexecutemorecode\.

Listing 6:Verbatim finalization prompt appended on the last CodeAct interaction round\.The instance\-specific user message is constructed separately from the system prompt\. Before dispatch, the evaluator projects the permitteddata\_pathand optionallabel\_pathinto the isolated execution environment and rewrites them as paths under/input/\. The user message then identifies these paths, presents the natural\-languageinstruction, and appends any additional fields explicitly included inagent\_input\. The instruction is augmented with a runtime rule stating that input files under/input/are read\-only and that requested artifacts must be saved under/workspace/file\_check/or a relative path underfile\_check/\.

The benchmark information boundary is enforced by the evaluator rather than encoded as additional task text in the system prompt\. CodeAct receives only the preparedagent\_inputand the permitted input\-file mounts\. The ground truth, Parser Prompt, Validation Configuration, metric definitions, metric weights, and evaluator\-side scores are not included in the model messages or execution workspace\. The model also has no interactive mechanism for requesting additional files during execution\. After CodeAct returns its final<solution\>, the external Parser Agent extracts the required fields, and the evaluator applies the hidden validation units and scoring configuration\.

Although the worked example includes<thought\>blocks, the executable protocol recognizes only<execute\>and<solution\>as actions\. Text outside these two blocks is neither executed nor treated as the final answer\. Accordingly, intermediate Python output and observations are used only to support subsequent interaction rounds, whereas the content enclosed by<solution\>constitutes the final response submitted to the evaluation pipeline\.

#### D\.3\.3Execution Controls

Table[9](https://arxiv.org/html/2608.04156#A4.T9)reports the CodeAct\-specific interaction controls\. Execution output is returned to the model as a merged textual observation\. This observation may contain standard\-stream text, expression results, textual display representations, or a Python traceback\. When an execution succeeds without textual output, the runtime returns an explicit success message; when it exceeds the per\-execution limit, a timeout observation is returned\. These observations provide the model with an opportunity to diagnose and revise failed analyses, while the repeated\-failure controls prevent unproductive execution loops\. On the final permitted round, CodeAct is instructed to stop executing code and return the best available result inside<solution\>\. If the hard instance deadline or an unrecoverable infrastructure error occurs earlier, the instance terminates with a structured error rather than receiving evaluator feedback or a reference answer\.

Table 9:CodeAct interaction and execution controls\.The fixed Python environment contains general numerical and scientific\-computing packages, including NumPy, SciPy, pandas, Matplotlib, seaborn, and scikit\-learn, together with EEG\- and physiological\-signal packages such as MNE, YASA, pyEDFlib, edfio, and WFDB\. CodeAct may compose these libraries freely but cannot assume access to uninstalled task\-specific software\. Variables and temporary files persist across interaction rounds only within the current instance; different instances receive independent kernels, message histories, and workspaces\. Requested artifacts remain available for evaluator\-side validation before cleanup, while the retained audit record stores action types, token usage, elapsed time, code hashes, observation summaries, and termination status\. CPU, memory, image, filesystem, and network isolation are described separately in Appendix[D\.4](https://arxiv.org/html/2608.04156#A4.SS4)\.

### D\.4Containerized Runtime and Failure Handling

This section specifies the containerized execution environment used by BrainAgent and CodeAct and defines how infrastructure\-induced interruptions are distinguished from target\-system failures\. Both paradigms operate through the same instance\-level input–output interface, while their container configurations differ where required by their execution mechanisms\.

#### D\.4\.1Container Configuration

Eachinstanceis executed in a fresh Docker container with an independent runtime state\. Both paradigms use a versioned Python 3\.9 scientific\-computing environment and receive the same instruction and permitted input files\. The containerization layer does not alter the expected report or artifact requirements of the instance\. Table[10](https://arxiv.org/html/2608.04156#A4.T10)summarizes the principal runtime settings\.

Table 10:Principal container configurations used for BrainAgent and CodeAct\.The larger memory allocation for BrainAgent accommodates its multi\-agent runtime, hierarchical intermediate state, and domain\-oriented analytical toolset\. CodeAct instead executes generated code in a smaller scientific\-computing sandbox\. Its execution container has no network access and receives no target\-model credentials, whereas BrainAgent requires network access because model requests and tool\-mediated planning are performed inside the container\. Within each execution paradigm, the corresponding resource and time constraints are held fixed across evaluated models\.

#### D\.4\.2Failure Classification and Retry Policy

We distinguish verified infrastructure interruptions from failures attributable to the target system\. Transient API or container\-runtime failures are retried under a predefined policy; if recovery is unsuccessful, the affected instance is rerun after the external condition is restored using the same instruction, inputs, model configuration, execution paradigm, and container constraints\. The interrupted execution is excluded from score aggregation and replaced by the valid rerun\. In contrast, failures arising from model planning, code generation, tool use, execution strategy, or incomplete outputs are retained as evaluation outcomes and are not rerun by the evaluator\. Any self\-correction performed by BrainAgent or CodeAct within their allotted execution budgets is considered part of the evaluated paradigm; once that budget is exhausted, the resulting report, missing output, or termination state is scored as produced\. Table[11](https://arxiv.org/html/2608.04156#A4.T11)summarizes the resulting policy described above\.

Table 11:Failure classification and instance\-level rerun policy\.

## Appendix EAdditional Experimental Results and Analyses

### E\.1Overall Score Distributions

Figure[7](https://arxiv.org/html/2608.04156#A5.F7)shows that aggregate model scores arise from highly heterogeneous instance\-level outcomes\. In both subsets, the distributions remain broad and contain substantial mass at both low and perfect scores, indicating that even the strongest models do not achieve uniformly reliable performance across tasks\. BrainAgent raises the performance floor and compresses the differences among models: the unweighted mean instance scores span \(63\.1\)–\(83\.4\) in Foundational Analysis and \(64\.2\)–\(77\.9\) in Sleep Assessment, compared with \(40\.9\)–\(86\.0\) and \(46\.0\)–\(77\.5\), respectively, under CodeAct\. This reduced dispersion suggests that structured workflows make performance less sensitive to the underlying model, particularly by supporting weaker models\. CodeAct exhibits greater model dependence, yet its strongest configurations reach the same performance range as the leading BrainAgent systems, demonstrating that autonomous coding retains a high ceiling when the model can plan and execute the analysis reliably\. Overall, the distributions distinguish the robustness advantage of structured agentic execution from the higher\-variance potential of autonomous coding\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x11.png)Figure 7:Instance\-level score distributions across models and execution paradigms\.Violin densities and translucent points show the unweighted normalized instance scores for Foundational Analysis \(top\) and Sleep Assessment \(bottom\) under BrainAgent \(left\) and CodeAct \(right\)\. Black horizontal markers denote arithmetic means\. Models are sorted independently within each panel by their mean instance score; task\-difficulty weights are not applied\.
### E\.2Difficulty\-Conditioned Effects of Execution Paradigms

Figure[8](https://arxiv.org/html/2608.04156#A5.F8)complements the aggregate difficulty results by jointly examining performance and cross\-instance stability for each paired model–task combination\. To make dispersion comparable across tasks with different mean scores and numbers of instances, within\-task score variability is normalized by its finite\-sample maximum under the bounded0–100100scoring range\. The difficulty\-level means lie predominantly to the right of zero, confirming that BrainAgent generally improves task scores, while the broad scatter across all four quadrants shows that this effect is not universal\. In Foundational Analysis, the mean performance advantage is strongest on Easy tasks and becomes smaller for Medium and Hard tasks\. In Sleep Assessment, the largest separation occurs on Medium tasks, whereas the Hard\-task mean remains close to the origin\. Easy and Medium means are also at or below zero on the variability axis, indicating that their score gains are achieved without a systematic loss of stability\. By contrast, the limited separation on Hard tasks suggests that structured workflows cannot fully compensate for the reasoning and evidence\-integration demands of the most complex analyses\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x12.png)Figure 8:Difficulty\-conditioned performance–stability trade\-off between BrainAgent \(BA\) and CodeAct \(CA\)\.Each translucent marker represents a paired model–task comparison in Foundational Analysis or Sleep Assessment, with color and shape denoting task difficulty; outlined markers indicate difficulty\-level means\. Positive horizontal values favor BA in mean task score, whereas negative vertical values indicate lower bounded\-adjusted within\-task variability under BA\. The lower\-right quadrant therefore represents simultaneous improvements in performance and stability\.
### E\.3Within\-Family Performance Consistency

Figure[9](https://arxiv.org/html/2608.04156#A5.F9)examines whether task\-level performance profiles remain stable across model variants from the same family\. All four comparisons exhibit strong linear and rank agreement\. For Qwen3\.7 Plus and Qwen3\.7 Max, the Pearson correlations are \(r=0\.855\) under BrainAgent and \(r=0\.932\) under CodeAct, with corresponding Spearman correlations of \( ho=0\.868\) and \( ho=0\.936\)\. GPT\-5\.6 Terra and GPT\-5\.6 Sol show similarly strong agreement, with Pearson \(r=0\.827/0\.890\) and Spearman \( ho=0\.867/0\.858\) under BrainAgent/CodeAct, respectively; all correlations are significant at \(p<0\.001\)\. The consistent trends across Foundational Analysis and Sleep Assessment indicate that tasks that challenge one family member generally remain challenging for another, even as absolute performance changes\. This suggests that BrainBench captures stable, task\-specific capability structure rather than rankings driven only by aggregate scores\. Nevertheless, the visible deviations from the identity line show that the stronger model variant does not improve every task uniformly; these results support consistency of the task profile, not equivalence of the paired models or guaranteed gains on individual tasks\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x13.png)Figure 9:Task\-level performance consistency within model families\.Each point compares the unweighted mean instance score of the same task for two variants from one model family\. Blue circles denote the 40 Foundational Analysis tasks and orange squares denote the 43 Sleep Assessment tasks\. The first two panels compare Qwen3\.7 Plus with Qwen3\.7 Max under \(a\) BrainAgent and \(b\) CodeAct; the last two panels compare GPT\-5\.6 Terra with GPT\-5\.6 Sol under the same execution paradigms\. The gray dashed line denotes identity, the red line is the ordinary least\-squares fit, and each inset reports Pearson and Spearman correlations over the 83 pooled tasks\. Task\-difficulty weights are not applied\.
### E\.4Token Usage and Execution Efficiency

Figure[10](https://arxiv.org/html/2608.04156#A5.F10)compares total token consumption under the two execution paradigms\. Across both subsets and all evaluated models, BrainAgent consistently uses more tokens than CodeAct, reflecting the additional communication and context required for agent coordination, tool selection, and intermediate result synthesis\. CodeAct is substantially more token\-efficient because much of the analytical computation is delegated directly to executable code\. Nevertheless, token consumption is not monotonically associated with benchmark performance: several highly ranked CodeAct configurations achieve strong scores with comparatively modest token budgets, whereas larger token usage does not necessarily yield a higher rank\. These results reveal a clear performance–efficiency trade\-off between the two paradigms: structured agentic execution generally provides stronger and more reliable EEG analysis at higher token overhead, while autonomous coding reduces token consumption but exhibits less consistent performance\. As token accounting and pricing differ across model providers, the comparison should be interpreted as execution overhead rather than a direct estimate of monetary cost\.

![Refer to caption](https://arxiv.org/html/2608.04156v1/x14.png)Figure 10:Token usage and execution efficiency\.Total tokens per instance are compared between BrainAgent and CodeAct for Foundational Analysis \(a\) and Sleep Assessment \(b\)\. Markers denote means, thick bars show the interquartile range, and thin whiskers indicate the 10th–90th percentiles\. Gray lines connect the two execution paradigms for the same model, and the colored annotations report the corresponding performance ranks\. The horizontal axis uses a logarithmic scale\.
### E\.5Parser Stability and Scoring Robustness

The black\-box pipeline in Appendix[D\.1](https://arxiv.org/html/2608.04156#A4.SS1)uses two evaluator\-side LLM components described in Section[3\.2\.3](https://arxiv.org/html/2608.04156#S3.SS2.SSS3)and Appendix[C\.2](https://arxiv.org/html/2608.04156#A3.SS2): the Parser Agent aligns free\-form reports with predefined fields, while the Semantic Judge evaluates conclusions that cannot be reduced to deterministic matching\. We therefore audit both components on the complete Sleep Assessment run of Qwen3\.7 Plus under BrainAgent\. A separate, stronger Qwen3\.7 Max model is used as the audit judge with deterministic decoding\. For each of the 1,025 instances, the auditor receives the parser prompt, target\-system report, and extracted JSON, and determines only whether the extraction faithfully represents the reported answer without assessing its scientific correctness\. For each of the 225 semantic validation units, it receives the original judge rubric, judged text, and recorded score, and determines whether the score is consistent with the rubric\. Internally inconsistent audit responses are subjected to a second adjudication pass; any remaining ambiguity is retained as requiring review rather than forced into an agreement or disagreement\.

Table 12:Independent audit of evaluator\-side LLM components on the Qwen3\.7 Plus–BrainAgent Sleep Assessment run\. Agreement is computed over accepted audit decisions; unresolved results are reported separately\.Table[12](https://arxiv.org/html/2608.04156#A5.T12)shows near\-perfect agreement for output parsing, with only two disputed extractions among 1,025 instances\. This indicates that converting flexible reports into structured fields contributes little evaluator\-side noise in this setting and supports the separation between analytical competence and format compliance\. Semantic scoring is more demanding: the auditor agrees with 214 of 224 accepted decisions, while ten scores are disputed and one remains unresolved\. The resulting 95\.54% agreement suggests that the Semantic Judge is broadly consistent but constitutes a larger source of evaluation uncertainty than mechanical extraction\. Because this experiment covers one subset, target model, execution paradigm, and LLM auditor, it should be interpreted as a focused robustness check rather than a human\-annotated estimate of absolute evaluator accuracy\. Retaining the original judge inputs, outputs, and audit decisions makes the remaining disagreements traceable and available for subsequent human review\.

Similar Articles

Benchmarking LLMs

Reddit r/AI_Agents

A study or report on benchmarking large language models, likely comparing performance across various tasks.

MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction

arXiv cs.CL

MedicalBench is a new benchmark for evaluating large language models on medical concept extraction from electronic health records, focusing on implicit reasoning and evidence grounding. It includes 823 expert-annotated examples and shows that current models perform modestly, highlighting the difficulty of extracting implicitly stated medical concepts.

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

arXiv cs.CL

RedBench introduces a universal dataset aggregating 37 benchmark datasets with 29,362 samples across 22 risk categories and 19 domains to enable standardized and comprehensive red teaming evaluation of large language models. The work addresses inconsistencies in existing red teaming datasets and provides baselines, evaluation code, and open-source resources for assessing LLM robustness against adversarial prompts.