EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection
Summary
EEG-VID introduces task-guided latent predictive pretraining to improve EEG decoding under session and subject shifts, achieving above-chance target selection in assistive robotics scenarios.
View Cached Full Text
Cached at: 09/02/26, 06:17 AM
# EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection
Source: [https://arxiv.org/html/2609.00566](https://arxiv.org/html/2609.00566)
Junyi MaYuxuan WuYanzi Miao††thanks:Guanzhong Sun and Yanzi Miao are with the School of Information and Control Engineering, China University of Mining and Technology, Xuzhou, China \(e\-mail: tb22060026a41@cumt\.edu\.cn; myz@cumt\.edu\.cn\)\.††thanks:Junyi Ma is with the School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, China \(e\-mail: junyi\.ma@sjtu\.edu\.cn\)\.††thanks:Yuxuan Wu is with the School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, China, and also with the Shanghai Innovation Institute, Shanghai, China \(e\-mail: furrygreen@sjtu\.edu\.cn\)\.
###### Abstract
Objective:We investigate whether task\-guided latent predictive pretraining improves EEG decoding under session and subject shifts and whether a four\-electrode spatial posterior can support assistive target selection\.
Methods:EEG\-VID predicts future latent EEG states from recent history using an exponential\-moving\-average \(EMA\) target encoder and weak task guidance, and then fine\-tunes the pretrained model for supervised decoding\. We apply the same Stage 1 objective to five EEG backbones and replace the proposed predictor with recurrent and Transformer alternatives to test transferability and predictor dependence\.
Results:On the 48\-region cross\-day VIG\-48 task, EEG\-VID reaches 6\.52% Top\-1 and 30\.50% Top\-5 accuracy\. Across VIG\-48 and BCI Competition IV\-2a/IV\-2b, Stage 1 improves mean accuracy in 41 of 42 matched backbone–dataset–protocol comparisons, including all 12 leave\-one\-subject\-out \(LOSO\) settings, with a maximum gain of 16\.22 percentage points\. Weak task guidance improves over latent\-only pretraining in all 18 pooled model–dataset comparisons\. Alternative predictors outperform direct supervision in 48 of 54 comparisons, while the proposed predictor performs best in 45 of 54 matched comparisons\. In a separate six\-participant offline robot\-scene study, candidate\-constrained target selection reaches 40\.24% versus a 25% chance level after subject\-specific calibration\. Because gaze is unconstrained in VIG\-48, its spatial result is interpreted at the wearable\-device level rather than as source\-specific cortical decoding\.
Conclusion:Weak task guidance improves the consistency of latent predictive pretraining across EEG decoders and distribution shifts, while the resulting spatial posterior supports above\-chance scene\-constrained target selection\.
Significance:Task\-guided latent prediction provides a transferable pretraining strategy for EEG decoding under session and subject shifts and connects low\-channel EEG spatial posteriors with scene constraints for assistive target selection\.
###### Index Terms:
Assistive robotics, brain\-computer interfaces, EEG decoding, representation learning\.
## IIntroduction
Reliable target selection is a basic requirement for assistive manipulation\. Before a robot can reach or grasp an object, it must infer which target the user intends to select\. EEG provides a non\-invasive control signal for users with limited speech or manual input, but practical decoding remains difficult because the signal is noisy and varies across sessions and users\. Established convolutional and transformer\-based EEG decoders can learn effective discriminative features\[[1](https://arxiv.org/html/2609.00566#bib.bib4),[2](https://arxiv.org/html/2609.00566#bib.bib5),[3](https://arxiv.org/html/2609.00566#bib.bib6),[4](https://arxiv.org/html/2609.00566#bib.bib7)\], yet those features may shift with electrode placement, contact impedance, fatigue, and other non\-stationary factors\.
Supervised EEG decoders optimize task discrimination but do not explicitly model cross\-window predictive structure\. Latent prediction can capture temporal dynamics, yet it may also preserve predictable task\-irrelevant components such as slow drift, ocular activity, or visually evoked responses\. EEG\-VID therefore adds weak task guidance to future latent\-state prediction so that pretraining emphasizes predictable components relevant to decoding \(Fig\.[1](https://arxiv.org/html/2609.00566#S1.F1)\)\.
Fig\. 1:Motivation and paradigm comparison of EEG\-VID\.EEG\-VID combines temporal predictive consistency with weak task guidance, and restricts the decoded grid\-region posterior to scene\-derived candidates for assistive target selection\.Based on this observation, we develop EEG\-VID with task\-guided latent predictive pretraining \(Stage 1\) followed by supervised EEG decoding \(Stage 2\)\. We focus on whether this pretraining objective transfers across different EEG decoders under session and subject shifts, rather than whether it benefits EEG\-VID alone\.
We evaluate this question on the in\-house VIG\-48 cross\-day dataset and BCI Competition IV\-2a/IV\-2b\[[5](https://arxiv.org/html/2609.00566#bib.bib22),[6](https://arxiv.org/html/2609.00566#bib.bib23),[7](https://arxiv.org/html/2609.00566#bib.bib21)\]using pooled cross\-session, within\-subject, and leave\-one\-subject\-out protocols\. Stage 1 is applied to five existing EEG backbones\[[1](https://arxiv.org/html/2609.00566#bib.bib4),[2](https://arxiv.org/html/2609.00566#bib.bib5),[3](https://arxiv.org/html/2609.00566#bib.bib6),[8](https://arxiv.org/html/2609.00566#bib.bib10),[9](https://arxiv.org/html/2609.00566#bib.bib11)\], and the proposed predictor is replaced by recurrent and Transformer alternatives to further test whether the observed benefit depends on a particular predictor architecture\.
A separate six\-participant offline robot\-scene study examines the assistive relevance of the learned spatial posterior under scene\-constrained target selection\. Across all evaluations, Stage 1 improves mean accuracy in 41 of 42 matched backbone–dataset–protocol comparisons, including all 12 LOSO settings\. Alternative predictors outperform direct supervision in 48 of 54 comparisons\. Robot\-scene selection reaches 40\.24% versus a 25% chance level after subject\-specific calibration\. Because gaze is unconstrained in VIG\-48, we interpret its spatial decoding result at the wearable\-device level rather than as source\-specific cortical decoding\.
The main contributions are:
- •We propose a task\-guided latent predictive pretraining objective that combines future\-state prediction with weak task supervision to learn predictable EEG representations that remain useful for downstream decoding\.
- •We formulate Stage 1 as a backbone\-agnostic pretraining objective and evaluate its transfer through matched comparisons across different EEG decoders, predictor architectures, and session and subject shift protocols\.
- •We evaluate the proposed strategy in a low\-channel cross\-day visual decoding setting and further test its use for scene\-constrained target selection in a six\-participant offline robot\-scene study\.
## IIRelated Work
### II\-AEEG\-based intention decoding
EEG has long been used to decode user intentions and cognitive states in BCIs and rehabilitation systems\. Classical paradigms include motor imagery\[[2](https://arxiv.org/html/2609.00566#bib.bib5),[1](https://arxiv.org/html/2609.00566#bib.bib4)\], steady\-state visual evoked potentials\[[10](https://arxiv.org/html/2609.00566#bib.bib8)\], P300 event\-related potentials\[[11](https://arxiv.org/html/2609.00566#bib.bib9)\], and visual stimulus classification\[[8](https://arxiv.org/html/2609.00566#bib.bib10),[9](https://arxiv.org/html/2609.00566#bib.bib11),[12](https://arxiv.org/html/2609.00566#bib.bib55)\]\. Recent studies extend EEG\-based visual interfaces to large\-scale natural\-image and object decoding\[[13](https://arxiv.org/html/2609.00566#bib.bib40),[8](https://arxiv.org/html/2609.00566#bib.bib10),[14](https://arxiv.org/html/2609.00566#bib.bib3)\], brain\-supervised image editing\[[15](https://arxiv.org/html/2609.00566#bib.bib12)\], and 3D visual perception\[[9](https://arxiv.org/html/2609.00566#bib.bib11)\]\. Beyond decoding alone, assistive BCI systems have combined neural intent with autonomous or shared robotic control\[[16](https://arxiv.org/html/2609.00566#bib.bib43),[17](https://arxiv.org/html/2609.00566#bib.bib44),[18](https://arxiv.org/html/2609.00566#bib.bib45),[19](https://arxiv.org/html/2609.00566#bib.bib18),[20](https://arxiv.org/html/2609.00566#bib.bib46),[21](https://arxiv.org/html/2609.00566#bib.bib47)\]\. Most of these systems are trained with discriminative objectives on fixed or aggregated observation windows\. Our work instead learns predictive structure in latent EEG dynamics before supervised decoding\.
### II\-BEEG representation learning and adaptation
EEG representation learning is challenged by low signal\-to\-noise ratio, non\-stationarity, and substantial subject/session variability\[[22](https://arxiv.org/html/2609.00566#bib.bib13)\]\. These distribution shifts have motivated transfer, alignment, and subject\-independent learning for cross\-session and cross\-subject EEG decoding\[[23](https://arxiv.org/html/2609.00566#bib.bib27),[24](https://arxiv.org/html/2609.00566#bib.bib28),[25](https://arxiv.org/html/2609.00566#bib.bib29),[26](https://arxiv.org/html/2609.00566#bib.bib30),[27](https://arxiv.org/html/2609.00566#bib.bib14),[28](https://arxiv.org/html/2609.00566#bib.bib26),[29](https://arxiv.org/html/2609.00566#bib.bib48),[30](https://arxiv.org/html/2609.00566#bib.bib49),[31](https://arxiv.org/html/2609.00566#bib.bib50),[32](https://arxiv.org/html/2609.00566#bib.bib15),[33](https://arxiv.org/html/2609.00566#bib.bib54)\]\. Existing approaches further include signal preprocessing and channel selection\[[4](https://arxiv.org/html/2609.00566#bib.bib7)\], spatiotemporal neural architectures\[[1](https://arxiv.org/html/2609.00566#bib.bib4),[2](https://arxiv.org/html/2609.00566#bib.bib5),[3](https://arxiv.org/html/2609.00566#bib.bib6),[34](https://arxiv.org/html/2609.00566#bib.bib25)\], and lightweight or subject\-specific calibration\.
Self\-supervised and pretrained EEG representation learning spans contrastive, temporal\-context, and other pretext objectives\[[35](https://arxiv.org/html/2609.00566#bib.bib31),[36](https://arxiv.org/html/2609.00566#bib.bib32),[37](https://arxiv.org/html/2609.00566#bib.bib20),[38](https://arxiv.org/html/2609.00566#bib.bib51)\], masked or predictive representation learning\[[39](https://arxiv.org/html/2609.00566#bib.bib19),[40](https://arxiv.org/html/2609.00566#bib.bib52)\], and larger cross\-dataset pretrained models\[[41](https://arxiv.org/html/2609.00566#bib.bib33),[42](https://arxiv.org/html/2609.00566#bib.bib34),[43](https://arxiv.org/html/2609.00566#bib.bib35),[44](https://arxiv.org/html/2609.00566#bib.bib36)\]\. Purely self\-supervised predictive objectives do not explicitly prioritize predictable EEG components that are informative for a specific downstream task\.
### II\-CLatent predictive modeling
Predictive representation learning models future or missing information in representation space rather than reconstructing raw observations\. Contrastive Predictive Coding learns representations by predicting future latent states\[[45](https://arxiv.org/html/2609.00566#bib.bib37)\], while data2vec predicts contextualized latent targets\[[46](https://arxiv.org/html/2609.00566#bib.bib38)\]\. Target\-encoder methods such as BYOL and I\-JEPA further demonstrate predictive learning directly in representation space\[[47](https://arxiv.org/html/2609.00566#bib.bib39),[48](https://arxiv.org/html/2609.00566#bib.bib16)\]\. For EEG, latent prediction avoids raw\-signal reconstruction, which may require the model to represent task\-irrelevant signal variation\. The key challenge is to retain predictable components that are informative for the downstream task\. We study whether weak task guidance can retain task\-relevant structure within predictive EEG representations across backbones and subject shifts\.
## IIIMethods

Fig\. 2:Overall architecture of EEG\-VID\.Stage 1 learns latent EEG\-state transition dynamics by predicting future EEG representations from historical EEG windows, supervised by an EMA target encoder and guided by a weak auxiliary classification term\. Stage 2 transfers the learned dynamic representation to supervised visual intention region decoding\.This section presents EEG\-VID, a two\-stage framework for region\-level visual intention decoding \(Fig\.[2](https://arxiv.org/html/2609.00566#S3.F2)\)\. EEG\-VID first learns latent transitions from historical to future EEG representations through latent predictive pretraining, and then transfers the learned dynamic representations to supervised region decoding\. The key idea is that future latent\-state prediction provides a temporal consistency constraint, while weak task guidance encourages the learned predictable structure to remain relevant to downstream decoding\.
### III\-AProblem Formulation
For each visual target selection trial, the user observes a visual scene divided into anN×MN\\times Mgrid, corresponding toNMNMcandidate regions\. Given a sequence of historical EEG windows
𝐗t=\{𝐱t−K\+1,…,𝐱t\},\\mathbf\{X\}\_\{t\}=\\\{\\mathbf\{x\}\_\{t\-K\+1\},\\ldots,\\mathbf\{x\}\_\{t\}\\\},where𝐱t∈ℝL×C\\mathbf\{x\}\_\{t\}\\in\\mathbb\{R\}^\{L\\times C\}denotes thett\-th EEG window andKK,LL,CCdenote the number of historical windows, the window length and the number of EEG channels, the goal is to predict the region labelht∈\{0,…,NM−1\}h\_\{t\}\\in\\\{0,\\ldots,NM\-1\\\}of the user’s intended target\.
Rather than decoding each EEG window independently, EEG\-VID models temporal transitions between historical and future latent EEG states\. During training, the model learns to predict future latent EEG representations from historical windows\. At inference, onlyXtX\_\{t\}is required\.
### III\-BMulti\-scale Temporal\-Statistical Hybrid EEG Encoder
The Online EncoderEθE\_\{\\theta\}maps each EEG window to a latent EEG\-state representation\. To combine learned temporal patterns with stable signal descriptors, we use a two\-branch encoder \(Fig\.[3](https://arxiv.org/html/2609.00566#S3.F3)\)\. The temporal branch models multi\-scale dynamics, while the statistical branch summarizes additional signal statistics\.
#### III\-B1Temporal branch
A multi\-scale temporal stem applies parallel one\-dimensional convolutions with kernel sizes\{3,7,11,15\}\\\{3,7,11,15\\\},
𝐮τ\(k\)=Conv1Dk\(𝐱τ\),k∈𝒦,\\mathbf\{u\}^\{\(k\)\}\_\{\\tau\}=\\mathrm\{Conv1D\}\_\{k\}\(\\mathbf\{x\}\_\{\\tau\}\),\\qquad k\\in\\mathcal\{K\},\(1\)whose outputs are concatenated along the channel dimension, normalized, activated by GELU and projected to a common embedding dimension by a1×11\\times 1convolution,
𝐔τ=FrontProj\(Concatk∈𝒦𝐮τ\(k\)\),\\mathbf\{U\}\_\{\\tau\}=\\mathrm\{FrontProj\}\\\!\\left\(\\mathrm\{Concat\}\_\{k\\in\\mathcal\{K\}\}\\mathbf\{u\}^\{\(k\)\}\_\{\\tau\}\\right\),\(2\)followed by BatchNorm, GELU and a channel squeeze\-and\-excitation module\. Two dilated residual blocks \(dilation11and22\) enlarge the receptive field, learnable positional embeddings are added, and a Transformer encoder captures long\-range dependencies within the window,
𝐇τ=Transformer\(ResBlock\(𝐔τ\)\+𝐏\)\.\\mathbf\{H\}\_\{\\tau\}=\\mathrm\{Transformer\}\\\!\\left\(\\mathrm\{ResBlock\}\(\\mathbf\{U\}\_\{\\tau\}\)\+\\mathbf\{P\}\\right\)\.\(3\)LayerNorm and attention pooling then yield the window\-level temporal representation𝐳τtemp=AttnPool\(LayerNorm\(𝐇τ\)\)\\mathbf\{z\}^\{\\mathrm\{temp\}\}\_\{\\tau\}=\\mathrm\{AttnPool\}\(\\mathrm\{LayerNorm\}\(\\mathbf\{H\}\_\{\\tau\}\)\)\.
#### III\-B2Statistical branch
For each window, we compute six channel\-wise temporal descriptors—mean, standard deviation, RMS, absolute mean, first\-difference standard deviation, and line length—together with pairwise channel correlations and binned log\-power spectral features\. Their concatenation𝐫τ=\[𝐫τtime‖𝐫τcorr‖𝐫τfreq\]\\mathbf\{r\}\_\{\\tau\}=\[\\mathbf\{r\}^\{\\mathrm\{time\}\}\_\{\\tau\}\\\|\\mathbf\{r\}^\{\\mathrm\{corr\}\}\_\{\\tau\}\\\|\\mathbf\{r\}^\{\\mathrm\{freq\}\}\_\{\\tau\}\]is mapped by a projection MLP to the statistical representation𝐳τstat=Pη\(𝐫τ\)\\mathbf\{z\}^\{\\mathrm\{stat\}\}\_\{\\tau\}=P\_\{\\eta\}\(\\mathbf\{r\}\_\{\\tau\}\)\.
#### III\-B3Interaction fusion
The two branches are fused through an explicit interaction feature
𝐪τ=\[𝐳τtemp∥𝐳τstat∥𝐳τtemp⊙𝐳τstat∥\|𝐳τtemp−𝐳τstat\|\],\\mathbf\{q\}\_\{\\tau\}=\\left\[\\mathbf\{z\}^\{\\mathrm\{temp\}\}\_\{\\tau\}\\;\\\|\\;\\mathbf\{z\}^\{\\mathrm\{stat\}\}\_\{\\tau\}\\;\\\|\\;\\mathbf\{z\}^\{\\mathrm\{temp\}\}\_\{\\tau\}\\odot\\mathbf\{z\}^\{\\mathrm\{stat\}\}\_\{\\tau\}\\;\\\|\\;\\left\|\\mathbf\{z\}^\{\\mathrm\{temp\}\}\_\{\\tau\}\-\\mathbf\{z\}^\{\\mathrm\{stat\}\}\_\{\\tau\}\\right\|\\right\],\(4\)where∥\\\|is concatenation and⊙\\odotelement\-wise multiplication\. The product term captures interactions between the two branches, while the absolute\-difference term captures differences between their features\. A fusion MLP \(LayerNorm–Linear–GELU–Dropout–Linear\) produces the window\-level latent EEG state
𝐒τ=Fθ\(𝐪τ\),𝐒τ∈ℝd,\\mathbf\{S\}\_\{\\tau\}=F\_\{\\theta\}\(\\mathbf\{q\}\_\{\\tau\}\),\\qquad\\mathbf\{S\}\_\{\\tau\}\\in\\mathbb\{R\}^\{d\},\(5\)and we write the complete mapping as𝐒τ=Eθ\(𝐱τ\)\\mathbf\{S\}\_\{\\tau\}=E\_\{\\theta\}\(\\mathbf\{x\}\_\{\\tau\}\)\. ApplyingEθE\_\{\\theta\}to each window with shared parameters gives the historical latent sequence
𝐒t−K\+1:t=Eθ\(𝐗t\)∈ℝK×d\.\\mathbf\{S\}\_\{t\-K\+1:t\}=E\_\{\\theta\}\(\\mathbf\{X\}\_\{t\}\)\\in\\mathbb\{R\}^\{K\\times d\}\.\(6\)

Fig\. 3:Architecture of the multi\-scale temporal\-statistical hybrid EEG encoder\.A temporal branch extracts multi\-scale temporal patterns, while a statistical branch extracts stable statistical descriptors\. The two branches are fused through interaction features and projected into the EEG latent representation\.
### III\-CTask\-Guided Latent Predictive Pretraining \(Stage 1\)
Given𝐒t−K\+1:t\\mathbf\{S\}\_\{t\-K\+1:t\}, an intention modeling branch aggregates the historical latent tokens and outputs an intention representation𝐜t\\mathbf\{c\}\_\{t\}together with auxiliary region logits𝐨taux\\mathbf\{o\}^\{\\mathrm\{aux\}\}\_\{t\}:
\(𝐜t,𝐨taux\)=Gψ\(𝐒t−K\+1:t\),𝐨taux∈ℝNM\.\(\\mathbf\{c\}\_\{t\},\\mathbf\{o\}^\{\\mathrm\{aux\}\}\_\{t\}\)=G\_\{\\psi\}\(\\mathbf\{S\}\_\{t\-K\+1:t\}\),\\qquad\\mathbf\{o\}^\{\\mathrm\{aux\}\}\_\{t\}\\in\\mathbb\{R\}^\{NM\}\.\(7\)GψG\_\{\\psi\}is an Intent Transformer operating only on theKKhistorical latent tokens\. A learnable summary token produces𝐜t\\mathbf\{c\}\_\{t\}, and masked self\-attention prevents access to the future EEG window used as the prediction target\. The summary\-token mask is defined so that it can aggregate all historical tokens while remaining independent of𝐱t\+1\\mathbf\{x\}\_\{t\+1\}\.
The latent predictor is conditioned on both the history and the intention representation,
\(𝐒^t\+1,𝐰t\)=Wϕ\(𝐒t−K\+1:t,𝐜t\),\(\\hat\{\\mathbf\{S\}\}\_\{t\+1\},\\mathbf\{w\}\_\{t\}\)=W\_\{\\phi\}\(\\mathbf\{S\}\_\{t\-K\+1:t\},\\mathbf\{c\}\_\{t\}\),\(8\)where𝐒^t\+1\\hat\{\\mathbf\{S\}\}\_\{t\+1\}is the predicted future latent EEG state and𝐰t\\mathbf\{w\}\_\{t\}is the EEG\-state dynamics representation\. Structurally,WϕW\_\{\\phi\}contains a latent\-token projection module, a stack of predictor blocks, and a lightweight predictor head\. Each predictor block combines rotary positional embeddings \(RoPE\), attention, and an MLP\.
Following the target\-encoder principle used in predictive representation learning\[[47](https://arxiv.org/html/2609.00566#bib.bib39),[48](https://arxiv.org/html/2609.00566#bib.bib16)\], the prediction target is produced by an EMA target encoderEθ¯E\_\{\\bar\{\\theta\}\}that shares the architecture ofEθE\_\{\\theta\}but is updated without gradients,
θ¯←λemaθ¯\+\(1−λema\)θ,𝐒t\+1=Eθ¯\(𝐱t\+1\),\\bar\{\\theta\}\\leftarrow\\lambda\_\{\\mathrm\{ema\}\}\\bar\{\\theta\}\+\(1\-\\lambda\_\{\\mathrm\{ema\}\}\)\\theta,\\qquad\\mathbf\{S\}\_\{t\+1\}=E\_\{\\bar\{\\theta\}\}\(\\mathbf\{x\}\_\{t\+1\}\),\(9\)where the target\-encoder EMA decay is fixed toλema=0\.995\\lambda\_\{\\mathrm\{ema\}\}=0\.995\. A stop\-gradientsg\(⋅\)\\mathrm\{sg\}\(\\cdot\)is applied to𝐒t\+1\\mathbf\{S\}\_\{t\+1\}\. Together with the EMA target encoder, this stabilizes the prediction target and helps prevent trivial co\-adaptation between the online encoder and the target representation\.
We separate the magnitude and directional components of future\-state prediction as
ℒlatent=‖𝐒^t\+1−sg\(𝐒t\+1\)‖22,\\mathcal\{L\}\_\{\\mathrm\{latent\}\}=\\left\\\|\\hat\{\\mathbf\{S\}\}\_\{t\+1\}\-\\mathrm\{sg\}\(\\mathbf\{S\}\_\{t\+1\}\)\\right\\\|\_\{2\}^\{2\},\(10\)
ℒcos=1−cos\(𝐒^t\+1,sg\(𝐒t\+1\)\)\.\\mathcal\{L\}\_\{\\mathrm\{cos\}\}=1\-\\cos\\left\(\\hat\{\\mathbf\{S\}\}\_\{t\+1\},\\mathrm\{sg\}\(\\mathbf\{S\}\_\{t\+1\}\)\\right\)\.\(11\)
#### Weak task guidance
Temporal prediction alone does not determine which predictable EEG components are useful for downstream decoding because slowly varying task\-irrelevant activity can also be predictable\. We therefore introduce a weak auxiliary intention\-classification term
ℒintent=CE\(𝐨taux,ht\)\.\\mathcal\{L\}\_\{\\mathrm\{intent\}\}=\\mathrm\{CE\}\(\\mathbf\{o\}^\{\\mathrm\{aux\}\}\_\{t\},h\_\{t\}\)\.\(12\)The complete Stage 1 objective is
ℒS1=λlatentℒlatent\+λcosℒcos\+λclsS1ℒintent\.\\mathcal\{L\}\_\{S1\}=\\lambda\_\{\\mathrm\{latent\}\}\\mathcal\{L\}\_\{\\mathrm\{latent\}\}\+\\lambda\_\{\\mathrm\{cos\}\}\\mathcal\{L\}\_\{\\mathrm\{cos\}\}\+\\lambda\_\{\\mathrm\{cls\}\}^\{S1\}\\mathcal\{L\}\_\{\\mathrm\{intent\}\}\.\(13\)Unless otherwise stated, we useλlatent=1\.0\\lambda\_\{\\mathrm\{latent\}\}=1\.0,λcos=0\.1\\lambda\_\{\\mathrm\{cos\}\}=0\.1,λclsS1=0\.1\\lambda\_\{\\mathrm\{cls\}\}^\{S1\}=0\.1, andλema=0\.995\\lambda\_\{\\mathrm\{ema\}\}=0\.995\. The small classification weight allows this term to guide the learned representation without replacing the latent prediction objective\. Accordingly, Section[V\-D](https://arxiv.org/html/2609.00566#S5.SS4)compares the default withλclsS1∈\{0,0\.2,0\.3\}\\lambda\_\{\\mathrm\{cls\}\}^\{S1\}\\in\\\{0,0\.2,0\.3\\\}while keeping the other Stage 1 coefficients fixed\. This sweep tests the balance between prediction and task supervision over these four settings\. It does not show that 0\.1 is globally optimal\.
Because Stage 1 uses the training\-split labels throughℒintent\\mathcal\{L\}\_\{\\mathrm\{intent\}\}, the procedure is a weakly supervised pretraining stage rather than a self\-supervised one\.
### III\-DSupervised Visual Intention Region Decoding \(Stage 2\)
Stage 2 transfers the online encoder, the intention branch and the latent predictor to supervised decoding\. Instead of relying on the current window alone, the decoder fuses current evidence, temporally aggregated intention context and predicted dynamics,
𝐮t=\[𝐒t‖𝐜t‖𝐰t\],ℓt=Cω\(𝐮t\)∈ℝNM,\\mathbf\{u\}\_\{t\}=\[\\mathbf\{S\}\_\{t\}\\;\\\|\\;\\mathbf\{c\}\_\{t\}\\;\\\|\\;\\mathbf\{w\}\_\{t\}\],\\qquad\\boldsymbol\{\\ell\}\_\{t\}=C\_\{\\omega\}\(\\mathbf\{u\}\_\{t\}\)\\in\\mathbb\{R\}^\{NM\},\(14\)and predictsh^t=argmaxiℓt,i\\hat\{h\}\_\{t\}=\\arg\\max\_\{i\}\\boldsymbol\{\\ell\}\_\{t,i\}\.
Because the labels live on a 2\-D grid, a coordinate consistency term exploits the spatial structure\. For class indexi∈\{0,…,NM−1\}i\\in\\\{0,\\ldots,NM\-1\\\}, we define its row and column coordinates asρi=⌊i/M⌋\\rho\_\{i\}=\\lfloor i/M\\rfloorandκi=imodM\\kappa\_\{i\}=i\\bmod M\. For the ground\-truth classhth\_\{t\}, letρt⋆=ρht\\rho\_\{t\}^\{\\star\}=\\rho\_\{h\_\{t\}\}andκt⋆=κht\\kappa\_\{t\}^\{\\star\}=\\kappa\_\{h\_\{t\}\}\. With𝐩t=softmax\(ℓt\)\\mathbf\{p\}\_\{t\}=\\mathrm\{softmax\}\(\\boldsymbol\{\\ell\}\_\{t\}\), the expected coordinates areρ^t=∑i𝐩t,iρi\\hat\{\\rho\}\_\{t\}=\\sum\_\{i\}\\mathbf\{p\}\_\{t,i\}\\rho\_\{i\}andκ^t=∑i𝐩t,iκi\\hat\{\\kappa\}\_\{t\}=\\sum\_\{i\}\\mathbf\{p\}\_\{t,i\}\\kappa\_\{i\},
ℒcoord=SmoothL1\(ρ^t,ρt⋆\)\+SmoothL1\(κ^t,κt⋆\)\.\\mathcal\{L\}\_\{\\mathrm\{coord\}\}=\\mathrm\{SmoothL1\}\(\\hat\{\\rho\}\_\{t\},\\rho\_\{t\}^\{\\star\}\)\+\\mathrm\{SmoothL1\}\(\\hat\{\\kappa\}\_\{t\},\\kappa\_\{t\}^\{\\star\}\)\.\(15\)
We retain a weak predictive\-consistency term during fine\-tuning to encourage preservation of the learned transition structure\. We define the retention loss using the same predictive weights as in Stage 1,
ℒkeep=λlatentℒlatent\+λcosℒcos,\\mathcal\{L\}\_\{\\mathrm\{keep\}\}=\\lambda\_\{\\mathrm\{latent\}\}\\mathcal\{L\}\_\{\\mathrm\{latent\}\}\+\\lambda\_\{\\mathrm\{cos\}\}\\mathcal\{L\}\_\{\\mathrm\{cos\}\},\(16\)whereλlatent=1\.0\\lambda\_\{\\mathrm\{latent\}\}=1\.0andλcos=0\.1\\lambda\_\{\\mathrm\{cos\}\}=0\.1are shared with Stage 1\.
The complete Stage 2 objective is
ℒstage2=λclsℒcls\+λcoordℒcoord\+λkeepℒkeep,\\mathcal\{L\}\_\{\\mathrm\{stage2\}\}=\\lambda\_\{\\mathrm\{cls\}\}\\mathcal\{L\}\_\{\\mathrm\{cls\}\}\+\\lambda\_\{\\mathrm\{coord\}\}\\mathcal\{L\}\_\{\\mathrm\{coord\}\}\+\\lambda\_\{\\mathrm\{keep\}\}\\mathcal\{L\}\_\{\\mathrm\{keep\}\},\(17\)whereℒcls=CE\(ℓt,ht\)\\mathcal\{L\}\_\{\\mathrm\{cls\}\}=\\mathrm\{CE\}\(\\boldsymbol\{\\ell\}\_\{t\},h\_\{t\}\)\. Unless otherwise stated, we useλcls=1\.0\\lambda\_\{\\mathrm\{cls\}\}=1\.0,λcoord=0\.1\\lambda\_\{\\mathrm\{coord\}\}=0\.1, andλkeep=0\.03\\lambda\_\{\\mathrm\{keep\}\}=0\.03\. The future EEG window is used only to computeℒkeep\\mathcal\{L\}\_\{\\mathrm\{keep\}\}during training and is never required at inference\. On datasets without a 2\-D label geometry \(Section[IV](https://arxiv.org/html/2609.00566#S4)\),λcoord\\lambda\_\{\\mathrm\{coord\}\}is set to zero andNMNMis replaced by the number of classes\.
### III\-ERobot Target Selection
The user selects one of four candidate objects without speech, gesture, or manual command\. No dedicated eye tracker is used\. For each robot\-camera observation, SAM\[[49](https://arxiv.org/html/2609.00566#bib.bib17)\]first segments the four objects present in the scene,𝒪=\{o1,…,oJ\}\\mathcal\{O\}=\\\{o\_\{1\},\\ldots,o\_\{J\}\\\}withJ=4J=4\. The image plane is divided into the same6×86\\times 8grid used by EEG\-VID, and each segmented object is mapped to a grid labelg\(oj\)g\(o\_\{j\}\)\. SAM therefore defines the scene\-dependent candidate label set
𝒞=\{g\(o1\),…,g\(oJ\)\},J=4\.\\mathcal\{C\}=\\\{g\(o\_\{1\}\),\\ldots,g\(o\_\{J\}\)\\\},\\qquad J=4\.\(18\)EEG\-VID independently produces logitsℓt∈ℝ48\\boldsymbol\{\\ell\}\_\{t\}\\in\\mathbb\{R\}^\{48\}from EEG\. The final target is the object whose SAM\-derived grid label receives the largest EEG logit,
o^=argmaxoj∈𝒪ℓt,g\(oj\)\.\\hat\{o\}=\\arg\\max\_\{o\_\{j\}\\in\\mathcal\{O\}\}\\boldsymbol\{\\ell\}\_\{t,g\(o\_\{j\}\)\}\.\(19\)SAM supplies only the four candidate spatial labels\. It does not infer the user’s intended target\. EEG\-VID makes the intention decision by restricting its posterior to those labels\. Because every evaluated scene contains four candidate objects, random selection within the candidate set has a 25% chance level\.
We evaluate this decision rule in an offline robot\-scene experiment using six tabletop configurations, each containing four objects at different spatial positions\. For each RGB observation, SAM produces the object masks, and each object is assigned to the grid cell containing the largest number of pixels from its mask\. The subject\-specific EEG\-VID model supplies the 48\-way spatial scores, after which \([19](https://arxiv.org/html/2609.00566#S3.E19)\) selects among the four occupied candidate cells\.
The selected mask can be fused with depth for downstream grasp planning \(AnyGrasp\[[50](https://arxiv.org/html/2609.00566#bib.bib24)\]\)\. The reported outcome is*target selection*before motion planning\. Grasp execution is not evaluated\.
### III\-FData Acquisition, Ethics, and Preprocessing
#### III\-F1Ethics statement
Seven healthy adults participated in total\. One contributed the VIG\-48 cross\-day corpus, and six contributed the robot\-scene set\. All participants provided written informed consent\. The study was approved by the Ethics Committee of Xuzhou Central Hospital \(No\. XZXY\-LJ\-20210513\-054\)\. This study was not a clinical trial\.
#### III\-F2Acquisition and preprocessing
Fig\. 4:Four\-electrode Muse S Athena montage used for the in\-house recordings at AF7, AF8, TP9, and TP10\. The layout is shown for acquisition reproducibility and is not used for source\-localization claims\.EEG was recorded at 256 Hz with a Muse S Athena using AF7, AF8, TP9, and TP10\. Participants viewed a6×86\\times 8grid\. A target cell at a randomly selected position was continuously highlighted for 3–6 s\. The target did not flicker, so no frequency\-tagged SSVEP cue was used\. No dedicated eye tracker and no gaze\-contingent hardware were used at any stage\. Participants were not instructed to fixate centrally\. Under this single\-device protocol, natural orienting toward the intended cell remains part of the recorded input rather than being explicitly suppressed\. The consequences of this choice for what can be claimed about the decoded signal are stated in Section[VI\-B](https://arxiv.org/html/2609.00566#S6.SS2)\.
Recordings were segmented into non\-overlapping 1\-s windows\. A fixed preprocessing pipeline—non\-finite\-value handling, outlier suppression, fourth\-order 0\.5–45 Hz band\-pass filtering, and temporal smoothing—was applied consistently across all in\-house experiments\. No ICA or dedicated ocular\-artifact rejection was applied\. With four scalp channels and no reference EOG, ocular components cannot be identified reliably\. Removing orienting\-related activity could also suppress information available to the device\-level interface\. We therefore retain a fixed preprocessing pipeline and interpret VIG\-48 at the device level rather than as cortical source decoding\.
## IVExperimental Setup
### IV\-ADatasets
To evaluate decoding performance and transferability, we use the in\-house VIG\-48 dataset and two public BCI Competition IV benchmarks\. We further conduct a separate robot\-scene experiment to assess scene\-informed target selection\.
#### IV\-A1In\-house EEG visual intention dataset \(VIG\-48\)
VIG\-48 contains approximately 15,000 window\-level samples from one participant recorded across multiple days\. Approximately 13,000 samples are used for model development, and the held\-out test split is collected on three days that do not overlap the training days, so that evaluation measures cross\-day distribution shift rather than within\-day generalization\. Within the model\-development portion, training and validation are separated by complete event identifiers before Stage 1 context construction\. All windows from a given cued event remain in one split\. The task is region\-level decoding over a6×86\\times 8grid with 48 candidate regions\. The Top\-1 and Top\-5 chance levels are2\.08%2\.08\\%and10\.42%10\.42\\%, respectively\. Because multiple windows originate from the same cued trial, individual windows are not treated as independent observations for a binomial significance test\.
#### IV\-A2Robot\-scene target\-selection set
A separate in\-house robot\-scene set is collected from six additional participants across six tabletop scene configurations\. Every test scene contains four candidate objects whose spatial positions vary across configurations\. Before robot\-scene evaluation, each participant completes 48 subject\-specific calibration trials using the standard visual\-intention acquisition paradigm\. Each calibration trial lasts 10 s with one target label, providing 480 s of calibration EEG per participant\. Robot\-scene test trials are collected separately and last 30 s\. During testing, each participant views the scene from the robot\-camera viewpoint\. RGB and depth images from this viewpoint are recorded together with the EEG trials\. Seven robot\-scene test trials are excluded because of data\-saving or processing failures\.
For each participant, EEG\-VID is initialized from the visual\-intention model and adapted using only that participant’s calibration data\. This adaptation is performed in Stage 2 with supervised classification fine\-tuning; Stage 1 is not re\-run on the robot calibration set\. The adapted model is then evaluated on the participant’s held\-out robot\-scene trials\. The reported endpoint is window\-level candidate\-constrained target\-selection accuracy, defined as the number of windows for which the correct object’s grid label has the largest EEG score among the four SAM\-derived candidate labels divided by the total number of valid test windows\.
#### IV\-A3Public benchmarks
To test whether the proposed pretraining principle transfers beyond our acquisition setup, we evaluate on BCI Competition IV dataset 2a\[[5](https://arxiv.org/html/2609.00566#bib.bib22)\]and dataset 2b\[[6](https://arxiv.org/html/2609.00566#bib.bib23)\], released as part of BCI Competition IV\[[7](https://arxiv.org/html/2609.00566#bib.bib21)\]\. IV\-2a is a four\-class motor\-imagery benchmark with 22 EEG channels, and IV\-2b is a two\-class motor\-imagery benchmark with three EEG channels\. EOG channels are excluded in both datasets\. These datasets are not visual\-intention tasks\. They are included to test whether Stage 1 transfers to standard EEG decoding settings\. Chance levels are 25% for IV\-2a and 50% for IV\-2b\. The three evaluation protocols applied to them are defined in Section[IV\-B](https://arxiv.org/html/2609.00566#S4.SS2)\.
### IV\-BEvaluation Protocols
Absolute accuracy on IV\-2a and IV\-2b depends strongly on how models are trained and how decisions are made\. Rather than reporting a single configuration, we evaluate every method under three protocols that differ in the amount of subject\-specific data available at training time\. All three use the same preprocessing, 1\-s windowing, optimizer settings, and the same five fixed random seeds\. In all three protocols, the classification decision is made per 1\-s window rather than per trial\.
For every protocol, raw trials/events retain an event identifier through preprocessing and windowing\. Training and validation are then separated by complete event identifiers, so all windows from one event remain in the same split\. Data\-dependent normalization statistics are estimated from training windows only and applied unchanged to validation and test data\. Stage 1 history–target contexts are constructed only after this split and independently within each split and each event; no context crosses an event boundary or a train/validation/test boundary\.
#### IV\-B1Pooled cross\-session \(subject\-overlapping\)
A single subject\-agnostic model is trained on the pooled training sessions of all nine subjects and evaluated on the pooled independent evaluation sessions of the same nine subjects\. For IV\-2a, the training pool is A01T–A09T and the test set is A01E–A09E\. For IV\-2b, the training pool contains Bxx01T, Bxx02T, and Bxx03T, while the test set contains Bxx04E and Bxx05E \(xx=01,…,09=01,\\ldots,09\)\. Training and testing therefore share subjects but not sessions\. For each seed, 20% of the pooled training events are held out for validation \(val\_ratio=0\.2\\mathrm\{val\\\_ratio\}=0\.2\), with all windows from an event kept together\. Evaluation sessions are never used for model selection\. After segmentation, IV\-2a contains 10,368 training\-pool windows and 10,368 test windows, and IV\-2b contains 14,720 and 11,360\. Each IV\-2a trial contributes exactly four consecutive 1\-s windows \(9subjects×288trials×4=10,3689\\ \\text\{subjects\}\\times 288\\ \\text\{trials\}\\times 4=10\{,\}368\), taken from the fixed 4\-s motor\-imagery interval\. This protocol measures session shift with subject identity held fixed and is used for the main comparison and for all ablations\.
#### IV\-B2Within\-subject
One model is trained per subject on that subject’s own training session\(s\) and evaluated on that subject’s own evaluation session\(s\), with no data from any other subject\. The validation subset is drawn at the event level from that subject’s training session\(s\), never from the evaluation session\(s\)\. This matches the conventional way IV\-2a and IV\-2b are reported, except that decisions remain window\-level\. It measures how each method behaves when subject\-specific data are available but scarce\.
#### IV\-B3Cross\-subject \(leave\-one\-subject\-out\)
For each held\-out subject, the model is trained on the remaining eight subjects and evaluated on the held\-out subject’s sessions, with no calibration data from that subject\. Validation events are selected only from the eight training subjects\. The held\-out subject is never used for normalization, model selection, or checkpoint selection\. This is the strictest of the three and the one most relevant to a deployable interface, since a new user supplies no labelled data\.
#### IV\-B4Interpreting the absolute values
All IV\-2a/IV\-2b results are 1\-s\-window accuracies, whereas standard competition protocols classify complete trials\. We also do not use sliding\-window augmentation, filter\-bank features, or covariance alignment\[[27](https://arxiv.org/html/2609.00566#bib.bib14)\]\. Absolute values therefore should not be compared directly with published competition results\. Because all methods share the same preprocessing and decision protocol, the within\-table comparisons test the training strategy under a common controlled setting\.
### IV\-CWindow Construction and Temporal Context
All datasets are segmented into consecutive, non\-overlapping 1\-s EEG windows\. IV\-2a/IV\-2b contain 250 samples per window \(250 Hz\), whereas the in\-house data contain 256 samples \(256 Hz\)\. A fourth\-order 0\.5–45 Hz band\-pass filter is used throughout\. We fix the history length toK=4K=4\.
For consecutive windows\{𝐱0,𝐱1,…\}\\\{\\mathbf\{x\}\_\{0\},\\mathbf\{x\}\_\{1\},\\ldots\\\}, Stage 1 uses
\[𝐱t−3,𝐱t−2,𝐱t−1,𝐱t\]⟶𝐱t\+1\.\[\\mathbf\{x\}\_\{t\-3\},\\mathbf\{x\}\_\{t\-2\},\\mathbf\{x\}\_\{t\-1\},\\mathbf\{x\}\_\{t\}\]\\longrightarrow\\mathbf\{x\}\_\{t\+1\}\.\(20\)Thus, each prediction uses at most 4 s of history to predict the next 1\-s window\. Missing history at the start of an event is left\-padded by repeating the first window\. Raw windows do not overlap, although adjacent history–target samples from the same event share three historical windows\. This overlap occurs only within a single event and a single data split\. Stage 1 contexts are generated after event\-level splitting and never cross event or train/validation/test boundaries\. The terminal window of each event is omitted whenever a future target is required\.
For IV\-2a and IV\-2b, where an event contributes only four windows, this scheme yields three Stage 1 samples per trial and the first two are partly left\-padded\. The effective history is therefore shorter than 4 s on these benchmarks\. On the in\-house recordings, where each highlighted target lasts 3–6 s, full\-length histories are available for most samples\. We report this asymmetry because it may explain part of the variation in Stage 1 gains across datasets\.
### IV\-DBaselines, Variants and Metrics
We compare EEG\-VID with EEGNet\[[2](https://arxiv.org/html/2609.00566#bib.bib5)\], DeepConvNet\[[1](https://arxiv.org/html/2609.00566#bib.bib4)\], EEG\-Conformer\[[3](https://arxiv.org/html/2609.00566#bib.bib6)\], TSConv\[[8](https://arxiv.org/html/2609.00566#bib.bib10)\], Neuro\-3D\[[9](https://arxiv.org/html/2609.00566#bib.bib11)\], EEG2Rep\[[39](https://arxiv.org/html/2609.00566#bib.bib19)\], and BENDR\[[37](https://arxiv.org/html/2609.00566#bib.bib20)\]\. EEG\-VID w/o Stage 1 retains the supervised EEG\-VID architecture but omits predictive pretraining\. All reported baseline numbers are produced by our local implementations under the common data splits and preprocessing pipeline rather than copied from the cited papers\. Unless stated otherwise, results use five fixed random seeds and are reported as mean±\\pmstandard deviation\. We report Top\-1 accuracy on all datasets and Top\-5 accuracy on the 48\-way in\-house task\. The public\-benchmark accuracies are reported at the 1\-s\-window level under each of the three protocols of Section[IV\-B](https://arxiv.org/html/2609.00566#S4.SS2)\. For the robot\-scene experiment we report SAM\-constrained target\-selection accuracy, i\.e\., the fraction of held\-out test cases for which the highest\-logit candidate object matches the intended object\. With four candidates, random chance is 25%\.
### IV\-EStage 1 Transfer to Existing EEG Backbones
To separate the training objective from the EEG\-VID encoder, Stage 1 is also applied to EEGNet, DeepConvNet, EEG\-Conformer, Neuro\-3D and TSConv\. For each backboneBB, the classifier output is replaced during Stage 1 by a latent adaptation layer, giving𝐳t=B\(𝐱t\)\\mathbf\{z\}\_\{t\}=B\(\\mathbf\{x\}\_\{t\}\)\. An EMA copy of the same backbone provides the stopped\-gradient future target𝐳t\+1target\\mathbf\{z\}^\{\\mathrm\{target\}\}\_\{t\+1\}, and a predictor estimates that target from theKKhistorical latent states\. A lightweight auxiliary classifier uses the historical representation and the same training labels as EEG\-VID\. Thus, the external\-backbone variants use the same three Stage 1 loss components and coefficients in \([13](https://arxiv.org/html/2609.00566#S3.E13)\), including latent distance, cosine consistency, and weak task guidance\.
After Stage 1, the pretrained backbone initializes its original supervised decoder and is fine\-tuned under the same Stage 2 data split and optimizer policy as the corresponding baseline\. We denote the variants pretrained with task\-guided latent prediction by the prefix TLP, giving TLP\-EEGNet, TLP\-DeepConvNet, TLP\-Conformer, TLP\-Neuro3D, and TLP\-TSConv\. These variants are not treated as additional competitors in the EEG\-VID architecture ranking\. Instead, each TLP variant is compared with its corresponding backbone without Stage 1 to test whether the training strategy transfers\.
To further separate the Stage 1 training principle from the proposed predictor architecture, we replace the predictor with LSTM, GRU, and standard Transformer alternatives for EEG\-VID and all five TLP variants under the pooled protocol\. All replacements use the same latent input/output dimensions, Stage 1 objective, task\-guidance weight, data splits, and training policy, with no predictor\-specific hyperparameter tuning\.
### IV\-FStatistical Analysis
Results from repeated optimization runs are reported as mean±\\pmstandard deviation over five matched seeds\. Seeds quantify optimization variability and are not treated as independent participants\. For the within\-subject and leave\-one\-subject\-out protocols, the primary summary is the mean across the nine subject\-specific evaluations\. Table[III](https://arxiv.org/html/2609.00566#S5.T3)additionally reports the number of subjects whose mean accuracy improves after Stage 1 \(k/9k/9\), exposing the paired direction of the effect without treating those counts as a population\-level significance test\.
We also report the direction of the accuracy change across matched comparisons\. Because these comparisons reuse datasets, subjects, seeds, and related model families, the resulting counts are not treated as independent observations or population\-level significance tests\. Table[V](https://arxiv.org/html/2609.00566#S5.T5)additionally reports the mean and standard deviation of paired accuracy differences across matched model–dataset settings\. For pooled backbone comparisons using matched seeds, Table[II](https://arxiv.org/html/2609.00566#S5.T2)reports the paired change directly\.
### IV\-GImplementation Details
All models are trained with AdamW\[[51](https://arxiv.org/html/2609.00566#bib.bib53)\]on a single NVIDIA RTX 4090 using a batch size of 64\. Training uses validation\-based early stopping with a patience of 100 and a maximum allowance of 10,000 epochs\. For the training\-budget sensitivity control in Section[V\-B](https://arxiv.org/html/2609.00566#S5.SS2), all six corresponding single\-stage models are additionally rerun with a fourfold larger maximum allowance under the same early\-stopping rule\. Weight decay is10−410^\{\-4\}for both stages; learning rates are3×10−43\\times 10^\{\-4\}for Stage 1 and1×10−41\\times 10^\{\-4\}for Stage 2\. Unless changed by an ablation, Stage 1 usesλlatent=1\.0\\lambda\_\{\\mathrm\{latent\}\}=1\.0,λcos=0\.1\\lambda\_\{\\mathrm\{cos\}\}=0\.1,λclsS1=0\.1\\lambda\_\{\\mathrm\{cls\}\}^\{S1\}=0\.1, and EMA decay 0\.995\. Stage 2 usesλcls=1\.0\\lambda\_\{\\mathrm\{cls\}\}=1\.0,λcoord=0\.1\\lambda\_\{\\mathrm\{coord\}\}=0\.1, andλkeep=0\.03\\lambda\_\{\\mathrm\{keep\}\}=0\.03, with cosine coefficient 0\.1 insideℒkeep\\mathcal\{L\}\_\{\\mathrm\{keep\}\}\. The robot target\-selection formulation uses candidate objects segmented from RGB\-D observations and matches them to the decoded spatial region as described in the Methods section\.
## VResults
### V\-APooled Cross\-Session Comparison
TABLE I:Pooled cross\-session comparison\. Best results are shown inbold\.- †\\dagger: Values are mean±\\pmstandard deviation over five fixed random seeds under the shared local pipeline\.
- ‡\\ddagger: IV\-2a and IV\-2b report 1\-s\-window accuracy\.
Table[I](https://arxiv.org/html/2609.00566#S5.T1)compares EEG\-VID with seven locally retrained methods and its no\-Stage 1 ablation\. Under the pooled protocol, EEG\-VID achieves6\.52±0\.95%6\.52\\pm 0\.95\\%/30\.50±1\.79%30\.50\\pm 1\.79\\%Top\-1/Top\-5 on VIG\-48,51\.30±0\.42%51\.30\\pm 0\.42\\%on IV\-2a, and75\.90±1\.23%75\.90\\pm 1\.23\\%on IV\-2b\. Relative to the identical architecture without Stage 1, Top\-1 improves by 1\.47, 11\.30, and 8\.23 points, respectively, isolating the effect of the two\-stage training procedure from decoder architecture\.
EEG2Rep and BENDR provide direct predictive/self\-supervised reference points\. Under the same local evaluation pipeline, both are below EEG\-VID on the three Top\-1 tasks and on VIG\-48 Top\-5\. Because both were designed for larger montages and larger pretraining corpora, this comparison should be read as evidence about behaviour in the present low\-channel, small\-data regime rather than as a general ranking of the methods\.
### V\-BStage 1 Transfer across EEG Backbones
TABLE II:Stage 1 gains across EEG backbones\.- †\\dagger: Entries are seed\-paired accuracy changes in percentage points after Stage 1 relative to the matched backbone, reported as mean±\\pmstandard deviation under the pooled protocol\.
- ‡\\ddagger: EEG\-VID is compared with the identical architecture without Stage 1\.
Table[II](https://arxiv.org/html/2609.00566#S5.T2)addresses a different question by comparing each decoder only with itself before and after Stage 1\. All 15 external\-backbone Top\-1 mean changes are non\-negative, and all five backbones also improve on VIG\-48 Top\-5\. The magnitude of the gain depends on the architecture\. For example, DeepConvNet gains19\.2919\.29points on IV\-2a whereas TSConv changes by only0\.05±1\.160\.05\\pm 1\.16points on IV\-2b\. The latter is best interpreted as no measurable effect\.
Some Stage 1\-enhanced external backbones exceed EEG\-VID in absolute accuracy \(e\.g\., TLP\-Neuro3D on VIG\-48 and TLP\-DeepConvNet on IV\-2a/IV\-2b\)\. This result further supports interpreting Stage 1 as a transferable training strategy rather than an architecture\-specific advantage\. EEG\-VID is an application\-oriented model with a grid\-structured decoder that maps decoded regions to candidate objects\. The gain also varies across architectures, showing that Stage 1 does not provide a fixed improvement across all models and datasets\.
As a training\-budget sensitivity control, all six single\-stage counterparts were rerun with a fourfold larger maximum epoch allowance under the same early\-stopping rule\. Stage 1 remained higher in 23 of 24 pooled model–metric comparisons\. The only exception was TLP\-TSConv on IV\-2b, where the difference was−0\.07\-0\.07percentage points\. Thus, the maximum epoch limit of the single\-stage models does not explain the main pooled gains, although this control does not exactly match the number of parameter updates or total computation\.
### V\-CGeneralization across Subject\-Wise Protocols
TABLE III:Subject\-wise public\-benchmark comparison\. Best results are shown inbold\.- †\\dagger: Entries are accuracy \(%\) followed by \(Δ\\Delta;k/9k/9\), whereΔ\\Deltadenotes the mean change versus the matched non\-Stage 1 backbone andkkthe number of subjects with a positive paired change\. “Within” uses subject\-specific training; “LOSO” uses no labelled data from the evaluated subject\.
The subject\-wise protocols separate absolute model ranking from Stage 1 transfer\. Under within\-subject evaluation, Stage 1 yields positive mean changes for 11 of 12 model–dataset pairs\. The only negative mean is TLP\-TSConv on IV\-2b at−0\.24\-0\.24points\. Thek/9k/9counts in Table[III](https://arxiv.org/html/2609.00566#S5.T3)show that the improvements are observed across multiple participants\. EEG\-VID improves for 8/9 subjects on both IV\-2a and IV\-2b, while TLP\-DeepConvNet and TLP\-Conformer improve for at least 7/9 and 8/9 subjects, respectively\.
The strongest pattern appears under leave\-one\-subject\-out \(LOSO\) evaluation\. Stage 1 improves the mean for all six models on both public datasets\. On IV\-2a, EEG\-VID, TLP\-EEGNet, and TLP\-DeepConvNet improve for all 9/9 held\-out subjects\. TLP\-Conformer and TLP\-Neuro3D improve for 8/9\. The largest mean gain is \+16\.22 points for DeepConvNet, while EEG\-VID improves from 33\.55% to 40\.80%\. On IV\-2b, all six models again have positive mean changes, with subject\-level improvement counts ranging from 6/9 to 8/9\. The LOSO results show that the Stage 1 benefit extends beyond subject\-overlapping cross\-session evaluation\.
Across pooled, within\-subject, and LOSO evaluations, Stage 1 produces a positive mean change in 41 of 42 matched decoder–dataset–protocol comparisons\. This count and thek/9k/9entries are descriptive rather than population\-level significance tests \(Section[IV\-F](https://arxiv.org/html/2609.00566#S4.SS6)\)\.
### V\-DStage 1 Objective Ablation
TABLE IV:Effect of weak task guidance\.- †\\dagger: Entries areAcc0\.1−Acc0\\mathrm\{Acc\}\_\{0\.1\}\-\\mathrm\{Acc\}\_\{0\}in percentage points under the pooled protocol\.
- ‡\\ddagger: Number of datasets for which 0\.1 gives the best of the three nonzero weights\{0\.1,0\.2,0\.3\}\\\{0\.1,0\.2,0\.3\\\}\.
The default weak task\-guidance setting \(λclsS1=0\.1\\lambda\_\{\\mathrm\{cls\}\}^\{S1\}=0\.1\) exceeds latent\-only pretraining in all 18 model–dataset comparisons \(Table[IV](https://arxiv.org/html/2609.00566#S5.T4)\)\. For EEG\-VID, latent\-only pretraining obtains 5\.90%, 48\.47% and 73\.35% on VIG\-48, IV\-2a and IV\-2b, compared with 6\.52%, 51\.30% and 75\.90% for the default\. Latent\-only EEG\-VID also outperforms the no\-Stage 1 model on all three datasets, indicating that latent prediction alone provides consistent gains even without weak task guidance\.
Weights 0\.2 and 0\.3 underperform the default in 34 of 36 comparisons, except for TLP\-DeepConvNet on VIG\-48 \(0\.3\) and TLP\-TSConv on IV\-2a \(0\.2\)\. Thus, a small nonzero task\-guidance weight is the most reliable choice within the tested range\. These results do not show that it is globally optimal\.
### V\-EPredictor Architecture Robustness
TABLE V:Predictor\-architecture robustness\.- †\\dagger: Counts use 18 matched model–dataset settings per predictor replacement under the pooled protocol \(54 overall\); S1 denotes Stage 1 and DS direct supervision\.
- ‡\\ddagger:ΔAcc=Accproposed−Accreplacement\\Delta\\mathrm\{Acc\}=\\mathrm\{Acc\}\_\{\\mathrm\{proposed\}\}\-\\mathrm\{Acc\}\_\{\\mathrm\{replacement\}\}, reported as mean±\\pmstandard deviation in percentage points\.
Table[V](https://arxiv.org/html/2609.00566#S5.T5)tests whether the Stage 1 benefit depends on the proposed predictor\. LSTM, GRU, and Transformer replacements retain an advantage over direct supervision in 48 of 54 pooled model–dataset comparisons, including 18 of 18 on IV\-2a\. Thus, the Stage 1 benefit is not tied to a specific predictor architecture\. The proposed predictor is nevertheless higher in 45 of 54 matched comparisons, with an overall mean advantage of 0\.56±\\pm0\.63 percentage points\. These results separate the transferable Stage 1 training principle from the additional, architecture\-dependent benefit of the proposed predictor\.
### V\-FElectrode\-Subset Analysis
TABLE VI:Electrode\-subset analysis on VIG\-48\.- †\\dagger: Values are mean±\\pmstandard deviation over five fixed random seeds\. Comparisons are descriptive within each model and do not imply source localization\.
For EEG\-VID, the full montage is best, followed by AF7/AF8 and then TP9/TP10\. The same frontal\-over\-temporo\-parietal ordering appears for the five supervised baselines\. Because AF7/AF8 are close to the eyes and gaze was not constrained, this pattern cannot be interpreted as cortical localization\. It only shows that the frontal pair carries stronger standalone target\-related information in the present device configuration\.
### V\-GOffline Robot\-Scene Target Selection
We next evaluate whether the decoded 48\-region posterior can support object\-level selection when scene information restricts the decision space\. For every robot\-camera frame, SAM identifies the four object masks and maps them to four candidate grid labels\. The subject\-specific adapted EEG\-VID model is applied to the held\-out robot\-scene EEG, and its 48\-way scores are restricted to these four SAM\-derived labels using \([19](https://arxiv.org/html/2609.00566#S3.E19)\)\. SAM therefore determines which spatial labels are eligible candidates, whereas EEG\-VID determines which candidate is selected\.
Fig\. 5:Qualitative visualization of offline robot target selection\. SAM segments the four scene objects and maps their positions to candidate grid labels\. EEG\-VID supplies the 48\-region EEG posterior, and the final target is the candidate label with the largest EEG score\.EEG\-VID achieves40\.24%window\-level candidate\-constrained target\-selection accuracy, exceeding the25%random chance level for four candidate objects by 15\.24 percentage points\. Because SAM only defines the four eligible candidate locations and does not infer the intended target, the above\-chance result indicates that the EEG\-derived spatial posterior retains useful target information after scene\-based restriction\.
This result should not be directly compared with the 48\-way VIG\-48 Top\-1 accuracy because the two evaluations use different decision spaces\. The robot\-scene experiment tests whether a calibrated 48\-region EEG posterior can support object selection after being restricted to four scene\-consistent candidates\. It therefore evaluates the complementary roles of visual scene constraints and EEG intention evidence rather than unconstrained 48\-way decoding\.
This experiment evaluates target selection before motion planning and does not measure grasp success\. A substantial gap remains to reliable closed\-loop assistive control\. Fig\.[5](https://arxiv.org/html/2609.00566#S5.F5)illustrates the candidate\-restriction process\.
## VIDiscussion
### VI\-AMain Findings
The main finding is that Stage 1 improves matched decoders across datasets and evaluation protocols, with the clearest gains under subject shift\. The subject\-wise results further show that these gains are broadly distributed rather than driven by a small subset of participants, supporting transfer beyond subject\-overlapping cross\-session evaluation\.
The ablations separate the contribution of the training objective from predictor design\. Latent prediction alone improves over direct supervision, and weak task guidance further strengthens the learned representation\. The benefit also persists with generic temporal predictors, indicating that the transfer effect is not tied to a single predictor architecture\.
Overall, the results identify task\-guided latent prediction as the primary transferable mechanism in the tested setting, while predictor design provides an additional architecture\-dependent gain\. The selected task\-guidance weight and predictor configuration work well in the tested settings but are not necessarily optimal for other datasets or models\.
### VI\-BPhysiological Interpretation and Assistive Relevance
We interpret VIG\-48 at the wearable\-device level rather than as source\-specific cortical decoding\. Natural gaze orienting is permitted, AF7/AF8 are close to the eyes, and eye movements can contribute strongly to scalp EEG during free viewing\[[52](https://arxiv.org/html/2609.00566#bib.bib41),[53](https://arxiv.org/html/2609.00566#bib.bib42),[54](https://arxiv.org/html/2609.00566#bib.bib2),[55](https://arxiv.org/html/2609.00566#bib.bib1)\]\. Without reference EOG or eye tracking, cortical, visually evoked, and ocular contributions cannot be separated\. The electrode\-subset analysis is therefore descriptive and is not used for source\-localization claims\. The transferability of Stage 1 is assessed independently on the motor\-imagery IV\-2a/IV\-2b benchmarks\.
As a stand\-alone 48\-way command channel, 6\.52% Top\-1 accuracy is insufficient for practical control\. Restricting the decision to four SAM\-derived scene candidates yields 40\.24% window\-level target\-selection accuracy versus a 25% chance level\. This supports an offline, calibrated proof of concept, not closed\-loop or clinical readiness\.
### VI\-CLimitations and Future Work
Several limitations define the scope of the current study\. VIG\-48 contains one participant, and the robot\-scene study uses subject\-specific calibration\. Future work will expand the visual\-intention dataset to multiple participants and evaluate cross\-subject transfer with less subject\-specific calibration\. The four\-channel montage and unconstrained gaze do not allow cortical and ocular contributions to be separated\. Future studies will add EOG or eye tracking and denser EEG to better distinguish these sources\. Public\-benchmark results use 1\-s window\-level decisions, so trial\-level aggregation should be evaluated for direct comparison with standard protocols\. Finally, Stage 1 requires additional computation compared with direct supervision\. Future work will use more closely compute\-matched comparisons and closed\-loop robot experiments with online spatial\-belief updating and EEG\-based correction\.
Accordingly, the current results support a transferable pretraining strategy and an offline scene\-informed target\-selection proof of concept rather than a deployable BCI or a source\-specific neural mechanism\.
## VIIConclusion
EEG\-VID combines task\-guided latent predictive pretraining with supervised EEG decoding\. Stage 1 improves matched decoders under cross\-session and cross\-subject shifts and remains effective when the proposed predictor is replaced by generic recurrent or Transformer alternatives\. In the robot\-scene study, restricting the EEG\-derived posterior over grid regions to scene\-derived candidates yields above\-chance window\-level target selection after subject\-specific calibration\. These results support task\-guided latent prediction as a transferable EEG pretraining strategy and motivate future closed\-loop evaluation\.
## References
- \[1\]R\. T\. Schirrmeister, J\. T\. Springenberg, L\. D\. J\. Fiederer, M\. Glasstetter, K\. Eggensperger, M\. Tangermann, F\. Hutter, W\. Burgard, and T\. Ball\(2017\)Deep learning with convolutional neural networks for EEG decoding and visualization\.Human Brain Mapping38\(11\),pp\. 5391–5420\.External Links:[Document](https://dx.doi.org/10.1002/hbm.23730)Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p1.1),[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1),[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.4.1.1)\.
- \[2\]V\. J\. Lawhern, A\. J\. Solon, N\. R\. Waytowich, S\. M\. Gordon, C\. P\. Hung, and B\. J\. Lance\(2018\)EEGNet: a compact convolutional neural network for EEG\-based brain–computer interfaces\.Journal of Neural Engineering15\(5\),pp\. 056013\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p1.1),[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1),[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.3.1.1)\.
- \[3\]Y\. Song, Q\. Zheng, B\. Liu, and X\. Gao\(2023\)EEG Conformer: convolutional transformer for EEG decoding and visualization\.IEEE Transactions on Neural Systems and Rehabilitation Engineering31,pp\. 710–719\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2022.3230250)Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p1.1),[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.5.1.1)\.
- \[4\]C\. Sun and C\. Mou\(2023\)Survey on the research direction of EEG\-based signal processing\.Frontiers in Neuroscience17,pp\. 1203059\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p1.1),[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[5\]C\. Brunner, R\. Leeb, G\. R\. Müller\-Putz, A\. Schlögl, and G\. Pfurtscheller\(2008\)BCI competition 2008—Graz data set A\.Technical reportInstitute for Knowledge Discovery, Graz University of Technology,Graz, Austria\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§IV\-A3](https://arxiv.org/html/2609.00566#S4.SS1.SSS3.p1.1)\.
- \[6\]R\. Leeb, C\. Brunner, G\. R\. Müller\-Putz, A\. Schlögl, and G\. Pfurtscheller\(2008\)BCI competition 2008—Graz data set B\.Technical reportInstitute for Knowledge Discovery, Graz University of Technology,Graz, Austria\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§IV\-A3](https://arxiv.org/html/2609.00566#S4.SS1.SSS3.p1.1)\.
- \[7\]M\. Tangermann, K\. Müller, A\. Aertsen, N\. Birbaumer, C\. Braun, C\. Brunner, R\. Leeb, C\. Mehring, K\. J\. Miller, G\. R\. Müller\-Putz,et al\.\(2012\)Review of the BCI competition IV\.Frontiers in Neuroscience6,pp\. 55\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§IV\-A3](https://arxiv.org/html/2609.00566#S4.SS1.SSS3.p1.1)\.
- \[8\]Y\. Song, B\. Liu, X\. Li, N\. Shi, Y\. Wang, and X\. Gao\(2024\)Decoding natural images from EEG for object recognition\.InInternational conference on learning representations,pp\. 47648–47665\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.7.1.1)\.
- \[9\]Z\. Guo, J\. Wu, Y\. Song, J\. Bu, W\. Mai, Q\. Zheng, W\. Ouyang, and C\. Song\(2025\)Neuro\-3D: towards 3D visual decoding from EEG signals\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 23870–23880\.Cited by:[§I](https://arxiv.org/html/2609.00566#S1.p4.1),[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.6.1.1)\.
- \[10\]N\. Waytowich, V\. J\. Lawhern, J\. O\. Garcia, J\. Cummings, J\. Faller, P\. Sajda, and J\. M\. Vettel\(2018\)Compact convolutional neural networks for classification of asynchronous steady\-state visual evoked potentials\.Journal of Neural Engineering15\(6\),pp\. 066031\.Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[11\]L\. A\. Farwell and E\. Donchin\(1988\)Talking off the top of your head: toward a mental prosthesis utilizing event\-related brain potentials\.Electroencephalography and Clinical Neurophysiology70\(6\),pp\. 510–523\.Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[12\]Y\. Yao, W\. De Swaef, S\. Geirnaert, and A\. Bertrand\(2025\)EEG\-based decoding of selective visual attention in superimposed videos\.IEEE Journal of Biomedical and Health Informatics29\(10\),pp\. 7248–7261\.External Links:[Document](https://dx.doi.org/10.1109/JBHI.2025.3580261)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[13\]A\. T\. Gifford, K\. Dwivedi, G\. Roig, and R\. M\. Cichy\(2022\)A large and rich EEG dataset for modeling human visual object recognition\.NeuroImage264,pp\. 119754\.External Links:[Document](https://dx.doi.org/10.1016/j.neuroimage.2022.119754)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[14\]H\. Wu, Q\. Li, C\. Zhang, Z\. He, and X\. Ying\(2025\)Bridging the vision\-brain gap with an uncertainty\-aware blur prior\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 2246–2257\.Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[15\]K\. M\. Davis, C\. De La Torre\-Ortiz, and T\. Ruotsalo\(2022\)Brain\-supervised image editing\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 18480–18489\.Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[16\]J\. d\. R\. Millán, R\. Rupp, G\. Müller\-Putz, R\. Murray\-Smith, C\. Giugliemma, M\. Tangermann, C\. Vidaurre, F\. Cincotti, A\. Kübler, R\. Leeb, C\. Neuper, K\. R\. Müller, and D\. Mattia\(2010\)Combining brain\-computer interfaces and assistive technologies: state\-of\-the\-art and challenges\.Frontiers in Neuroscience4,pp\. 161\.External Links:[Document](https://dx.doi.org/10.3389/fnins.2010.00161)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[17\]I\. Iturrate, J\. M\. Antelis, A\. Kubler, and J\. Minguez\(2009\)A noninvasive brain\-actuated wheelchair based on a P300 neurophysiological protocol and automated navigation\.IEEE Transactions on Robotics25\(3\),pp\. 614–627\.External Links:[Document](https://dx.doi.org/10.1109/TRO.2009.2020347)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[18\]T\. Carlson and J\. del R\. Millan\(2013\)Brain\-controlled wheelchairs: a robotic architecture\.IEEE Robotics & Automation Magazine20\(1\),pp\. 65–73\.External Links:[Document](https://dx.doi.org/10.1109/MRA.2012.2229936)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[19\]R\. Zhang, S\. Lee, M\. Hwang, A\. Hiranaka, C\. Wang, W\. Ai, J\. J\. R\. Tan, S\. Gupta, Y\. Hao, G\. Levine, R\. Gao, A\. Norcia, L\. Fei\-Fei, and J\. Wu\(2023\)NOIR: neural signal operated intelligent robots for everyday activities\.InProceedings of The 7th Conference on Robot Learning,J\. Tan, M\. Toussaint, and K\. Darvish \(Eds\.\),Proceedings of Machine Learning Research, Vol\.229,pp\. 1737–1760\.External Links:[Link](https://proceedings.mlr.press/v229/zhang23f.html)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[20\]Y\. Xu, C\. Ding, X\. Shu, K\. Gui, Y\. Bezsudnova, X\. Sheng, and D\. Zhang\(2019\)Shared control of a robotic arm using non\-invasive brain–computer interface and computer vision guidance\.Robotics and Autonomous Systems115,pp\. 121–129\.External Links:[Document](https://dx.doi.org/10.1016/j.robot.2019.02.014)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[21\]Y\. Zhou, T\. Yu, W\. Gao, W\. Huang, Z\. Lu, Q\. Huang, and Y\. Li\(2023\)Shared three\-dimensional robotic arm control based on asynchronous BCI and computer vision\.IEEE Transactions on Neural Systems and Rehabilitation Engineering31,pp\. 3163–3175\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2023.3299350)Cited by:[§II\-A](https://arxiv.org/html/2609.00566#S2.SS1.p1.1)\.
- \[22\]X\. Zhou, C\. Liu, J\. Zhou, Z\. Wang, L\. Zhai, Z\. Jia, C\. Guan, and Y\. Liu\(2023\)Interpretable and robust AI in EEG systems: a survey\.arXiv preprint arXiv:2304\.10755\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[23\]V\. Jayaram, M\. Alamgir, Y\. Altun, B\. Scholkopf, and M\. Grosse\-Wentrup\(2016\)Transfer learning in brain\-computer interfaces\.IEEE Computational Intelligence Magazine11\(1\),pp\. 20–31\.External Links:[Document](https://dx.doi.org/10.1109/MCI.2015.2501545)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[24\]P\. Zanini, M\. Congedo, C\. Jutten, S\. Said, and Y\. Berthoumieu\(2018\)Transfer learning: a riemannian geometry framework with applications to brain–computer interfaces\.IEEE Transactions on Biomedical Engineering65\(5\),pp\. 1107–1116\.External Links:[Document](https://dx.doi.org/10.1109/TBME.2017.2742541)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[25\]P\. L\. C\. Rodrigues, C\. Jutten, and M\. Congedo\(2019\)Riemannian procrustes analysis: transfer learning for brain–computer interfaces\.IEEE Transactions on Biomedical Engineering66\(8\),pp\. 2390–2401\.External Links:[Document](https://dx.doi.org/10.1109/TBME.2018.2889705)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[26\]D\. Wu, Y\. Xu, and B\. Lu\(2022\)Transfer learning for EEG\-based brain\-computer interfaces: a review of progress made since 2016\.IEEE Transactions on Cognitive and Developmental Systems14\(1\),pp\. 4–19\.External Links:[Document](https://dx.doi.org/10.1109/TCDS.2020.3007453)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[27\]H\. He and D\. Wu\(2020\)Transfer learning for brain–computer interfaces: a euclidean space data alignment approach\.IEEE Transactions on Biomedical Engineering67\(2\),pp\. 399–410\.External Links:[Document](https://dx.doi.org/10.1109/TBME.2019.2913914)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1),[§IV\-B4](https://arxiv.org/html/2609.00566#S4.SS2.SSS4.p1.1)\.
- \[28\]C\. Flores, M\. Contreras, I\. Macedo, and J\. Andreu\-Perez\(2024\)Transfer learning with active sampling for rapid training and calibration in BCI\-P300 across health states and multi\-centre data\.IEEE Transactions on Neural Systems and Rehabilitation Engineering32,pp\. 3794–3803\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2024.3420960)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[29\]Q\. She, T\. Chen, F\. Fang, J\. Zhang, Y\. Gao, and Y\. Zhang\(2023\)Improved domain adaptation network based on wasserstein distance for motor imagery EEG classification\.IEEE Transactions on Neural Systems and Rehabilitation Engineering31,pp\. 1137–1148\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2023.3241846)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[30\]S\. Sartipi and M\. Cetin\(2024\)Subject\-independent deep architecture for EEG\-based motor imagery classification\.IEEE Transactions on Neural Systems and Rehabilitation Engineering32,pp\. 718–727\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2024.3360194)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[31\]Y\. Zhong, L\. Yao, G\. Pan, and Y\. Wang\(2024\)Cross\-subject motor imagery decoding by transfer learning of tactile ERD\.IEEE Transactions on Neural Systems and Rehabilitation Engineering32,pp\. 662–671\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2024.3358491)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[32\]Y\. Zhou, T\. Luo, X\. Zhang, and T\. Han\(2023\)Spatial feature regularization and label decoupling based cross\-subject motor imagery EEG decoding\.InChinese Conference on Pattern Recognition and Computer Vision \(PRCV\),pp\. 407–423\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[33\]P\. Chen, X\. Liu, C\. Ma, H\. Wang, X\. Yang, C\. Grebogi, X\. Gu, and Z\. Gao\(2025\)Unsupervised domain adaptation with synchronized self\-training for cross\-domain motor imagery recognition\.IEEE Journal of Biomedical and Health Informatics29\(5\),pp\. 3664–3677\.External Links:[Document](https://dx.doi.org/10.1109/JBHI.2025.3525577)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[34\]Y\. Li, L\. Guo, Y\. Liu, J\. Liu, and F\. Meng\(2021\)A temporal\-spectral\-based squeeze\-and\-excitation feature fusion network for motor imagery eeg decoding\.IEEE Transactions on Neural Systems and Rehabilitation Engineering29,pp\. 1534–1545\.External Links:[Document](https://dx.doi.org/10.1109/TNSRE.2021.3099908)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p1.1)\.
- \[35\]M\. N\. Mohsenvand, M\. R\. Izadi, and P\. Maes\(2020\)Contrastive representation learning for electroencephalogram classification\.InProceedings of the Machine Learning for Health NeurIPS Workshop,Proceedings of Machine Learning Research, Vol\.136,pp\. 238–253\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[36\]H\. Banville, O\. Chehab, A\. Hyvärinen, D\. Engemann, and A\. Gramfort\(2021\)Uncovering the structure of clinical EEG signals with self\-supervised learning\.Journal of Neural Engineering18\(4\),pp\. 046020\.External Links:[Document](https://dx.doi.org/10.1088/1741-2552/abca18)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[37\]D\. Kostas, S\. Aroca\-Ouellette, and F\. Rudzicz\(2021\)BENDR: using transformers and a contrastive self\-supervised learning task to learn from massive amounts of EEG data\.Frontiers in Human Neuroscience15,pp\. 653659\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.9.1.1)\.
- \[38\]M\. H\. Rafiei, L\. V\. Gauthier, H\. Adeli, and D\. Takabi\(2024\)Self\-supervised learning for electroencephalography\.IEEE Transactions on Neural Networks and Learning Systems35\(2\),pp\. 1457–1471\.External Links:[Document](https://dx.doi.org/10.1109/TNNLS.2022.3190448)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[39\]N\. Mohammadi Foumani, G\. Mackellar, S\. Ghane, S\. Irtza, N\. Nguyen, and M\. Salehi\(2024\)EEG2Rep: enhancing self\-supervised EEG representation through informative masked inputs\.InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 5544–5555\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1),[§IV\-D](https://arxiv.org/html/2609.00566#S4.SS4.p1.1),[TABLE I](https://arxiv.org/html/2609.00566#S5.T1.4.8.1.1)\.
- \[40\]Z\. Fu, H\. Zhu, Y\. Zhao, R\. Huan, Y\. Zhang, S\. Chen, and Y\. Pan\(2024\)GMAEEG: a self\-supervised graph masked autoencoder for EEG representation learning\.IEEE Journal of Biomedical and Health Informatics28\(11\),pp\. 6486–6497\.External Links:[Document](https://dx.doi.org/10.1109/JBHI.2024.3443651)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[41\]C\. Yang, M\. Westover, and J\. Sun\(2023\)BIOT: biosignal transformer for cross\-data learning in the wild\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Document](https://dx.doi.org/10.52202/075280-3420)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[42\]W\. Jiang, L\. Zhao, and B\. Lu\(2024\)Large brain model for learning generic representations with tremendous EEG data in BCI\.InInternational Conference on Learning Representations,pp\. 16405–16426\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[43\]G\. Wang, W\. Liu, Y\. He, C\. Xu, L\. Ma, and H\. Li\(2024\)EEGPT: pretrained transformer for universal and reliable representation of EEG signals\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 39249–39280\.External Links:[Document](https://dx.doi.org/10.52202/079017-1239)Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[44\]J\. Wang, S\. Zhao, Z\. Luo, Y\. Zhou, H\. Jiang, S\. Li, T\. Li, and G\. Pan\(2025\)CBraMod: a criss\-cross brain foundation model for EEG decoding\.InInternational conference on learning representations,Vol\.2025,pp\. 75310–75346\.Cited by:[§II\-B](https://arxiv.org/html/2609.00566#S2.SS2.p2.1)\.
- \[45\]A\. v\. d\. Oord, Y\. Li, and O\. Vinyals\(2018\)Representation learning with contrastive predictive coding\.arXiv preprint arXiv:1807\.03748\.Cited by:[§II\-C](https://arxiv.org/html/2609.00566#S2.SS3.p1.1)\.
- \[46\]A\. Baevski, W\. Hsu, Q\. Xu, A\. Babu, J\. Gu, and M\. Auli\(2022\)Data2vec: a general framework for self\-supervised learning in speech, vision and language\.InInternational conference on machine learning,Proceedings of Machine Learning Research, Vol\.162,pp\. 1298–1312\.Cited by:[§II\-C](https://arxiv.org/html/2609.00566#S2.SS3.p1.1)\.
- \[47\]J\. Grill, F\. Strub, F\. Altché, C\. Tallec, P\. Richemond, E\. Buchatskaya, C\. Doersch, B\. Avila Pires, Z\. Guo, M\. Gheshlaghi Azar,et al\.\(2020\)Bootstrap your own latent: a new approach to self\-supervised learning\.Advances in neural information processing systems33,pp\. 21271–21284\.Cited by:[§II\-C](https://arxiv.org/html/2609.00566#S2.SS3.p1.1),[§III\-C](https://arxiv.org/html/2609.00566#S3.SS3.p3.1)\.
- \[48\]M\. Assran, Q\. Duval, I\. Misra, P\. Bojanowski, P\. Vincent, M\. Rabbat, Y\. LeCun, and N\. Ballas\(2023\)Self\-supervised learning from images with a joint\-embedding predictive architecture\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 15619–15629\.Cited by:[§II\-C](https://arxiv.org/html/2609.00566#S2.SS3.p1.1),[§III\-C](https://arxiv.org/html/2609.00566#S3.SS3.p3.1)\.
- \[49\]A\. Kirillov, E\. Mintun, N\. Ravi, H\. Mao, C\. Rolland, L\. Gustafson, T\. Xiao, S\. Whitehead, A\. C\. Berg, W\. Lo,et al\.\(2023\)Segment anything\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 4015–4026\.Cited by:[§III\-E](https://arxiv.org/html/2609.00566#S3.SS5.p1.1)\.
- \[50\]H\. Fang, C\. Wang, H\. Fang, M\. Gou, J\. Liu, H\. Yan, W\. Liu, Y\. Xie, and C\. Lu\(2023\)AnyGrasp: robust and efficient grasp perception in spatial and temporal domains\.IEEE Transactions on Robotics39\(5\),pp\. 3929–3945\.External Links:[Document](https://dx.doi.org/10.1109/TRO.2023.3281153)Cited by:[§III\-E](https://arxiv.org/html/2609.00566#S3.SS5.p3.1)\.
- \[51\]I\. Loshchilov and F\. Hutter\(2019\)Decoupled weight decay regularization\.External Links:1711\.05101,[Link](https://arxiv.org/abs/1711.05101)Cited by:[§IV\-G](https://arxiv.org/html/2609.00566#S4.SS7.p1.1)\.
- \[52\]M\. Plöchl, J\. P\. Ossandón, and P\. König\(2012\)Combining EEG and eye tracking: identification, characterization, and correction of eye movement artifacts in electroencephalographic data\.Frontiers in Human Neuroscience6,pp\. 278\.External Links:[Document](https://dx.doi.org/10.3389/fnhum.2012.00278)Cited by:[§VI\-B](https://arxiv.org/html/2609.00566#S6.SS2.p1.1)\.
- \[53\]O\. Dimigen\(2020\)Optimizing the ICA\-based removal of ocular EEG artifacts from free viewing experiments\.NeuroImage207,pp\. 116117\.External Links:[Document](https://dx.doi.org/10.1016/j.neuroimage.2019.116117)Cited by:[§VI\-B](https://arxiv.org/html/2609.00566#S6.SS2.p1.1)\.
- \[54\]C\. Lin, C\. Zhang, J\. Xu, R\. Liu, Y\. Leng, and C\. Fu\(2023\)Neural correlation of EEG and eye movement in natural grasping intention estimation\.IEEE Transactions on Neural Systems and Rehabilitation Engineering31,pp\. 4329–4337\.Cited by:[§VI\-B](https://arxiv.org/html/2609.00566#S6.SS2.p1.1)\.
- \[55\]T\. V\. Afonso and F\. Heinrichs\(2025\)EEG\-EyeTrack: a benchmark for time series and functional data analysis with open challenges and baselines\.arXiv preprint arXiv:2504\.03760\.Cited by:[§VI\-B](https://arxiv.org/html/2609.00566#S6.SS2.p1.1)\.Similar Articles
Interpretable EEG Microstate Discovery via Variational Deep Embedding: A Systematic Architecture Search with Multi-Quadrant Evaluation
This paper presents Conv-VaDE, a variational deep embedding model for interpretable EEG microstate discovery that jointly learns topographic reconstruction and probabilistic soft clustering. It includes a systematic architecture search evaluated on resting-state EEG data to determine optimal model configurations for stability and interpretability.
STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning
Introduces STST-JEPA, a self-supervised transformer for EEG that predicts masked-token representations, pretrained on 47,703 sessions for brain age regression across ages 5–81.
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
ACE-EGO-0 is a unified Vision-Language-Action pretraining framework that leverages egocentric human videos and robot trajectories via a reliability-aware training objective, achieving state-of-the-art on embodied AI benchmarks.
Behavioral Latency as Weak Event-Time Supervision for EEG Reaction-Time Decoding
This paper reformulates EEG-based reaction-time decoding as event-time posterior modeling, using behavioral latency as weak supervision to improve prediction accuracy.
Visual Prompts in Video Models (8 minute read)
Visual prompt engineering (VIPE) automatically modifies task images to improve video model reasoning performance, often more effective than text-based prompting or test-time scaling.