Bio-MF:低延迟高保真EEG到fNIRS跨模态生成用于混合运动想象脑机接口
摘要
Bio-MF是一个一步式生成框架,用于EEG到fNIRS跨模态生成,实现混合运动想象脑机接口的低延迟高保真合成。
arXiv:2609.20904v1 Announce Type: new
Abstract: Hybrid motor-imagery brain-computer interfaces (MI-BCIs) combining EEG and fNIRS can outperform EEG-only systems by exploiting complementary electrophysiological and hemodynamic information. To obtain such hybrid information when paired EEG-fNIRS acquisition is unavailable or inconvenient, recent studies have focused on EEG-to-fNIRS cross-modal generation. However, existing methods still suffer from slow generation and often require pretraining, limiting their use in real-time MI-BCI scenarios. Although one-step generative models offer an attractive route to low-latency synthesis, removing the iterative refinement process can reduce generation fidelity and introduce non-physiological artifacts. To address these problems, this paper proposes Bio-MF, a latent-free one-step MeanFlow framework for EEG-conditioned fNIRS generation. Bio-MF performs direct signal-space x-prediction, converts this signal-space output into MeanFlow velocity supervision, and completes inference with one network evaluation. To preserve task-relevant hemodynamic structure under heterogeneous sensor layouts, Bio-MF integrates Spatial-Temporal Interactive 4D Encoding, cross-modal classifier-free guidance, and noise-level-gated FFT regularization. On Dataset 1, EEG + synthetic fNIRS improves ACC over EEG-only by 3.37 and 4.15 percentage points for HbR and HbO, respectively. On Dataset 2, the corresponding gains remain 2.98 and 2.50 percentage points under the unseen 64-channel EEG montage. On an RTX PRO 6000 GPU, Bio-MF generates one fNIRS trial in 7.0 ms, corresponding to an 857x speedup over the 1000-step SCDM latency. These results show that Bio-MF enables fast EEG-to-fNIRS synthesis while preserving task-relevant generation quality for downstream hybrid MI decoding. Our code is available at https://github.com/psychosiwa/Bio-MF.
查看缓存全文
缓存时间: 2026/09/21 09:12
# Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain–Computer Interfaces
Source: [https://arxiv.org/html/2609.20904](https://arxiv.org/html/2609.20904)
Sifan ZhangLuping Chen††thanks:Boyuan Zhao and Luping Chen are with the Key Laboratory of Modern Teaching Technology, Shaanxi Normal University, Xi’an 710062, China \(e\-mail: 20243259@snnu\.edu\.cn; 2023301681@snnu\.edu\.cn\)\.††thanks:Sifan Zhang is with Microsoft, Beijing, China \(e\-mail: sifanzhang@microsoft\.com\)\.††thanks:Corresponding author: Boyuan Zhao\.
###### Abstract
Hybrid motor\-imagery brain\-computer interfaces \(MI\-BCIs\) combining EEG and fNIRS can outperform EEG\-only systems by exploiting complementary electrophysiological and hemodynamic information\. To obtain such hybrid information when paired EEG\-fNIRS acquisition is unavailable or inconvenient, recent studies have focused on EEG\-to\-fNIRS cross\-modal generation\. However, existing methods still suffer from slow generation and often require pretraining, limiting their use in hybrid MI\-BCI scenarios\. Although one\-step generative models offer an attractive route to low\-latency synthesis, removing the iterative refinement process can reduce generation fidelity and introduce non\-physiological artifacts\. To address these problems, this paper proposes Bio\-MF, a latent\-free one\-step MeanFlow framework for EEG\-conditioned fNIRS generation\. Bio\-MF directly predicts clean fNIRS signals rather than noise in the raw fNIRS space, converts this clean\-signal output into MeanFlow velocity supervision, and completes inference with one network evaluation\. To preserve task\-relevant hemodynamic structure under heterogeneous sensor layouts, Bio\-MF integrates Spatial\-Temporal Interactive 4D Encoding, cross\-modal classifier\-free guidance, and noise\-level\-gated FFT regularization\. On Dataset 1, EEG \+ synthetic fNIRS improves ACC over EEG\-only by 3\.37 and 4\.15 percentage points for HbR and HbO, respectively\. On Dataset 2, the corresponding gains remain 2\.98 and 2\.50 percentage points under the unseen 64\-channel EEG montage\. On an RTX PRO 6000 GPU, Bio\-MF generates one fNIRS trial in 7\.0 ms, corresponding to an 857x speedup over the 1000\-step SCDM latency\. These results show that Bio\-MF enables fast EEG\-to\-fNIRS synthesis while preserving task\-relevant generation quality for hybrid MI\-BCIs\. Our code is available at[https://github\.com/psychosiwa/Bio\-MF](https://github.com/psychosiwa/Bio-MF)\.
###### Index Terms:
Brain\-computer interface, EEG\-to\-fNIRS generation, MeanFlow, multimodal neuroimaging, motor imagery, single\-step generation
## IIntroduction
Motor imagery brain\-computer interfaces \(MI\-BCIs\) decode sensorimotor activity induced by imagined limb movements and provide a non\-muscular control pathway for rehabilitation, assistive control, and human\-machine interaction \[1\]–\[3\]\. Electroencephalography \(EEG\) is widely used in MI\-BCIs because it is non\-invasive, portable, relatively inexpensive, and has millisecond\-level temporal resolution\. Most EEG\-based MI\-BCI methods therefore focus on extracting discriminative sensorimotor rhythm features from EEG signals, using spatial filtering, time\-frequency analysis, or neural classifiers to decode left\- and right\-hand motor imagery \[4\]–\[7\]\.
Despite these advantages, EEG\-only MI decoding remains challenging\. EEG signals have a low signal\-to\-noise ratio and are affected by ocular artifacts, muscle activity, volume conduction, inter\-subject variability, and session\-to\-session non\-stationarity \[4\]–\[7\]\. These factors are particularly relevant for single\-trial MI decoding, where event\-related desynchronization and synchronization patterns can be weak and variable across users and recording sessions\. In addition, EEG has limited spatial localization, which restricts its ability to characterize local cortical activation associated with motor imagery\. These limitations have motivated multimodal BCI studies that combine EEG with complementary neural or physiological signals\.
Functional near\-infrared spectroscopy \(fNIRS\) is a complementary modality\. It records task\-related changes in oxygenated and deoxygenated hemoglobin \[8\]–\[11\]\. Although fNIRS has lower temporal resolution than EEG because of the delayed hemodynamic response, it can provide spatial and metabolic information related to cortical activation\. Hybrid EEG\-fNIRS systems can therefore use EEG to capture fast electrophysiological dynamics and fNIRS to provide complementary hemodynamic information\. Previous studies have reported that EEG\-fNIRS fusion can improve MI decoding performance and stability compared with EEG\-only systems in several settings \[12\]–\[18\]\.
However, simultaneous EEG\-fNIRS acquisition is less convenient than EEG\-only recording\. EEG electrodes and fNIRS sources/detectors must share the scalp surface, which leads to sensor\-layout conflicts and increases preparation time, optical coupling requirements, and sensitivity to hair obstruction and source\-detector spacing \[19\]–\[23\]\. The hybrid setup can also be inconvenient for repeated calibration, portable BCI use, and feedback\-based settings where preparation time and user comfort are important\. These constraints motivate EEG\-conditioned synthesis of task\-related fNIRS signals for hybrid MI decoding when only EEG is available\.
Cross\-modal generation has also been extended to neural\-signal synthesis, including EEG\-to\-fMRI generation \[24\]\. For EEG\-to\-fNIRS synthesis, SCDM models spatial and multi\-scale temporal relations \[25\], whereas TADM combines unified pretraining with latent diffusion to improve transfer across tasks and sensor configurations \[26\]\. These studies establish EEG\-conditioned generation as a route to auxiliary representations for BCI decoding\.
Despite these advances, existing EEG\-to\-fNIRS generation methods still have limitations for latency\-sensitive hybrid MI\-BCI use\. SCDM uses an iterative diffusion paradigm with a U\-Net\-like denoising network, spatial cross\-modal generation modules, multi\-scale temporal representation modules, and precomputed correlation matrices to guide EEG\-to\-fNIRS mapping \[25\]\. This design supports EEG\-to\-fNIRS synthesis, but diffusion inference requires a long reverse denoising chain, which increases generation latency\. In addition, precomputed correlation matrices and projection\-based spatial representations may make the model dependent on a specific sensor layout, limiting adaptation to different EEG montages or fNIRS geometries\. TADM reduces part of the modeling burden by using unified pretraining and latent diffusion \[26\]\. However, unified pretraining can increase training and deployment cost, while VAE latent compression introduces an additional encoder\-decoder reconstruction path, may attenuate subtle HbO/HbR variations, and still retains iterative sampling latency because generation is performed in a latent diffusion process\. Moreover, convolutional U\-Net architectures are effective for regular\-grid signals but usually require input channels to be arranged on a fixed grid or mapped through predefined spatial transformations\. Such limitations can make the model dependent on a specific montage and less flexible when the number or physical placement of sensors changes\. In contrast, a token\-based Transformer formulation can represent EEG and fNIRS patches together with their physical coordinates and auxiliary controls in a shared sequence, allowing the model to process heterogeneous sampling rates and non\-uniform sensor layouts more naturally \[32\]–\[34\]\.
To address the above issues, we propose Bio\-MF, a Transformer\-based latent\-free single\-step MeanFlow framework for EEG\-conditioned fNIRS generation in hybrid MI\-BCIs\. Bio\-MF replaces multi\-step diffusion sampling with a flow\-matching one\-step generation paradigm\. MeanFlow learns an interval\-average velocity field that enables 1\-NFE generation without progressive distillation or a long reverse denoising trajectory \[30\], while Pixel Mean Flows show that latent\-free signal\-space outputs can be combined with velocity\-space supervision \[31\]\. Following this idea, Bio\-MF directly predicts the clean fNIRS signal rather than a noise residual in the raw fNIRS signal space and converts this clean\-signal output into MeanFlow velocity supervision\. This avoids VAE compression and reconstruction, keeps the generated output in the physiological signal domain, and removes both latent decoding and iterative denoising from the inference path\. Because the network output remains an fNIRS signal, physiological and spectral constraints can be applied directly to the generated content, making it possible to optimize signal quality while retaining one\-step inference speed\. To further mitigate the fidelity loss caused by removing iterative refinement, Bio\-MF uses a two\-stage noise\-level\-gated FFT regularizer that applies spectral alignment only in low\-noise states, thereby reducing frequency\-domain artifacts while avoiding conflicts with MeanFlow velocity learning\.
Bio\-MF further combines geometry\-aware token encoding with controllable cross\-modal guidance to improve EEG\-to\-fNIRS generation under heterogeneous sensor layouts\. First, inspired by REVE \[36\], Bio\-MF represents EEG electrodes and fNIRS channels using continuous four\-dimensional coordinates\(x,y,z,t\)\(x,y,z,t\), where\(x,y,z\)\(x,y,z\)encodes the physical sensor location andttencodes the patch\-center physical time\. To further capture location\-dependent temporal structure, we introduce Spatial\-Temporal Interactive 4D Encoding \(Interactive 4D\), which extends the original 4D space\-time encoding with a multiplicative interaction between spatial and temporal features\. This design aligns EEG and fNIRS tokens by physical time and provides the model with location\-dependent temporal information, rather than relying on fixed channel indices tied to a specific montage\. Second, Bio\-MF adopts cross\-modal classifier\-free guidance \(CFG\) \[37\] to balance EEG\-conditioned responses with the learned fNIRS prior\. Since EEG conditions can be noisy, overly strong conditioning may transfer non\-physiological EEG perturbations into fNIRS, whereas overly weak conditioning may lead to an average hemodynamic template\. CFG provides a controllable mechanism for combining conditional trial information with the marginal fNIRS prior\.
Experiments on paired EEG\-fNIRS Dataset 1 and EEG\-only Dataset 2 indicate that Bio\-MF can generate task\-related fNIRS signals for downstream hybrid MI decoding\. Compared with SCDM, Bio\-MF obtains comparable generated\-modality decoding performance in most evaluated settings while reducing generation latency from seconds to milliseconds\. On an RTX PRO 6000 GPU, Bio\-MF generates one fNIRS trial in 7\.0 ms, corresponding to an 857x speedup over the hardware\-matched 1000\-step SCDM baseline\. These results suggest that the Bio\-MF framework provides a feasible path toward fast and high\-quality EEG\-to\-fNIRS signal generation for hybrid MI\-BCIs\.
The main contributions are summarized as follows:
1. 1\.We propose Bio\-MF, a Transformer\-based latent\-free single\-step EEG\-to\-fNIRS generation framework for hybrid MI\-BCIs\. Bio\-MF directly predicts clean fNIRS signals rather than noise in the raw fNIRS signal space and uses MeanFlow velocity supervision to learn the interval\-average denoising direction from a noisy fNIRS state toward the clean fNIRS signal, enabling 1\-NFE inference without VAE compression or multi\-step reverse diffusion\.
2. 2\.We introduce Spatial\-Temporal Interactive 4D Encoding and cross\-modal CFG for EEG\-conditioned fNIRS generation\. Interactive 4D uses continuous coordinates for EEG electrodes, fNIRS channels, and physical time while modeling spatial\-temporal interactions\. CFG controls the balance between EEG\-conditioned responses and the fNIRS prior\.
3. 3\.We use a two\-stage noise\-level\-gated FFT regularizer to address the speed\-fidelity trade\-off of one\-step generation\. By applying spectral alignment only in low\-noise states, the training strategy is designed to reduce frequency\-domain artifacts while keeping the millisecond\-scale inference path\.
4. 4\.We evaluate Bio\-MF using downstream classification, cross\-device zero\-shot evaluation, signal\-level visualizations, ablation studies, and measured generation latency\. The results suggest that Bio\-MF can provide task\-useful synthetic fNIRS while reducing EEG\-to\-fNIRS generation latency compared with multi\-step diffusion baselines, supporting hybrid MI\-BCIs in rehabilitation training, assistive control, and feedback\-based motor\-imagery practice\.
## IIDatasets
### II\-ADataset 1
Dataset 1 is a publicly available hybrid motor imagery dataset introduced by Shin et al\. \[27\], containing temporally aligned EEG and fNIRS recordings\. This dataset was used as the paired source domain for training the proposed EEG\-to\-fNIRS generator and for evaluating in\-distribution downstream classification performance\. It includes 29 healthy participants performing left\- and right\-hand motor imagery \(LMI/RMI\) tasks\. Each participant completed three sessions, with 20 trials per session, resulting in 1740 valid MI trials\. Each trial consisted of a preparation period, a task execution interval, and a subsequent resting period, allowing the recorded fNIRS signals to cover the delayed hemodynamic response induced by motor imagery\.
EEG was recorded from 30 active electrodes arranged according to the international 10\-5 system\. The EEG signals were resampled to 160 Hz, band\-pass filtered between 0\.5 and 50 Hz using a fourth\-order Chebyshev type\-II filter, and further processed with independent component analysis \(ICA\) to reduce ocular artifacts\. fNIRS was acquired using 14 light sources and 16 detectors, forming 36 measurement channels\. The fNIRS signals were downsampled to 10 Hz, filtered with a sixth\-order zero\-phase Butterworth band\-pass filter within 0\.01\-0\.1 Hz, and converted into oxy\-hemoglobin \(HbO\) and deoxy\-hemoglobin \(HbR\) concentration changes\. Before model input, EEG and fNIRS tensors were z\-score normalized to reduce subject\-, channel\-, and session\-level scale differences\. For model training and evaluation, each EEG trial was represented as a30×400030\\times 4000sequence, and each fNIRS trial was represented as a two\-channel tensor of size2×36×2562\\times 36\\times 256, where the two channels correspond to HbO and HbR\.
### II\-BDataset 2
Dataset 2 was introduced to assess whether Bio\-MF can generalize to an unseen EEG acquisition setting without paired fNIRS supervision\. This dataset was derived from a PhysioNet motor imagery EEG resource recorded with BCI2000 \[28\], \[29\] and contains only EEG recordings\. Therefore, it was not used for generator training or fNIRS reconstruction supervision\. Instead, it provides a strict zero\-shot evaluation scenario in which the trained generator receives EEG from a different device layout and produces synthetic fNIRS signals for downstream hybrid MI\-BCI classification\.
Following the SCDM evaluation protocol \[25\], we selected data from 100 subjects performing LMI/RMI tasks\. Each subject contributed 42 trials with a balanced LMI/RMI ratio, yielding 4200 samples\. Unlike Dataset 1, which uses a 30\-electrode EEG montage based on the international 10\-5 system, Dataset 2 uses a 64\-channel high\-density EEG montage following the international 10\-10 system\. The EEG signals were sampled at 160 Hz, filtered within 8\-30 Hz using a finite impulse response \(FIR\) band\-pass filter, and z\-score normalized before zero\-shot fNIRS generation and downstream classification\. For Dataset 2, Bio\-MF uses the 64\-channel EEG coordinates as the condition tokens and generates synthetic HbO/HbR fNIRS on the 36\-channel fNIRS layout learned from Dataset 1\.
## IIIMethodology
Fig\. 1:Overall architecture of Bio\-MF\. \(a\) One\-step fNIRS generation from the sampled noise endpoint and EEG condition\. \(b\) Multimodal tokenization and shared Transformer blocks\. \(c\) Transformer block structure\. \(d\) Dual\-branch MeanFlow velocity supervision and noise\-level\-gated FFT regularization\. \(e\) CFG velocity target construction using the same shared\-stem structure as in \(b\)\.### III\-AOverall Architecture of the Proposed Framework
As shown in Fig\. 1, Bio\-MF uses a tokenized Transformer backbone shared by training and one\-step inference\. EEG and noised fNIRS are divided into physically aligned multi\-rate patches and augmented with continuous sensor\-time coordinates and auxiliary control tokens\. During training, a shared stem and separateuu\- andvv\-heads are optimized with MeanFlow velocity supervision, CFG target construction, and noise\-level\-gated spectral regularization\. Theuu\-head predicts a clean\-signal field that is converted to MeanFlow velocity, while the training\-onlyvv\-head supplies auxiliary velocity supervision and CFG targets \[30\], \[31\], \[37\]\. At inference,t=1t=1andr=0r=0, and the shared stem plusuu\-head generate the final fNIRS signal with one network evaluation\.
Each Transformer block contains RMSNorm, multi\-head self\-attention, SwiGLU, and vector\-gated residual connections \[32\], \[38\], \[39\]\. Interactive 4D encoding represents sensor geometry and physical patch time \[35\], \[36\], while the second training stage applies a gated FFT loss to constrain low\-frequency fNIRS morphology\. The following subsections define these components\.
### III\-BSpatial\-Temporal Interactive 4D Encoding and Multi\-Rate Tokenization
Conventional EEG deep learning studies often treat electrode channels as a one\-dimensional array or project them onto a two\-dimensional planar grid, such as a16×1616\\times 16matrix\. These channel representations can distort the native anatomical geometry of the brain and disrupt true spatial distances between recording sites, making models less robust when electrode layouts change, such as from 30 to 64 channels\. The REVE EEG foundation model shows that jointly encoding three\-dimensional electrode coordinates and time indices as 4D positional encoding can reduce fixed\-montage dependence and improve generalization to unseen recording setups \[36\]\. Inspired by this idea, Bio\-MF extends the 4D representation from EEG representation learning to EEG\-to\-fNIRS cross\-modal generation\.
To avoid such geometric distortion, Bio\-MF discards discrete channel indices and instead represents EEG and fNIRS measurements with continuous spatial coordinates and physical time\. Because EEG and fNIRS have substantially different sampling rates, we first apply a multi\-rate tokenization strategy: EEG signals at 160 Hz and fNIRS signals at 10 Hz are divided according to matched physical time windows, with fNIRS represented by 0\.8\-s temporal patches\. EEG tokens are partitioned using the corresponding physical intervals, and each EEG/fNIRS token is assigned its physical patch\-center timestamp\. The two modalities are thereby aligned by physical time rather than by raw sample indices and represented as unified sequential features\.
Before entering the denoising network, EEG tokens and noisy fNIRS tokens each incorporate the proposed Interactive 4D positional encoding\. The guidance scaleω\\omega, guidance interval \[tmint\_\{\\min\},tmaxt\_\{\\max\}\], and MeanFlow time pair\(t,r\)\(t,r\)are embedded as auxiliary control tokens\. The three types of tokens are then concatenated and fed into the Transformer denoising network\.
After patching, each token is assigned an explicit physical coordinate vector
𝒄i=\(xi,yi,zi,ti\),\\boldsymbol\{c\}\_\{i\}=\(x\_\{i\},y\_\{i\},z\_\{i\},t\_\{i\}\),\(1\)
where\(xi,yi,zi\)\(x\_\{i\},y\_\{i\},z\_\{i\}\)denotes the three\-dimensional Euclidean position of the corresponding electrode or optical measurement channel in a standard head model, andtit\_\{i\}denotes the physical center timestamp of the temporal patch in seconds\.
The original 4D encoding was introduced for EEG\-centric representation learning, where tokens are mainly organized within a single electrophysiological modality\. It represents each token using a joint\(x,y,z,t\)\(x,y,z,t\)coordinate, which provides continuous spatial\-temporal information but does not explicitly separate spatial and temporal factors or parameterize their multiplicative interaction\. For EEG\-to\-fNIRS generation, however, the model must bridge two heterogeneous modalities with different sensor layouts, sampling rates, and neurophysiological dynamics\. Therefore, the original joint 4D design may be suboptimal for representing cross\-modal spatial\-temporal coupling between EEG conditions and fNIRS responses\. To address this limitation, Interactive 4D decouples spatial and temporal Fourier features and introduces an explicit interaction term, making the coordinate representation better suited to multimodal EEG\-to\-fNIRS generation\.
Following the general design principle of REVE \[36\], the positional representation contains a structured Fourier branch and a lightweight learnable coordinate branch, followed by normalization after fusion\. In the Fourier branch, we decouple the spatial coordinates and the temporal coordinate before introducing the explicit spatial\-temporal interaction term\. Let𝒑i=\(xi,yi,zi\)\\boldsymbol\{p\}\_\{i\}=\(x\_\{i\},y\_\{i\},z\_\{i\}\)be the spatial coordinate\. We first compute separate spatial and temporal Fourier features:
𝒔i=Φs\(𝒑i\),𝒒i=Φt\(ti\),\\boldsymbol\{s\}\_\{i\}=\\Phi\_\{s\}\(\\boldsymbol\{p\}\_\{i\}\),\\quad\\boldsymbol\{q\}\_\{i\}=\\Phi\_\{t\}\(t\_\{i\}\),\(2\)
whereΦs\(⋅\)\\Phi\_\{s\}\(\\cdot\)jointly encodes the three\-dimensional sensor position using spatial Fourier bases, andΦt\(⋅\)\\Phi\_\{t\}\(\\cdot\)encodes the physical patch\-center time\. The Fourier mapping follows the standard sinusoidal form\. For a scalar coordinate projectionccand frequencyfkf\_\{k\}, one basis component is
ψk\(c\)=\[cos\(cfk2πW\),sin\(cfk2πW\)\],\\psi\_\{k\}\(c\)=\\left\[\\cos\\left\(cf\_\{k\}\\frac\{2\\pi\}\{W\}\\right\),\\sin\\left\(cf\_\{k\}\\frac\{2\\pi\}\{W\}\\right\)\\right\],\(3\)
whereWWis a normalization constant for the coordinate range\. In the spatial branch,ccis obtained by projecting𝒑i\\boldsymbol\{p\}\_\{i\}onto a spatial Fourier direction; in the temporal branch,c=tic=t\_\{i\}\.
To model location\-dependent temporal effects, Interactive 4D introduces a multiplicative interaction between the spatial and temporal Fourier features:
𝒓i=\(Ws𝒔i\)⊙\(Wt𝒒i\),\\boldsymbol\{r\}\_\{i\}=\(W\_\{s\}\\boldsymbol\{s\}\_\{i\}\)\\odot\(W\_\{t\}\\boldsymbol\{q\}\_\{i\}\),\(4\)
whereWsW\_\{s\}andWtW\_\{t\}are learnable linear projections and⊙\\odotdenotes element\-wise multiplication\. The Fourier\-based positional branch is then defined as
𝑭pe,i=𝒔i\+𝒒i\+𝒓i\.\\boldsymbol\{F\}\_\{pe,i\}=\\boldsymbol\{s\}\_\{i\}\+\\boldsymbol\{q\}\_\{i\}\+\\boldsymbol\{r\}\_\{i\}\.\(5\)
Here,𝑭pe,i\\boldsymbol\{F\}\_\{pe,i\}is not a second Fourier transform; it is the fused output of the spatial Fourier, temporal Fourier, and interactive Fourier\-based features\.
Following the REVE positional\-fusion design \[36\], we also include a learnable coordinate branch to adapt the continuous coordinates to the target generation setting:
𝑭lin,i=LayerNorm\(GELU\(Wl𝒄i\+bl\)\)\.\\boldsymbol\{F\}\_\{lin,i\}=\\mathrm\{LayerNorm\}\\left\(\\mathrm\{GELU\}\\left\(W\_\{l\}\\boldsymbol\{c\}\_\{i\}\+b\_\{l\}\\right\)\\right\)\.\(6\)
The final 4D positional representation is obtained by fusing the Fourier\-based branch and the learnable branch:
𝒑i4D=LayerNorm\(𝑭pe,i\+𝑭lin,i\)\.\\boldsymbol\{p\}^\{4D\}\_\{i\}=\\mathrm\{LayerNorm\}\\left\(\\boldsymbol\{F\}\_\{pe,i\}\+\\boldsymbol\{F\}\_\{lin,i\}\\right\)\.\(7\)
Finally, the positional representation is added to the corresponding signal token:
𝒉i=𝒉isignal\+𝒑i4D\.\\boldsymbol\{h\}\_\{i\}=\\boldsymbol\{h\}\_\{i\}^\{signal\}\+\\boldsymbol\{p\}^\{4D\}\_\{i\}\.\(8\)
### III\-CLatent\-Free MeanFlow for 1\-Step Generation
Bio\-MF does not use a VAE encoder/decoder and does not primarily output a noise residual\. Instead, it directly predicts the clean fNIRS signal in the raw fNIRS signal space, and then converts this signal\-form output into velocity space for supervision\. This design avoids reconstruction bias introduced by latent compression and keeps the model output interpretable as an fNIRS signal\.
Unlikeϵ\\epsilon\-prediction or score\-prediction, direct clean\-signal prediction parameterizes the target data itself, so the network output remains in an interpretable physiological signal space\. This design allows structural and task\-related constraints to be imposed directly on the generated content, including channel consistency, temporal smoothness, frequency\-domain consistency, or prior mask constraints, without indirectly mapping these constraints into noise or score space\. For EEG\-to\-fNIRS generation, where the synthetic signal should preserve clear physiological and task semantics, direct signal prediction therefore provides a more natural way to jointly optimize generation quality and one\-step inference speed\.
During training, given a real fNIRS signal𝐱0\\mathbf\{x\}\_\{0\}and Gaussian noiseϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\), the forward interpolation state is defined as
𝐳t=\(1−t\)𝐱0\+tϵ,t∈\[0,1\]\.\\mathbf\{z\}\_\{t\}=\(1\-t\)\\mathbf\{x\}\_\{0\}\+t\\boldsymbol\{\\epsilon\},\\qquad t\\in\[0,1\]\.\(9\)
Here,𝐱0∼pdata\\mathbf\{x\}\_\{0\}\\sim p\_\{\\mathrm\{data\}\}andϵ∼pprior\\boldsymbol\{\\epsilon\}\\sim p\_\{\\mathrm\{prior\}\}; thus𝐳0=𝐱0∼pdata\\mathbf\{z\}\_\{0\}=\\mathbf\{x\}\_\{0\}\\sim p\_\{\\mathrm\{data\}\}and𝐳1=ϵ∼pprior\\mathbf\{z\}\_\{1\}=\\boldsymbol\{\\epsilon\}\\sim p\_\{\\mathrm\{prior\}\}\. MeanFlow does not integrate along the full ODE trajectory step by step\. Instead, it learns the interval average velocity𝐮\\mathbf\{u\}from the current timettto the target timerr\. The velocity expressions below are evaluated fort\>0t\>0in practice, withttsampled away from zero to avoid the degenerate endpoint\.
The Bio\-MFuu\-head first outputs the signal\-form field𝐱u\\mathbf\{x\}\_\{u\}, where𝐱u\\mathbf\{x\}\_\{u\}denotes the network\-generated generalized clean\-signal field induced by theuu\-head under the current\(𝐳t,t,r\)\(\\mathbf\{z\}\_\{t\},t,r\)condition\. Whenr=0r=0, this generalized output corresponds to the clean endpoint signal𝐱0\\mathbf\{x\}\_\{0\}\. This signal\-form output is then converted into average velocity:
𝐮=𝐳t−𝐱ut\.\\mathbf\{u\}=\\frac\{\\mathbf\{z\}\_\{t\}\-\\mathbf\{x\}\_\{u\}\}\{t\}\.\(10\)
Because𝐮\\mathbf\{u\}represents the network\-output interval average velocity while the training target is defined in velocity space, Bio\-MF uses the MeanFlow identity to construct a supervised composite velocity \[31\]:
𝐕=𝐮\+\(t−r\)⋅𝐉𝐕𝐏sg\.\\mathbf\{V\}=\\mathbf\{u\}\+\(t\-r\)\\cdot\\mathbf\{JVP\}\_\{\\mathrm\{sg\}\}\.\(11\)
where JVP denotes the Jacobian\-vector product used to computed𝐮/dtd\\mathbf\{u\}/dt, and “sg” denotes stop\-gradient;𝐕\\mathbf\{V\}is the MeanFlow consistency velocity used to align the velocity target during training, rather than a new network output\. The guided velocity target𝐯g\\mathbf\{v\}\_\{g\}is defined by the cross\-modal guidance mechanism in the next subsection\.
### III\-DCross\-Modal CFG
Bio\-MF uses cross\-modal CFG as a statistical conditioning mechanism rather than a causal model of neurovascular coupling\. During training, the auxiliaryvv\-head evaluates real and zero EEG under the zero\-interval settingr=tr=tto obtain conditional and unconditional velocities:
𝐯c=fθv\(𝐳t,t,t,ω,0,1,𝐲\),𝐯u=fθv\(𝐳t,t,t,1,0,1,𝟎\)\.\\displaystyle\\mathbf\{v\}\_\{c\}=f\_\{\\theta\}^\{v\}\(\\mathbf\{z\}\_\{t\},t,t,\\omega,0,1,\\mathbf\{y\}\),\\quad\\mathbf\{v\}\_\{u\}=f\_\{\\theta\}^\{v\}\(\\mathbf\{z\}\_\{t\},t,t,1,0,1,\\mathbf\{0\}\)\.
\(12\)
Here,𝐯c\\mathbf\{v\}\_\{c\}is the EEG\-conditioned direction and𝐯u\\mathbf\{v\}\_\{u\}is the marginal fNIRS\-prior direction\. Given the base velocity
𝐯t=𝐳t−𝐱0t,\\mathbf\{v\}\_\{t\}=\\frac\{\\mathbf\{z\}\_\{t\}\-\\mathbf\{x\}\_\{0\}\}\{t\},\(13\)
the guided velocity target is defined as
𝐯g=𝐯t\+\(1−1ωint\)\(𝐯c−𝐯u\)\.\\mathbf\{v\}\_\{g\}=\\mathbf\{v\}\_\{t\}\+\\left\(1\-\\frac\{1\}\{\\omega\_\{int\}\}\\right\)\(\\mathbf\{v\}\_\{c\}\-\\mathbf\{v\}\_\{u\}\)\.\(14\)
whereωint=ω\\omega\_\{int\}=\\omegafort∈\[tmin,tmax\]t\\in\[t\_\{\\min\},t\_\{\\max\}\]andωint=1\\omega\_\{int\}=1otherwise\. The offset𝐯c−𝐯u\\mathbf\{v\}\_\{c\}\-\\mathbf\{v\}\_\{u\}captures EEG conditioning relative to the learned fNIRS prior, not a causal hemoglobin increment\. Theuu\-head learns the JVP\-corrected average velocity𝐕\\mathbf\{V\}, whereas thevv\-head directly regresses the same target\. Their per\-sample errors are
ℓu=‖𝐕−𝐯g‖22,ℓv=‖𝐯−𝐯g‖22\.\\ell\_\{u\}=\\left\\\|\\mathbf\{V\}\-\\mathbf\{v\}\_\{g\}\\right\\\|\_\{2\}^\{2\},\\qquad\\ell\_\{v\}=\\left\\\|\\mathbf\{v\}\-\\mathbf\{v\}\_\{g\}\\right\\\|\_\{2\}^\{2\}\.\(15\)
The MeanFlow loss combines both velocity errors:
ℒMF=𝔼\[ℓu\+ℓv\]\.\\mathcal\{L\}\_\{MF\}=\\mathbb\{E\}\\left\[\\ell\_\{u\}\+\\ell\_\{v\}\\right\]\.\(16\)
Thevv\-head and CFG target construction are omitted at inference\.
### III\-ETwo\-Stage Optimization with Noise\-Level\-Gated Spectral Regularization
To suppress high\-frequency artifacts without disturbing high\-noise velocity learning, Stage I optimizes onlyℒMF\\mathcal\{L\}\_\{MF\}, whereas Stage II forms𝐱u=𝐳t−t𝐮\\mathbf\{x\}\_\{u\}=\\mathbf\{z\}\_\{t\}\-t\\mathbf\{u\}and applies FFT regularization only fort<0\.4t<0\.4:
ℒfft=‖\|ℱt\(𝐱u\)\|−\|ℱt\(𝐱0\)\|‖22\.\\mathcal\{L\}\_\{fft\}=\\left\\\|\\left\|\\mathcal\{F\}\_\{t\}\(\\mathbf\{x\}\_\{u\}\)\\right\|\-\\left\|\\mathcal\{F\}\_\{t\}\(\\mathbf\{x\}\_\{0\}\)\\right\|\\right\\\|\_\{2\}^\{2\}\.\(17\)
Here,ℱt\\mathcal\{F\}\_\{t\}is a one\-dimensional FFT along time, and the norm spans HbO/HbR, spatial channels, and frequencies\. The Stage II objective is
ℒtotal=𝔼\[ℓu\+ℓv\+λfft⋅𝕀\(t<0\.4\)ℒfft\]\.\\mathcal\{L\}\_\{total\}=\\mathbb\{E\}\\left\[\\ell\_\{u\}\+\\ell\_\{v\}\+\\lambda\_\{fft\}\\cdot\\mathbb\{I\}\(t<0\.4\)\\mathcal\{L\}\_\{fft\}\\right\]\.\(18\)
The low\-noise gate limits conflicts between spectral alignment and high\-noise velocity learning\. Fig\. 2 contrasts the nonsmooth outputs obtained without FFT loss with the periodic artifacts produced by globally applying the spectral term\.
Fig\. 2:Diagnostic visualizations motivating two\-stage noise\-level\-gated FFT regularization\. \(a\) Scalp topographies without FFT loss show locally nonsmooth synthetic fNIRS responses compared with real fNIRS\. \(b\) HRF curves without FFT loss show high\-frequency jitter and morphology deviation in synthetic fNIRS\. \(c\) HRF curves under global FFT regularization, without the two\-stage schedule and noise\-level gate, show periodic or impulse\-like artifacts\.Algorithm 1Training Phase of Bio\-MFDefine:𝐱0\\mathbf\{x\}\_\{0\}: real fNIRS;𝐲\\mathbf\{y\}: EEG condition;ϵ\\boldsymbol\{\\epsilon\}: Gaussian noiseℱθ=\(ℱθ,u,ℱθ,v\)\\mathcal\{F\}\_\{\\theta\}=\(\\mathcal\{F\}\_\{\\theta,u\},\\mathcal\{F\}\_\{\\theta,v\}\): Bio\-MF denoising networkInput:𝐱0,𝐲\\mathbf\{x\}\_\{0\},\\mathbf\{y\}State Construction:Samplet\>0,r,ω,tmin,tmaxt\>0,r,\\omega,t\_\{\\min\},t\_\{\\max\}andϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)𝐳t←\(1−t\)𝐱0\+tϵ\\mathbf\{z\}\_\{t\}\\leftarrow\(1\-t\)\\mathbf\{x\}\_\{0\}\+t\\boldsymbol\{\\epsilon\},𝐯t←\(𝐳t−𝐱0\)/t\\mathbf\{v\}\_\{t\}\\leftarrow\(\\mathbf\{z\}\_\{t\}\-\\mathbf\{x\}\_\{0\}\)/tMeanFlow Denoising:𝐱u←ℱθ,ux\(𝐳t,𝐲,t,r,ω,tmin,tmax\)\\mathbf\{x\}\_\{u\}\\leftarrow\\mathcal\{F\}\_\{\\theta,u\}^\{x\}\(\\mathbf\{z\}\_\{t\},\\mathbf\{y\},t,r,\\omega,t\_\{\\min\},t\_\{\\max\}\)𝐮←\(𝐳t−𝐱u\)/t\\mathbf\{u\}\\leftarrow\(\\mathbf\{z\}\_\{t\}\-\\mathbf\{x\}\_\{u\}\)/t,𝐯←ℱθ,v\(𝐳t,𝐲,t,r,ω,tmin,tmax\)\\mathbf\{v\}\\leftarrow\\mathcal\{F\}\_\{\\theta,v\}\(\\mathbf\{z\}\_\{t\},\\mathbf\{y\},t,r,\\omega,t\_\{\\min\},t\_\{\\max\}\)Guided Target:ωint←ω\\omega\_\{int\}\\leftarrow\\omegaift∈\[tmin,tmax\]t\\in\[t\_\{\\min\},t\_\{\\max\}\], otherwise11\(𝐯c,𝐯u\)←ℱθ,v\(𝐳t,𝐲,t,t,ωint,0,1\),ℱθ,v\(𝐳t,𝟎,t,t,1,0,1\)\(\\mathbf\{v\}\_\{c\},\\mathbf\{v\}\_\{u\}\)\\leftarrow\\mathcal\{F\}\_\{\\theta,v\}\(\\mathbf\{z\}\_\{t\},\\mathbf\{y\},t,t,\\omega\_\{int\},0,1\),\\mathcal\{F\}\_\{\\theta,v\}\(\\mathbf\{z\}\_\{t\},\\mathbf\{0\},t,t,1,0,1\)𝐯g←𝐯t\+\(1−1/ωint\)\(𝐯c−𝐯u\)\\mathbf\{v\}\_\{g\}\\leftarrow\\mathbf\{v\}\_\{t\}\+\(1\-1/\\omega\_\{int\}\)\(\\mathbf\{v\}\_\{c\}\-\\mathbf\{v\}\_\{u\}\)If condition dropout is applied,\(𝐲,𝐯g\)←\(𝟎,𝐯t\)\(\\mathbf\{y\},\\mathbf\{v\}\_\{g\}\)\\leftarrow\(\\mathbf\{0\},\\mathbf\{v\}\_\{t\}\)Loss/Update:Let𝒰θ,u\\mathcal\{U\}\_\{\\theta,u\}be the velocity field induced byℱθ,ux\\mathcal\{F\}\_\{\\theta,u\}^\{x\}through𝐮=\(𝐳t−𝐱u\)/t\\mathbf\{u\}=\(\\mathbf\{z\}\_\{t\}\-\\mathbf\{x\}\_\{u\}\)/t𝐕←𝐮\+\(t−r\)sg\[JVP\(𝒰θ,u,\(𝐳t,t,r\),\(𝐯c,1,0\)\)\]\\mathbf\{V\}\\leftarrow\\mathbf\{u\}\+\(t\-r\)\\operatorname\{sg\}\[\\operatorname\{JVP\}\(\\mathcal\{U\}\_\{\\theta,u\};\(\\mathbf\{z\}\_\{t\},t,r\),\(\\mathbf\{v\}\_\{c\},1,0\)\)\]ℒmf←‖𝐕−𝐯g‖22\+‖𝐯−𝐯g‖22\\mathcal\{L\}\_\{mf\}\\leftarrow\\\|\\mathbf\{V\}\-\\mathbf\{v\}\_\{g\}\\\|\_\{2\}^\{2\}\+\\\|\\mathbf\{v\}\-\\mathbf\{v\}\_\{g\}\\\|\_\{2\}^\{2\}ℒfft←0\\mathcal\{L\}\_\{fft\}\\leftarrow 0; if Stage II andt<0\.4t<0\.4, setℒfft←‖\|RFFT\(𝐱u\)\|−\|RFFT\(𝐱0\)\|‖22\\mathcal\{L\}\_\{fft\}\\leftarrow\\left\\\|\|\\operatorname\{RFFT\}\(\\mathbf\{x\}\_\{u\}\)\|\-\|\\operatorname\{RFFT\}\(\\mathbf\{x\}\_\{0\}\)\|\\right\\\|\_\{2\}^\{2\}ℒtotal←𝔼\[ℒmf\+λfftℒfft\]\\mathcal\{L\}\_\{total\}\\leftarrow\\mathbb\{E\}\[\\mathcal\{L\}\_\{mf\}\+\\lambda\_\{fft\}\\mathcal\{L\}\_\{fft\}\]Updateθ\\thetaby minimizingℒtotal\\mathcal\{L\}\_\{total\}
Algorithm 2Inference Phase of Bio\-MFDefine:𝒟θu\\mathcal\{D\}\_\{\\theta\}^\{u\}: trained Bio\-MF denoising networkInput:EEG𝐲\\mathbf\{y\}, CFG scaleω\\omega, interval\[tmin,tmax\]\[t\_\{\\min\},t\_\{\\max\}\]Initialize Generation:t←1t\\leftarrow 1,r←0r\\leftarrow 0Sampling Process:Sampleϵ∼𝒩\(𝟎,𝐈\)\\boldsymbol\{\\epsilon\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\mathbf\{I\}\)Denoising Process:𝐱^0←𝒟θu\(ϵ,𝐲,t,r,ω,tmin,tmax\)\\hat\{\\mathbf\{x\}\}\_\{0\}\\leftarrow\\mathcal\{D\}\_\{\\theta\}^\{u\}\(\\boldsymbol\{\\epsilon\},\\mathbf\{y\},t,r,\\omega,t\_\{\\min\},t\_\{\\max\}\)Output:Reconstruct final fNIRS signal𝐱^0\\hat\{\\mathbf\{x\}\}\_\{0\}
Bio\-MF was implemented as a Transformer denoising network with hidden dimension 384 and 8 Transformer blocks\. The architecture contains a 4\-block shared stem followed by separate 4\-blockuu\- andvv\-heads, with the generation branch used at inference and thevv\-head used only during training for velocity supervision\. Each block uses 6\-head self\-attention with 64 channels per head, RMSNorm, SwiGLU feed\-forward layers with an expansion ratio of8/38/3, vector\-valued residual gating, and spatiotemporal Fourier positional encoding that jointly encodes\(x,y,z\)\(x,y,z\)sensor coordinates and separately encodes the temporal coordinatett\.
Using this configuration, all Bio\-MF experiments were trained on a single NVIDIA RTX PRO 6000 GPU\. The model was optimized for 240 epochs, with a 10\-epoch warmup period at the beginning of training\. Following the two\-stage objective described above, Stage I occupied the first 150 epochs and optimized onlyℒMF\\mathcal\{L\}\_\{MF\}, while Stage II used the remaining 90 epochs and enabled the noise\-level\-gated FFT regularizer\. The complete training run took 8 h\. By comparison, SCDM is reported to use a 30,000\-epoch training schedule \[25\], indicating that Bio\-MF uses a substantially shorter optimization schedule in addition to its one\-step inference advantage\.
## IVResults
### IV\-AClassification
To evaluate whether generated fNIRS provides useful complementary information for MI\-BCI decoding, we performed downstream left/right motor imagery classification under single\-modality and hybrid\-modality settings\. For Dataset 1, paired EEG\-fNIRS trials provide the source\-domain supervision for EEG\-conditioned fNIRS synthesis\. For Dataset 2, which contains only EEG recordings with an unseen 64\-channel montage, Bio\-MF was applied in a zero\-shot manner to synthesize 36\-channel fNIRS signals on the source fNIRS layout\. Table 1 compares Bio\-MF with the SCDM classification results reported under the same classifier family and ratio\-based evaluation protocol \[25\]\.
To keep the comparison aligned with prior EEG\-to\-fNIRS generation studies, the same classifier family was used for each modality setting: ESNet for EEG\-only input, FSNet for fNIRS\-only input, and FGANet for EEG\-fNIRS hybrid input \[15\], \[25\]\. Following the SCDM evaluation protocol \[25\], classifiers were evaluated under seven LMI/RMI training ratios, from 1:4 to 4:1, with the remaining samples used for testing\. For Dataset 1, these ratios correspond to 200/800, 300/700, 400/600, 500/500, 600/400, 700/300, and 800/200 LMI/RMI training samples drawn from 870 trials per class\. The balanced 1:1 setting used 10\-fold cross\-validation repeated 10 times, whereas the other ratios used 20 random train/test repetitions; Dataset 2 followed the same ratio\-based protocol\. We report the mean and standard deviation over all runs using ACC, SPE, PRE, and SEN, where SEN and SPE measure LMI and RMI recall, respectively, and PRE measures LMI precision\.
The results are summarized in Table 1\. For each dataset, the Bio\-MF EEG\-only baseline is retained as the reference for interpreting synthetic fNIRS\-only and EEG \+ synthetic fNIRS performance\. The Bio\-MF fusion rows correspond to the Full Bio\-MF configuration with Interactive 4D encoding\.
TABLE I:Bio\-MF and SCDM classification results\.Across the generated\-modality ACC comparisons in Table 1, Bio\-MF is ahead of SCDM in nearly all Dataset 1 settings\. The only ACC exception is EEG \+ synthetic fNIRS \(HbR\), where Bio\-MF is lower by 0\.09 percentage points \(75\.29% vs\. 75\.38%\); in HbO fusion, Bio\-MF remains higher by 1\.04 percentage points \(76\.07% vs\. 75\.03%\)\. Averaged over HbR/HbO fusion, Bio\-MF still leads SCDM by 0\.47 percentage points \(75\.68% vs\. 75\.21%\)\. Within the HbR fusion row, SCDM is higher only in ACC and SEN, whereas Bio\-MF remains higher in SPE and PRE and leads all fNIRS\-only values and all HbO\-fusion values\.
Dataset 2 provides a more challenging zero\-shot evaluation because it contains only EEG recordings, uses a different 64\-channel electrode montage, and has no paired fNIRS measurements for adaptation\. In this cross\-device setting, Bio\-MF outperforms SCDM in every generated\-modality ACC, SPE, PRE, and SEN comparison\. For fusion decoding, Bio\-MF exceeds SCDM by 1\.58 and 1\.97 percentage points in HbR and HbO ACC, with smaller ACC standard deviations \(5\.45/5\.25 vs\. 9\.79/10\.56\)\. The Bio\-MF fusion gains over EEG\-only reach 2\.98 and 2\.50 percentage points, compared with 1\.40 and 0\.53 for SCDM\.
TABLE II:Generation latency comparison\.†TADM latency is the paper\-reported value for one fNIRS sequence of length 1000 under the reported NVIDIA A800 experimental setting and is not hardware\-matched to our RTX PRO 6000 profiling;CfNIRSC\_\{fNIRS\}denotes the dataset\-specific fNIRS channel count\.
### IV\-BGeneration Latency
To assess generation latency relevant to hybrid MI\-BCI systems, we conducted an inference profiling experiment \[13\], \[14\]\. For the hardware\-matched comparison, Bio\-MF and SCDM were evaluated under the same input setting, where one EEG trial was used to generate one fNIRS sequence with shape2×36×2562\\times 36\\times 256\. For SCDM, the diffusion noising\-step search described above selected a Wasserstein\-minimum schedule withT=1000T=1000and a linear beta range from1×10−51\\times 10^\{\-5\}to0\.0150\.015\[25\]\. During inference, SCDM requires 1000 serial denoising evaluations for each generated fNIRS trial\. In contrast, Bio\-MF follows the single\-step MeanFlow generation procedure and synthesizes each trial with one network function evaluation \(1\-NFE\) \[31\]\. We additionally include TADM \[26\] as an external reported reference: TADM performs diffusion in a 512\-dimensional latent representation and reports 0\.5 s for generating one fNIRS sequence of length 1000, with the final fNIRS channel count depending on the dataset\-specific sensor layout\.
We profiled the implemented Bio\-MF and SCDM inference pipelines on the same NVIDIA RTX PRO 6000 GPU, and report the average per\-trial generation time from repeated runs under the same input configuration\. As summarized in Table 2, the latency gap mainly comes from the number of serial denoising evaluations: SCDM repeats its denoising pass 1000 times, whereas Bio\-MF generates each trial with one network function evaluation\.
The profiling results indicate that latency is dominated by the number and location of serial denoising evaluations rather than by model size alone\. SCDM performs a long reverse chain directly in the raw fNIRS signal space, causing the 1000\-step schedule to accumulate to 6\.0 s per trial on the RTX PRO 6000\. TADM reduces the per\-step state to a compact 512\-dimensional latent representation, which explains why its reported 1000\-step latency is much lower despite still using iterative diffusion\. Bio\-MF avoids both the latent reconstruction path and the long reverse chain: it generates the complete raw fNIRS trial with one network function evaluation, reducing measured total generation time to 7\.0 ms per trial on the RTX PRO 6000\. This corresponds to an 857x generation\-speed advantage over the hardware\-matched SCDM baseline and provides a practical latency margin for hybrid MI\-BCI feedback\.
### IV\-CPhysiological Visualization
We further examine whether Bio\-MF preserves plausible fNIRS morphology from spatial, temporal, and topographic perspectives\. For spatial correspondence, EEG signals are resampled to the fNIRS temporal length, and each fNIRS channel is connected to the EEG electrode with the highest mean absolute Pearson correlation over Dataset 1\. The spatial visualization shows that the synthetic fNIRS correspondences broadly follow those of real fNIRS, indicating that generation does not collapse to arbitrary channel patterns\.
Fig\. 3:EEG\-fNIRS spatial correspondence visualization on Dataset 1\. Blue markers denote EEG electrodes, orange markers denote fNIRS recording sites, and purple lines connect each fNIRS site to the EEG electrode with the highest average absolute Pearson correlation\.For temporal plausibility, trial\-averaged HRF curves are computed for real and synthetic fNIRS under LMI/RMI conditions\. The HRF visualization shows that synthetic curves follow the slow low\-frequency trends and cross\-trial variation ranges of real fNIRS in the main hemodynamic response interval, without obvious high\-frequency oscillation or mean\-curve collapse\.
Fig\. 4:HRF\-curve visualization on Dataset 1\. The panels compare average hemodynamic response curves and cross\-trial variation ranges of real and synthetic fNIRS under LMI/RMI conditions\.For scalp topology, the topography visualization compares real and synthetic fNIRS across consecutive hemodynamic windows\. Synthetic fNIRS preserves the major activation regions, spatial gradients, and temporal evolution trends of real fNIRS, although local amplitude differences remain\. Together with Table 1, these visualizations support the use of Bio\-MF\-generated fNIRS as an auxiliary modality for hybrid MI\-BCI decoding\.
Fig\. 5:Scalp\-topography visualization on Dataset 1\. This visualization compares real and synthetic fNIRS topographies for HbR/HbO under LMI/RMI conditions, evaluating whether synthetic fNIRS preserves the main activation regions, spatial gradients, and temporal evolution trends of real fNIRS\.Additional analyses of coordinate sensitivity and CFG strength are provided in the supplementary appendix\.
## VAblation Study
To evaluate the contribution of each core Bio\-MF module to final performance, we compare Full Bio\-MF with variants that modify one design choice at a time: Original 4D replaces Interactive 4D with the original 4D encoding; w/o 4D PE removes physical coordinate encoding; w/o CFG disables the conditional guidance offset; w/o FFT removes spectral regularization; and Global FFT applies FFT regularization at all noise levels\. Table 3 reports downstream classification results under the EEG \+ Synthetic fNIRS setting, averaged over HbR/HbO\. Detailed variant definitions are provided in the supplementary appendix\.
TABLE III:Classification ablation results \(%, EEG \+ Synthetic fNIRS; HbR/HbO average\)\.The ablation results show that the three modules contribute complementary benefits to downstream EEG \+ Synthetic fNIRS decoding\. Replacing Interactive 4D with the original 4D encoding reduces ACC on both datasets, indicating that explicit spatial\-temporal interaction improves in\-domain and cross\-montage decoding\. Removing physical 4D coordinates further degrades performance, especially for Dataset 2\. Removing CFG leads to a consistent decline, suggesting that the conditional offset between EEG responses and the fNIRS prior helps preserve task\-relevant information\. Removing FFT loss or applying global FFT also reduces classification performance, with Global FFT showing a larger decline than the two\-stage low\-noise\-masked FFT regularization\. These results support the use of Interactive 4D encoding, CFG, and noise\-level\-gated spectral regularization in the final Bio\-MF design\.
## VIConclusion and Discussion
This study proposed Bio\-MF for one\-step EEG\-to\-fNIRS cross\-modal generation in hybrid motor imagery BCIs\. The generated fNIRS signals provide useful complementary information for EEG\-based decoding, indicating that synthetic hemodynamic responses can serve as a lightweight auxiliary modality for hybrid MI\-BCI settings when simultaneous EEG\-fNIRS acquisition is unavailable\. Across Table 1, Bio\-MF improves nearly all generated\-modality ACC comparisons over SCDM after incorporating Interactive 4D encoding, and on Dataset 2 also improves SPE, PRE, and SEN with smaller standard deviations\. Bio\-MF generates one fNIRS trial in 7\.0 ms on an RTX PRO 6000 GPU, corresponding to an 857x speedup over 1000\-step SCDM sampling\. Ablation and visualization results further indicate that Interactive 4D encoding, cross\-modal CFG, and noise\-level\-gated FFT regularization jointly preserve task\-useful fNIRS structure while keeping the one\-step inference path\. Together, these findings support Bio\-MF as a low\-latency and high\-quality EEG\-to\-fNIRS signal generation route for hybrid MI\-BCIs\.
Several limitations remain\. First, Dataset 2 contains only EEG recordings, so the zero\-shot experiments mainly assess whether generated fNIRS improves downstream decoding under an unseen EEG montage rather than directly measuring channel\-wise synthetic\-real agreement\. Second, Bio\-MF is trained on a single paired EEG\-fNIRS source dataset, which may restrict the diversity of neurovascular correspondence patterns learned by the generator\. Third, although the visualizations show preserved HRF trends and scalp\-topographic structure, local amplitude differences remain and may require subject\-specific physiological calibration\. Future work should evaluate paired cross\-device EEG\-fNIRS datasets, broader cohorts and paradigms, longitudinal sessions, uncertainty\-aware generation, and user\-centered closed\-loop validation\.
## References
- \[1\]J\. R\. Wolpaw, N\. Birbaumer, D\. J\. McFarland, G\. Pfurtscheller, and T\. M\. Vaughan, ”Brain\-computer interfaces for communication and control,” Clinical Neurophysiology, vol\. 113, no\. 6, pp\. 767\-791, 2002, doi: 10\.1016/S1388\-2457\(02\)00057\-3\.
- \[2\]G\. Pfurtscheller and C\. Neuper, ”Motor imagery and direct brain\-computer communication,” Proceedings of the IEEE, vol\. 89, no\. 7, pp\. 1123\-1134, 2001, doi: 10\.1109/5\.939829\.
- \[3\]G\. Pfurtscheller and F\. H\. Lopes da Silva, ”Event\-related EEG/MEG synchronization and desynchronization: Basic principles,” Clinical Neurophysiology, vol\. 110, no\. 11, pp\. 1842\-1857, 1999, doi: 10\.1016/S1388\-2457\(99\)00141\-8\.
- \[4\]F\. Lotte et al\., ”A review of classification algorithms for EEG\-based brain\-computer interfaces: A 10 year update,” Journal of Neural Engineering, vol\. 15, no\. 3, Art\. no\. 031005, 2018, doi: 10\.1088/1741\-2552/aab2f2\.
- \[5\]T\. Fang et al\., ”Noninvasive neuroimaging and spatial filter transform enable ultra low delay motor imagery EEG decoding,” Journal of Neural Engineering, vol\. 19, no\. 6, Art\. no\. 066034, 2022, doi: 10\.1088/1741\-2552/aca82d\.
- \[6\]M\. Hassan and F\. Wendling, ”Electroencephalography source connectivity: Aiming for high resolution of brain networks in time and space,” IEEE Signal Processing Magazine, vol\. 35, no\. 3, pp\. 81\-96, 2018, doi: 10\.1109/MSP\.2017\.2777518\.
- \[7\]R\. Sitaram et al\., ”Temporal classification of multichannel near\-infrared spectroscopy signals of motor imagery for developing a brain\-computer interface,” NeuroImage, vol\. 34, no\. 4, pp\. 1416\-1427, 2007, doi: 10\.1016/j\.neuroimage\.2006\.11\.005\.
- \[8\]N\. Naseer and K\.\-S\. Hong, ”fNIRS\-based brain\-computer interfaces: A review,” Frontiers in Human Neuroscience, vol\. 9, Art\. no\. 3, 2015, doi: 10\.3389/fnhum\.2015\.00003\.
- \[9\]S\. B\. Borgheai et al\., ”Enhancing communication for people in late\-stage ALS using an fNIRS\-based BCI system,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol\. 28, no\. 5, pp\. 1198\-1207, 2020, doi: 10\.1109/TNSRE\.2020\.2980772\.
- \[10\]J\. Lu et al\., ”An fNIRS\-based dynamic functional connectivity analysis method to signify functional neurodegeneration of Parkinson’s disease,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol\. 31, pp\. 1199\-1207, 2023, doi: 10\.1109/TNSRE\.2023\.3242263\.
- \[11\]D\. J\. Heeger and D\. Ress, ”What does fMRI tell us about neuronal activity?” Nature Reviews Neuroscience, vol\. 3, no\. 2, pp\. 142\-151, 2002, doi: 10\.1038/nrn730\.
- \[12\]S\. Fazli et al\., ”Enhanced performance by a hybrid NIRS\-EEG brain computer interface,” NeuroImage, vol\. 59, no\. 1, pp\. 519\-529, 2012, doi: 10\.1016/j\.neuroimage\.2011\.07\.084\.
- \[13\]G\. Pfurtscheller, B\. Z\. Allison, G\. Bauernfeind, C\. Brunner, T\. Solis\-Escalante, R\. Scherer, T\. O\. Zander, G\. Mueller\-Putz, C\. Neuper, and N\. Birbaumer, ”The hybrid BCI,” Frontiers in Neuroscience, vol\. 4, Art\. no\. 3, 2010, doi: 10\.3389/fnpro\.2010\.00003\.
- \[14\]C\.\-H\. Han, K\.\-R\. Mueller, and H\.\-J\. Hwang, ”Enhanced performance of a brain switch by simultaneous use of EEG and NIRS data for asynchronous brain\-computer interface,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol\. 28, no\. 10, pp\. 2102\-2112, 2020, doi: 10\.1109/TNSRE\.2020\.3017167\.
- \[15\]Y\. Kwak, W\.\-J\. Song, and S\.\-E\. Kim, ”FGANet: fNIRS\-guided attention network for hybrid EEG\-fNIRS brain\-computer interfaces,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol\. 30, pp\. 329\-339, 2022, doi: 10\.1109/TNSRE\.2022\.3149899\.
- \[16\]Z\. Wang et al\., ”Incorporating EEG and fNIRS patterns to evaluate cortical excitability and MI\-BCI performance during motor training,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol\. 31, pp\. 2872\-2882, 2023, doi: 10\.1109/TNSRE\.2023\.3281855\.
- \[17\]Y\. Gao, B\. Jia, M\. Houston, and Y\. Zhang, ”Hybrid EEG\-fNIRS brain computer interface based on common spatial pattern by using EEG\-informed general linear model,” IEEE Transactions on Instrumentation and Measurement, vol\. 72, pp\. 1\-10, 2023, doi: 10\.1109/TIM\.2023\.3276509\.
- \[18\]W\. Huang, X\. Song, and D\. Kuang, ”A hierarchical diffusion\-convolutional network with node\-wise localization for EEG\-NIRS\-based brain\-computer interface,” in Proc\. International Winter Conference on Brain\-Computer Interface, pp\. 1\-6, 2024, doi: 10\.1109/BCI60775\.2024\.10480493\.
- \[19\]J\. Uchitel, E\. E\. Vidal\-Rosas, R\. J\. Cooper, and H\. Zhao, ”Wearable, integrated EEG\-fNIRS technologies: A review,” Sensors, vol\. 21, no\. 18, Art\. no\. 6106, 2021, doi: 10\.3390/s21186106\.
- \[20\]H\. Khan et al\., ”Analysis of human gait using hybrid EEG\-fNIRS\-based BCI system: A review,” Frontiers in Human Neuroscience, vol\. 14, Art\. no\. 613254, 2021, doi: 10\.3389/fnhum\.2020\.613254\.
- \[21\]V\. Kaiser et al\., ”Cortical effects of user training in a motor imagery based brain\-computer interface measured by fNIRS and EEG,” NeuroImage, vol\. 85, pp\. 432\-444, 2014, doi: 10\.1016/j\.neuroimage\.2013\.04\.097\.
- \[22\]R\. K\. Almajidy, K\. Mankodiya, M\. Abtahi, and U\. G\. Hofmann, ”A newcomer’s guide to functional near infrared spectroscopy experiments,” IEEE Reviews in Biomedical Engineering, vol\. 13, pp\. 292\-308, 2020, doi: 10\.1109/RBME\.2019\.2944351\.
- \[23\]A\. M\. Chiarelli et al\., ”Fiberless, multi\-channel fNIRS\-EEG system based on silicon photomultipliers: Towards sensitive and ecological mapping of brain activity and neurovascular coupling,” Sensors, vol\. 20, no\. 10, Art\. no\. 2831, 2020, doi: 10\.3390/s20102831\.
- \[24\]W\. Yao, Z\. Lyu, M\. Mahmud, N\. Zhong, B\. Lei, and S\. Wang, ”CATD: Unified representation learning for EEG\-to\-fMRI cross\-modal generation,” IEEE Transactions on Medical Imaging, vol\. 44, no\. 7, pp\. 2757\-2767, 2025, doi: 10\.1109/TMI\.2025\.3550206\.
- \[25\]Y\. Li, Y\. Wang, B\. Lei, and S\. Wang, ”SCDM: Unified representation learning for EEG\-to\-fNIRS cross\-modal generation in MI\-BCIs,” IEEE Transactions on Medical Imaging, vol\. 44, no\. 6, pp\. 2384\-2394, 2025, doi: 10\.1109/TMI\.2025\.3532480\.
- \[26\]B\. Yuan, Y\. Li, S\. Wang, Z\. Zhang, Z\. Huang, H\. Kuai, X\. Zhao, N\. Zhong, M\. K\.\-P\. Ng, and K\.\-F\. Tsang, ”TADM: Unified pre\-trained framework for EEG\-to\-fNIRS cross\-modal generation in BCI decoding,” IEEE Transactions on Consumer Electronics, vol\. 72, no\. 2, pp\. 3752\-3763, 2026, doi: 10\.1109/TCE\.2026\.3682400\.
- \[27\]J\. Shin et al\., ”Open access dataset for EEG\+NIRS single\-trial classification,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol\. 25, no\. 10, pp\. 1735\-1745, 2017, doi: 10\.1109/TNSRE\.2016\.2628057\.
- \[28\]A\. L\. Goldberger et al\., ”PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals,” Circulation, vol\. 101, no\. 23, pp\. e215\-e220, 2000, doi: 10\.1161/01\.CIR\.101\.23\.e215\.
- \[29\]G\. Schalk, D\. J\. McFarland, T\. Hinterberger, N\. Birbaumer, and J\. R\. Wolpaw, ”BCI2000: A general\-purpose brain\-computer interface \(BCI\) system,” IEEE Transactions on Biomedical Engineering, vol\. 51, no\. 6, pp\. 1034\-1043, 2004, doi: 10\.1109/TBME\.2004\.827072\.
- \[30\]Z\. Geng, M\. Deng, X\. Bai, J\. Z\. Kolter, and K\. He, ”Mean flows for one\-step generative modeling,” arXiv preprint arXiv:2505\.13447, 2025, doi: 10\.48550/arXiv\.2505\.13447\.
- \[31\]Y\. Lu et al\., ”One\-step latent\-free image generation with Pixel Mean Flows,” arXiv preprint arXiv:2601\.22158, 2026, doi: 10\.48550/arXiv\.2601\.22158\.
- \[32\]A\. Vaswani et al\., ”Attention is all you need,” in Advances in Neural Information Processing Systems, vol\. 30, pp\. 5998\-6008, 2017\.
- \[33\]A\. Dosovitskiy et al\., ”An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc\. International Conference on Learning Representations, 2021\.
- \[34\]W\. Peebles and S\. Xie, ”Scalable diffusion models with Transformers,” in Proc\. IEEE/CVF International Conference on Computer Vision, pp\. 4195\-4205, 2023\.
- \[35\]M\. Tancik et al\., ”Fourier features let networks learn high frequency functions in low dimensional domains,” in Advances in Neural Information Processing Systems, vol\. 33, pp\. 7537\-7547, 2020\.
- \[36\]Y\. El Ouahidi et al\., ”REVE: A foundation model for EEG – adapting to any setup with large\-scale pretraining on 25,000 subjects,” arXiv preprint arXiv:2510\.21585, 2025, doi: 10\.48550/arXiv\.2510\.21585\.
- \[37\]J\. Ho and T\. Salimans, ”Classifier\-free diffusion guidance,” arXiv preprint arXiv:2207\.12598, 2022, doi: 10\.48550/arXiv\.2207\.12598\.
- \[38\]B\. Zhang and R\. Sennrich, ”Root mean square layer normalization,” in Advances in Neural Information Processing Systems, vol\. 32, 2019\.
- \[39\]N\. Shazeer, ”GLU variants improve Transformer,” arXiv preprint arXiv:2002\.05202, 2020, doi: 10\.48550/arXiv\.2002\.05202\.相似文章
FRIST:基于fMRI表示学习的共享空间训练提升仅EEG信号的个体手指脑机接口解码
FRIST是一种两阶段EEG解码框架,利用fMRI数据增强仅基于EEG的个体手指脑机接口解码能力,在运动执行和运动想象任务中均提升了准确率。
MEL:用于fMRI翻译的坐标保持型EEG标记化
本文介绍了MEL,一个用于EEG到fMRI翻译的坐标保持型EEG标记化框架,通过显式建模血流动力学延迟和频谱空间动态,解决表示接口不匹配问题,并改善预测性能优于基线方法。
基于小波图像变换和谱流匹配的功能磁共振时间序列生成用于脑疾病识别
本文提出DSFM,一种新颖的生成框架,利用小波分解和谱流匹配合成逼真的fMRI时间序列,用于脑疾病识别,解决了数据稀缺和非平稳性挑战。
BrainG3N: 一种用于可控3D脑部MRI生成的双用途分词器
介绍了BrainG3N,一种用于3D脑部MRI潜在扩散的双用途分词器,它使用冻结的掩码自编码器(MAE)编码器生成临床信息丰富的嵌入表示,并使用CNN解码器进行重建,在23个任务的基准测试中达到了最先进性能,并实现了可控生成和纵向预测。
从胎儿-母亲心电图到胎儿多普勒波形的跨模态生成式信号转换框架
本文提出了一种跨模态生成框架,利用跨模态注意力和膨胀卷积从胎儿-母亲心电图合成胎儿多普勒超声波形,提高了合成质量,并量化了母胎耦合的影响。