MedTVL: Harnessing Vision and Language for Medical Time Series Classification
Summary
MedTVL is a text-guided dual-pathway architecture for medical time series classification that synergizes temporal and visual modalities with textual semantics, demonstrating superiority in various learning settings.
View Cached Full Text
Cached at: 09/01/26, 12:30 PM
# MedTVL: Harnessing Vision and Language for Medical Time Series Classification Source: [https://arxiv.org/html/2608.28605](https://arxiv.org/html/2608.28605) Jiexia Ye[jye324@connect\.hkust\-gz\.edu\.cn](mailto:[email protected])The Hong Kong University of Science and Technology \(Guangzhou\)Data Science and Analytics ThrustGuangzhouChinaJia Li[jialee@hkust\-gz\.edu\.cn](mailto:[email protected])The Hong Kong University of Science and Technology \(Guangzhou\)Data Science and Analytics ThrustGuangzhouChinaandFugee Tsung[season@ust\.hk](mailto:[email protected])The Hong Kong University of Science and TechnologyDepartment of Industrial Engineering & Decision AnalyticsHong Kong SARChina \(2026\) ###### Abstract\. Recent advancements in multimodal learning for medical time series \(MedTS\) classification highlight the benefits of integrating complementary modalities for clinical decision\. However, existing methods typically focus on bi\-modal interactions \(e\.g\., time series and text\), leaving the tri\-modal synergy between time series, vision, and language largely unexplored\. Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, we introduce MedTVL, a text\-guided dual\-pathway architecture tailored for MedTS classification\. Specifically, it synergizes a convolution\-based temporal pathway for fine\-grained temporal dynamics from raw numerical sequences and a transformer\-based visual pathway for holistic morphological structures from time\-series\-derived images\. Such combination of cross\-modal and architectural heterogeneity provides a comprehensive diagnostic perspective\. To further resolve potential diagnostic ambiguity, both pathways are guided by adaptive medical textual semantics\. Finally, a Mixture\-of\-Experts mechanism dynamically routes each instance to specialized fusion experts, capturing instance\-specific reliance on the temporal and visual pathway outputs\. In addition, MedTVL supports multimodal contrastive learning to mitigate the clinical label scarcity challenge\. Extensive experiments across multiple medical datasets and tasks, spanning supervised, few\-shot, and contrastive learning settings, demonstrate the superiority and transferability of MedTVL, highlighting its potential for robust clinical decision support\. Medical Time Series Classification; Multimodal Learning; Contrastive Learning ††journalyear:2026††copyright:cc††conference:Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2; August 09–13, 2026; Jeju Island, Republic of Korea††booktitle:Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2 \(KDD ’26\), August 09–13, 2026, Jeju Island, Republic of Korea††doi:10\.1145/3770855\.3818883††isbn:979\-8\-4007\-2259\-2/2026/08††ccs:Applied computing Health informatics††ccs:Computing methodologies Temporal reasoning## 1\.Introduction Figure 1\.Illustration of the multi\-view diagnostic process for MedTS, highlighting the complementary strengths of numerical, visual, and textual views to support clinical decisions\.Medical time series \(MedTS\) are widely applied in critical healthcare scenarios such as clinical monitoring, disease screening, and diagnostic decision\-making\(Guet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib21)\), encompassing diverse physiological signals including electrocardiograms \(ECG\)\(Dinget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib22)\), electroencephalograms \(EEG\)\(Sharma and Meena,[2024](https://arxiv.org/html/2608.28605#bib.bib23)\), and other vital signs\(Fatourechiet al\.,[2007](https://arxiv.org/html/2608.28605#bib.bib26)\)\. The nature of MedTS is complex across several dimensions: First, significant data heterogeneity, characterized by multi\-source signal modalities, disparate sampling rates, and diverse diagnostic tasks\(Rimet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib25); Wooet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib20)\)\. Second, prominent multi\-scale patterns, where certain patterns manifest as transient fluctuations within short time windows, while others are reflected in holistic morphological structures\(Wanget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib67)\)\. Third, inherent label scarcity, as incomplete annotations in real\-world clinical environments often compromise the generalizability of models\(Liet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib84)\)\. These multifaceted complexities pose challenges to developing robust and generalizable models for MedTS classification\. Traditional methods mostly rely on unimodal numerical signals\(Wanget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib67); Yeet al\.,[2026](https://arxiv.org/html/2608.28605#bib.bib69)\), with a few attempts to convert time series into images\(Wuet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib2)\)\. However, unimodal modeling struggles to capture the full spectrum of diagnostic evidence, leading to sub\-optimal performance\. With the advancement of multimodal learning, an increasing number of studies have integrated time series with text, validating the significance of textual semantics in supporting clinical decision\-making\(Liuet al\.,[2024a](https://arxiv.org/html/2608.28605#bib.bib83); Chanet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib77)\)\. However, due to the lack of explicit visual modeling, these methods fail to capture critical morphological patterns, which deviates from clinical practice where clinicians heavily rely on visual reasoning\. Recently, time\-series\-vision fusion has emerged within the general time series domain\(Shenet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib119); Lyuet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib113)\); however, these methods lack textual grounding and are not tailored to the unique properties of medical data\. While interest in tri\-modal modeling \(time series, vision, and language\) is growing, existing studies remain scarce and are restricted to either general\-purpose forecasting\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\)or specific medical modalities\(Lanet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib124)\)\. For instance, TimeVLM\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\)focuses on time series forecasting with temporal\-centric attention, treating vision and text as auxiliary, a limitation that prevents full exploitation of these modalities\. GEM\(Lanet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib124)\), specifically designed for ECG, maps time series and images into a shared textual space, which may buffer and dilute fine\-grained physiological details\. In real\-world clinical practice, the diagnosis of MedTS is inherently a multi\-view process\(Lanet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib124); Fanet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib125)\)\. Clinicians often examine raw temporal signals to identify fine\-grained pathological patterns \(e\.g\. premature ventricular contractions\(Kleweret al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib19)\)\)\. Complementarily, visual inspection of signal enables them to assess holistic morphology \(e\.g\.,ST\-segment elevation\(McLarenet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib18)\)\)\. Meanwhile, clinical text provides essential semantic context by incorporating patient history and prior assessments, allowing clinicians to confirm, refine, or revise their judgments\. Together, these views reflect how clinicians integrate precise temporal details, global visual cues, and semantic information into a coherent diagnostic decision\. Building upon and transcending these clinical practices, we proposeMedTVL, a unified framework that integratesTime series,Vision, andLanguage, tailored for MedTS classification\. MedTVL employs a text\-guided dual\-pathway architecture to model diverse diagnostic patterns across multiple scales\. Specifically, the temporal pathway utilizes a convolutional backbone to extract fine\-grained dynamics from raw numerical signals, while the visual pathway adopts a transformer\-based backbone to capture holistic morphology from time series\-derived Continuous Wavelet Transform \(CWT\) images\. This synergistic design goes beyond modality\-level complementarity by explicitly introducing architectural heterogeneity, where convolutional and transformer\-based backbones offer complementary inductive biases for multi\-scale diagnostic modeling\. Subsequently, shared textual semantics are adaptively injected into the pathway subspaces to constrain modality\-specific learning with clinical context\. Finally, to address the varying importance of temporal and visual cues across samples, MedTVL introduces a Mixture\-of\-Experts \(MoE\) mechanism that dynamically assigns each instance to specialized fusion experts\. This fine\-grained fusion allows the model to adaptively reconcile heterogeneous signals, providing robust and flexible integration across diverse clinical scenarios\. Additionally, the dual\-pathway structure naturally forms cross\-modal positive pairs, facilitating multimodal contrastive learning to mitigate the label scarcity challenge in clinical settings\. - •To the best of our knowledge, MedTVL is the first tri\-modal framework that encapsulates clinical diagnostic perspectives for general\-purpose medical time series classification\. - •MedTVL introduces a dual\-pathway architecture to capture multi\-scale patterns from synergistic temporal\-visual perspectives, using adaptive textual guidance and MoE\-based instance\-level fusion to handle medical data heterogeneity\. It also supports multimodal contrastive learning\. - •We conduct comprehensive experiments on datasets spanning multiple medical tasks, covering supervised, few\-shot, and contrastive learning settings\. Our results demonstrate that MedTVL outperforms SOTA baselines, while also showcasing its capability to address the label scarcity challenge\. The code link ishttps://github\.com/start2020/MedTVL Figure 2\.Overview of the MedTVL framework\. More details are in section[3](https://arxiv.org/html/2608.28605#S3)\. ## 2\.Related Work Unimodal Medical Time Series Classification\.Unimodal approaches predominantly rely on numerical medical signal modeling, with a limited number of studies exploring image\-based representations of time series\(Wuet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib2); Pratiheret al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib70)\)\. Many recent deep learning approaches adopt architectures such as convolutional neural networks \(CNNs\)\(Lawhernet al\.,[2018](https://arxiv.org/html/2608.28605#bib.bib65)\), recurrent neural networks \(RNNs\)\(Salloum and Kuo,[2017](https://arxiv.org/html/2608.28605#bib.bib64)\), graph neural networks \(GNNs\)\(Tanget al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib44)\), and Transformers\(Wanget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib67); Yeet al\.,[2026](https://arxiv.org/html/2608.28605#bib.bib69)\)\. While such methods have shown promising results, they are often tailored for specific signals \(e\.g\., ECG\) with limited cross\-signal generalization\. Moreover, relying on a unimodal perspective constrains their capacity to fully capture the variability and heterogeneity of real\-world clinical data, thereby motivating increasing interest in multimodal learning\. Bi\-modal Time Series Representation LearningCurrent multimodal research focuses predominantly on time\-series–text alignment, particularly within general time\-series domains\(Chenget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib74); Yeet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib68)\)\(e\.g\. Time\-LLM\(Jinet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib72)\), UniTS\(Gaoet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib73)\)\)\. In the medical domain, some works leverage Large Language Models \(LLMs\) to integrate time series and text \(e\.g\., MedTsLLM\(Chanet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib77)\), MedualTime\(Yeet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib78)\)\) for supervised learning while other focus on report\-guided ECG contrastive learning\(Liet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib84); Liuet al\.,[2024a](https://arxiv.org/html/2608.28605#bib.bib83)\)\. However, these methods typically overlook explicit visual modeling, failing to capture clinically meaningful morphological patterns essential for diagnostic practice\. Although time series–vision approaches have recently emerged\(Shenet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib119); Lyuet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib113)\), they primarily focus on forecasting and remain unoptimized for the unique challenges of clinical diagnostics\. A notable exception is MedViA\(Fanet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib125)\), which is specifically designed for medical data; however, it relies on static fusion and lacks clinical semantic guidance, limiting its flexibility to handle MedTS variability\. Tri\-modal Time Series Representation Learning\.The exploration of tri\-modal frameworks remains more scarce, to our knowledge, with only Time\-VLM\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\)and GEM\(Lanet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib124)\)reported to date\. Time\-VLM\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\)prioritizes forecasting via temporal\-dominant attention, limiting full exploitation of the visual and textual inputs\. Its gated fusion applies a shared weighting across instances, potentially limiting the model’s flexibility in integrating heterogeneous pathway outputs\. Conversely, GEM\(Lanet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib124)\)collapses both numerical signals and visual images into a unified textual space, which inevitably erodes fine\-grained waveform dynamics and morphological details\. Contrastive Learning for Time Series\.Contrastive learning is a powerful paradigm for mitigating label scarcity\. In the general time series domain, it has evolved from methods defining positive pairs\(Tonekaboniet al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib95)\), to augmentation\-based multi\-view alignment \(e\.g\., TS\-TCC\(Eldeleet al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib92)\), TS2Vec\(Yueet al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib94)\)\), and further to approaches exploiting intrinsic temporal properties, such as frequency\-domain consistency and seasonal–trend disentanglement\(Wooet al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib96); Zhanget al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib97)\)\. Medical time series contrastive learning has progressed from early adoption of general\-domain techniques to deeper adaptation to medical characteristics while most of them are tailored to specific signal types \(e\.g\., EEG\)\(Liuet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib79)\)\. All these methods above remain confined to the time series modality\. Recently, some works focus on report\-guided ECG contrastive learning\(Liet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib84); Liuet al\.,[2024a](https://arxiv.org/html/2608.28605#bib.bib83)\)\. However, the absence of paired text in most public datasets limits the application of report\-guided contrastive learning\. AimTS\(Chenet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib89)\), a recent pioneering work, introduces time series–image contrastive learning to enhance representation generalization\. Compared with MedTVL, it is designed for general time series and lacks textual semantic support\. ## 3\.Methodology Problem Formulation\.Consider a medical dataset withNNsamples𝒟=\{\(𝐗i,si,yi\)\}i=1N\\mathcal\{D\}=\\\{\(\\mathbf\{X\}\_\{i\},s\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}, where𝐗∈ℝL×C\\mathbf\{X\}\\in\\mathbb\{R\}^\{L\\times C\}denotes the raw temporal signal with lengthLLandCCchannels;y∈𝒴=\{1,2,…,M\}y\\in\\mathcal\{Y\}=\\\{1,2,\\dots,M\\\}is its label andMMis the number of classes;sis\_\{i\}represents its associated clinical semantics\. To ensure broad applicability across clinical settings,ssis defined with flexible granularity, ranging from sample\-level diagnostic reports to dataset\-level semantic descriptions, or their combination\. For𝐗\\mathbf\{X\}, a corresponding visual representation𝐕=g\(𝐗\)\\mathbf\{V\}=g\(\\mathbf\{X\}\)is derived\. Our objective is to develop a frameworkf\(⋅\)f\(\\cdot\)that leverages the complementary strengths of the tri\-modal information for MedTS classification:y¯=f\(𝐗,𝐕,s;Θ\)\\bar\{y\}=f\(\\mathbf\{X\},\\mathbf\{V\},s;\\Theta\)whereΘ\\Thetadenotes the trainable parameters\. The framework under supervised learning is optimized to by minimizing the cross\-entropy loss:ℒce=−1N∑i=1Nyilog\(y^i\)\\mathcal\{L\}\_\{ce\}=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}y\_\{i\}\\log\(\\hat\{y\}\_\{i\}\)\. Overview\.MedTVL is a tri\-modal framework inspired by the holistic diagnostics in clinical practice, designed for comprehensive medical time\-series \(MedTS\) modeling\. As illustrated in Figure 1, MedTVL comprises four key components\. The temporal pathway adopts InceptionTime\(Ismail Fawazet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib8)\)as backbone to process raw time series and extract multi\-scale patterns from a numerical perspective\. In parallel, the visual pathway employs a Swin Transformer to capture structural representations from CWT scalograms, complementing temporal dynamics with morphological cues\. Within each pathway, shared textual information is adaptively projected into modality\-specific representations to provide clinical semantic guidance\. Finally, the text\-enhanced dual\-pathway features are routed through a Mixture\-of\-Experts module, enabling adaptive fusion of temporal and visual cues on a per\-sample basis\. ### 3\.1\.Convolutional Network\-based Temporal Pathway The temporal pathway aims to characterize the fine\-grained multi\-scale temporal dynamics from a numerical perspective\. Medical time series often exhibit clinically informative patterns that are localized in time and vary in temporal extent, such as brief waveform distortions or short\-lived abnormal rhythms\(Liuet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib79)\)\. Convolutional neural networks\(Cuiet al\.,[2016](https://arxiv.org/html/2608.28605#bib.bib9); Zhaoet al\.,[2017](https://arxiv.org/html/2608.28605#bib.bib7)\)are well suited for modeling such intrinsic numerical dynamics due to their strong locality bias and temporal shift invariance\. InceptionTime as Backbone\.Among CNN\-based architectures, we adopt InceptionTime\(Ismail Fawazet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib8)\)as the backbone of the temporal pathway, owing to its ability to model temporal patterns across multiple receptive fields simultaneously\. Unlike conventional 1D CNNs with fixed kernel sizes\(Heet al\.,[2016](https://arxiv.org/html/2608.28605#bib.bib6); Wanget al\.,[2017](https://arxiv.org/html/2608.28605#bib.bib5)\), InceptionTime employs parallel convolutional filters with different temporal spans, allowing it to robustly capture pathological patterns occurring at diverse time scales, which is particularly important for heterogeneous medical signals\. Moreover, InceptionTime incorporates bottleneck convolutions to reduce computational complexity and residual connections to facilitate stable optimization, making it well suited for long medical time series and data\-limited clinical settings\. For brevity, we present a simplified formulation of the InceptionTime encoder, focusing on its multi\-scale integration: \(1\)𝐡=GAP\(σ\(Concatk∈𝒦∪\{p\}\(𝐙k\)\+ϕ\(𝐗\)\)\)\\mathbf\{h\}=\\text\{GAP\}\\left\(\\sigma\\left\(\\text\{Concat\}\_\{k\\in\\mathcal\{K\}\\cup\\\{p\\\}\}\(\\mathbf\{Z\}\_\{k\}\)\+\\phi\(\\mathbf\{X\}\)\\right\)\\right\)whereGAP\(⋅\)\\text\{GAP\}\(\\cdot\)denotes the global average pooling;σ\(⋅\)\\sigma\(\\cdot\)refers to ReLU;𝒦\\mathcal\{K\}represents the set of parallel convolutional kernel sizes;𝐙k=Convk\(Conv1\(𝐗\)\)\\mathbf\{Z\}\_\{k\}=\\text\{Conv\}\_\{k\}\(\\text\{Conv\}\_\{1\}\(\\mathbf\{X\}\)\)is the feature map generated by thekk\-th scale branch, while𝐙p=Conv1\(MaxPool\(𝐗\)\)\\mathbf\{Z\}\_\{p\}=\\text\{Conv\}\_\{1\}\(\\text\{MaxPool\}\(\\mathbf\{X\}\)\)signifies the max\-pooling branch;ϕ\(𝐗\)\\phi\(\\mathbf\{X\}\)denotes the residual connection\.𝐡ts∈ℝD=fMLP\(𝐡\)\\mathbf\{h\}\_\{ts\}\\in\\mathbb\{R\}^\{D\}=f\_\{\\text\{MLP\}\}\(\\mathbf\{h\}\)denotes the final embedding of temporal pathway\. Although InceptionTime can aggregate local features into a global representation, such aggregation is intrinsically driven by localized temporal patterns, leaving global structural morphology under\-modeled at the temporal pathway, thereby motivating a complementary visual pathway\. ### 3\.2\.Transformer\-based Visual Pathway The visual pathway is designed to mirror and deepen clinical visual inspection to model global spectral\-spatial rhythms\. CWT\-based Visual Representation\.Converting time series into images has gained increasing attention in recent time\-series studies, including line plots\(Liet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib116)\), heatmaps\(Chenet al\.,[2025a](https://arxiv.org/html/2608.28605#bib.bib115)\), and spectrograms derived from the Short\-Time Fourier Transform \(STFT\)\([Dixitet al\.,](https://arxiv.org/html/2608.28605#bib.bib133)\)and the Continuous Wavelet Transform \(CWT\)\(Almanza\-Conejoet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib11)\)\. In practice, clinicians primarily rely on 1D tracings \(i\.e\., line plots\) to identify physiological patterns\. However, subtle pathological signals are often obscured by complex temporal fluctuations\. While STFT\-based spectrograms are constrained by their fixed time–frequency resolution, limiting its ability to capture global spatiotemporal correlations, the Continuous Wavelet Transform \(CWT\) inherently provides a multi\-resolution analysis\. This property allows CWT to preserve global structural information while effectively capturing transient spectral variations\. In this paper, we apply CWT to process the raw time series data\. Specifically, for a multivariate time series sample𝐗\\mathbf\{X\}withCCchannels, we first apply CWT to each channel independently, producing a set of scalograms𝒱i=CWT\(𝐗\)=\{vi,1,vi,2,…,vi,C\}\\mathcal\{V\}\_\{i\}=CWT\(\\mathbf\{X\}\)=\\\{v\_\{i,1\},v\_\{i,2\},\\dots,v\_\{i,C\}\\\}, where eachvi,c∈ℝh×wv\_\{i,c\}\\in\\mathbb\{R\}^\{h\\times w\}denotes the time\-frequency representation of thecc\-th channel\. Grid\-based Image Creation\.Furthermore, inspired by\(Liet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib116)\), we arrange the scalograms of all channels into a structured grid following the standard clinical ordering of leads\. This layout facilitates the modeling of inter\-channel correlations, closely mirroring the way clinicians cross\-reference diagnostic patterns across multiple leads\. Specifically, each scalogram is first rendered as a three\-channel image by applying a fixed colormap\. The rendered scalogram images are then arranged into a single grid\-structured image𝐕i=𝒢\(𝒱i\)\\mathbf\{V\}\_\{i\}=\\mathcal\{G\}\(\\mathcal\{V\}\_\{i\}\)using a predefined grid layout\. By default, we adopt a square grid and organize theCCchannel\-wise scalograms into a grid of sizel×ll\\times lwhen\(l−1\)2<C≤l2\(l\-1\)^\{2\}<C\\leq l^\{2\}, with unused grid positions left empty\. To ensure structural consistency across samples, the grid layout and channel ordering are fixed within each dataset\. This grid\-based construction preserves channel\-wise independence while enabling cross\-channel interactions to be implicitly modeled by visual backbones\. To summarize, the CWT\-based grid visual representation𝐕\\mathbf\{V\}is constructed as follows: \(2\)𝐕=𝒢\(CWT\(𝐗\)\)\\mathbf\{V\}=\\mathcal\{G\}\(\\text\{CWT\}\(\\mathbf\{X\}\)\)where𝒢\(⋅\)\\mathcal\{G\}\(\\cdot\)is grid mapping operation\.𝐗\\mathbf\{X\}is the raw multi\-channel time series and𝐕∈ℝH×W×3\\mathbf\{V\}\\in\\mathbb\{R\}^\{H\\times W\\times 3\}is its 2D grid image\. Swin Transformer as Backbone\.To extract latent embeddings from the 2D scalograms, we employ the Swin Transformer\(Liuet al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib100)\)as vision backbone\. Unlike conventional CNNs\(Heet al\.,[2016](https://arxiv.org/html/2608.28605#bib.bib6); Huanget al\.,[2016](https://arxiv.org/html/2608.28605#bib.bib98)\)that primarily capture local patterns or standard Vision Transformers\(Hanet al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib99); Dosovitskiyet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib103)\)that suffer from quadratic complexity, the Swin Transformer provides a hierarchical architecture with shifted window partitioning, aligning with the intrinsic properties of CWT scalograms\. First, its progressive downsampling produces hierarchical feature maps, supporting multi\-scale modeling of both transient and macro\-level patterns\. Second, its local window multi\-head self\-attention \(W\-MSA\) ensures efficient computation, while its shifted version \(SW\-MSA\) enables cross\-window interactions to capture long\-range spatial\-frequency correlations across the CWT grid\. Together, W\-MSA and SW\-MSA form consecutive Swin Transformer blocks: \(3\)𝐡¯l=W\-MSA\(LN\(𝐡l−1\)\)\+𝐡l−1,\\displaystyle\\bar\{\\mathbf\{h\}\}^\{l\}=\\text\{W\-MSA\}\(\\text\{LN\}\(\\mathbf\{h\}^\{l\-1\}\)\)\+\\mathbf\{h\}^\{l\-1\},𝐡l=fMLP\(LN\(𝐡¯l\)\)\+𝐡¯l,\\displaystyle\\mathbf\{h\}^\{l\}=f\_\{\\text\{MLP\}\}\(\\text\{LN\}\(\\bar\{\\mathbf\{h\}\}^\{l\}\)\)\+\\bar\{\\mathbf\{h\}\}^\{l\},𝐡¯l\+1=SW\-MSA\(LN\(𝐡l\)\)\+𝐡l,\\displaystyle\\bar\{\\mathbf\{h\}\}^\{l\+1\}=\\text\{SW\-MSA\}\(\\text\{LN\}\(\\mathbf\{h\}^\{l\}\)\)\+\\mathbf\{h\}^\{l\},𝐡l\+1=fMLP\(LN\(𝐡¯l\+1\)\)\+𝐡¯l\+1\\displaystyle\\mathbf\{h\}^\{l\+1\}=f\_\{\\text\{MLP\}\}\(\\text\{LN\}\(\\bar\{\\mathbf\{h\}\}^\{l\+1\}\)\)\+\\bar\{\\mathbf\{h\}\}^\{l\+1\}wherellis the block index\.LNandfMLPf\_\{\\text\{MLP\}\}denote normalization and fully\-connected layers\. Finally, the pooled visual features derived from Swin Transformer are projected into a lower\-dimensional latent space and normalized as𝐡img∈ℝD\\mathbf\{h\}\_\{\\text\{img\}\}\\in\\mathbb\{R\}^\{D\}\. Notably, MedTVL is flexible, supporting alternative transformer\-based vision backbones\. ### 3\.3\.Adaptive Textual Guidance In clinical practice, medical text is often used as auxiliary context to disambiguate numerical signals and imaging findings\(Chanet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib77)\)\. Motivated by this, we leverage medical text to guide dual\-pathway learning, aligning representations with clinical semantics and mitigating overfitting to superficial patterns\. However, due to the heterogeneous distributions of temporal and visual modalities, directly sharing a single textual embedding may cause feature degradation\. Thus, we propose adaptive textual guidance that projects shared text into modality\-specific subspaces and integrates it into each pathway via gated fusion, enabling modality\-aware constraints\. Specifically, a frozen medical language model encodes the textssinto a shared embedding,𝐡txt=fLM\(s\)\\mathbf\{h\}\_\{\\text\{txt\}\}=f\_\{\\text\{LM\}\}\(s\)\. This embedding is projected into pathway\-specific subspaces,𝐡txtp∈ℝD=fMLPp\(LN\(𝐡txt\)\)\\mathbf\{h\}^\{p\}\_\{\\text\{txt\}\}\\in\\mathbb\{R\}^\{D\}=f\_\{\\text\{MLP\}\}^\{p\}\(\\text\{LN\}\(\\mathbf\{h\}\_\{\\text\{txt\}\}\)\),p∈\{ts,img\}p\\in\\\{\\text\{ts\},\\text\{img\}\\\}\. To selectively integrate textual evidence, a adaptive gating vector is computed as𝐠∈ℝD=σ\(𝐖g\[𝐡p;𝐡txtp\]\)\\mathbf\{g\}\\in\\mathbb\{R\}^\{D\}=\\sigma\(\\mathbf\{W\}\_\{g\}\[\\mathbf\{h\}\_\{p\};\\mathbf\{h\}^\{p\}\_\{\\text\{txt\}\}\]\), and the enhanced representation𝐡¯p\\bar\{\\mathbf\{h\}\}\_\{p\}is obtained via weighted fusion: \(4\)𝐡¯p=\(1−𝐠\)⊙𝐡p\+𝐠⊙𝐡txtp\\bar\{\\mathbf\{h\}\}\_\{p\}=\(1\-\\mathbf\{g\}\)\\odot\\mathbf\{h\}\_\{p\}\+\\mathbf\{g\}\\odot\\mathbf\{h\}^\{p\}\_\{\\text\{txt\}\}where⊙\\odotdenotes element\-wise multiplication\. For temporal pathway𝐡¯p=𝐡¯ts\\bar\{\\mathbf\{h\}\}\_\{p\}=\\bar\{\\mathbf\{h\}\}\_\{\\text\{ts\}\}and for visual pathway𝐡¯p=𝐡¯img\\bar\{\\mathbf\{h\}\}\_\{p\}=\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}\. By injecting pathway\-specific clinical semantics, the module provides enriched representations for the subsequent MoE fusion, facilitating instance\-level expert selection\. ### 3\.4\.MoE\-based Instance\-Adaptive Fusion The inherent diversity of medical samples poses a significant challenge for dual\-pathway fusion, as diagnostic cues often prioritize either temporal dynamics or morphological patterns, depending on the underlying pathology\(Yunet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib13)\)\. While simple concatenation is static and gated fusion relies on shared weights, these methods may struggle to capture complex interactions across heterogeneous temporal and visual features, potentially leading to suboptimal integration in MedTS\. Mixture of Experts \(MoE\)\(Mu and Lin,[2025](https://arxiv.org/html/2608.28605#bib.bib14)\)has recently been introduced into time\-series modeling to enable fine\-grained adaptive specialization via parameter\-level expert selection\(Liuet al\.,[2025a](https://arxiv.org/html/2608.28605#bib.bib16),[b](https://arxiv.org/html/2608.28605#bib.bib15)\)\. By routing individual instances to specialized sub\-networks, MoE provides a more expressive mechanism for heterogeneous fusion, allowing diverse fusion patterns to be modeled by distinct expert parameters\. Following\(Shiet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib17)\), we reconfigure a shared pool of sparsely activated experts for instance\-adaptive pathway\-level fusion\. Specifically, we first apply modality\-specific projection to align heterogeneous pathways into a unified space:𝐳img=fMLPimg\(LN\(𝐡¯img\)\)\\mathbf\{z\}\_\{\\text\{img\}\}=f\_\{\\text\{MLP\}\}^\{\\text\{img\}\}\(\\text\{LN\}\(\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}\)\)and𝐳ts=fMLPts\(LN\(𝐡¯ts\)\)\\mathbf\{z\}\_\{\\text\{ts\}\}=f\_\{\\text\{MLP\}\}^\{\\text\{ts\}\}\(\\text\{LN\}\(\\bar\{\\mathbf\{h\}\}\_\{\\text\{ts\}\}\)\)\. The projected representations are then concatenated as a joint embedding𝐳=\[𝐳img;𝐳ts\]\\mathbf\{z\}=\[\\mathbf\{z\}\_\{\\text\{img\}\};\\mathbf\{z\}\_\{\\text\{ts\}\}\], which serves as the input to MoE fusion module\. This module consists of a shared expertEshared\(⋅\)\\text\{E\}\_\{\\text\{shared\}\}\(\\cdot\)for modeling global physiological patterns shared across instances andOOexpertsEi\(⋅\)i=1O\{\\text\{E\}\_\{i\}\(\\cdot\)\}\_\{i=1\}^\{O\}to for capturing instance\-specific variations\. In practice, each expert mirrors the architecture of a standard FFN\(Mu and Lin,[2025](https://arxiv.org/html/2608.28605#bib.bib14)\)\. A sigmoid gateσ\(Ws𝐳\)\\sigma\(W\_\{s\}\\mathbf\{z\}\)controls the contribution of the shared expert, and a sparse gating functionG\(𝐳\)\\text\{G\}\(\\mathbf\{z\}\)selects the top\-KKrouted experts\. The fused representation𝐳¯\\bar\{\\mathbf\{z\}\}is denoted as: \(5\)𝐳¯=σ\(Ws𝐳\)⋅Eshared\(𝐳\)\+∑k=1KG\(𝐳\)k⋅Ek\(𝐳\)\\bar\{\\mathbf\{z\}\}=\\sigma\(W\_\{s\}\\mathbf\{z\}\)\\cdot\\text\{E\}\_\{shared\}\(\\mathbf\{z\}\)\+\\sum\_\{k=1\}^\{K\}G\(\\mathbf\{z\}\)\_\{k\}\\cdot\\text\{E\}\_\{k\}\(\\mathbf\{z\}\)whereG\(𝐳\)=TopK\(Softmax\(Wg𝐳\)\)G\(\\mathbf\{z\}\)=\\text\{TopK\}\(\\text\{Softmax\}\(W\_\{g\}\\mathbf\{z\}\)\)represents the gating weights that route inputs to the topKKmost relevant specialized sub\-networks\. It effectively mitigates pathway conflicts and enables MedTVL to prioritize the most informative features based on the sample properties\. Finally, the fused representation𝐳¯\\bar\{\\mathbf\{z\}\}is fed into a MLP classifier to produce the final diagnostic predictions:y^=fMLP\(𝐳¯\)\\hat\{y\}=f\_\{\\text\{MLP\}\}\(\\bar\{\\mathbf\{z\}\}\)\. ### 3\.5\.Multimodal Contrastive Learning MedTVL’s dual pathways provide aligned multimodal views of the same physiological signals, namely temporal dynamics and visual morphology\. In this section, we describe how MedTVL is leveraged for multimodal contrastive learning\. Dual\-level Contrastive Loss\.To enhance robustness, we employ a dual\-level contrastive objective: an inter\-modality loss to harmonize cross\-modal semantics and an intra\-modality loss to preserve individual discriminability\. As detailed in Section[3\.3](https://arxiv.org/html/2608.28605#S3.SS3),𝐡¯tsi\\bar\{\\mathbf\{h\}\}\_\{\\text\{ts\}\}^\{i\}and𝐡¯imgi\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}^\{i\}denote the text\-enhanced embeddings of the temporal and visual pathways for sampleiiwithin a batch\. For the inter\-modality objective, we treat the cross\-pathway pair\(𝐡¯tsi,𝐡¯imgi\)\(\\bar\{\\mathbf\{h\}\}\_\{\\text\{ts\}\}^\{i\},\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}^\{i\}\)as a natural positive pair, while instances𝐡¯imgj\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}^\{j\}\(wherej≠ij\\neq i\) serve as cross\-modal negative samples\. For the intra\-modality objective, we construct positive pairs\(𝐡¯pi,𝐡¯p,augi\)\(\\bar\{\\mathbf\{h\}\}\_\{p\}^\{i\},\\bar\{\\mathbf\{h\}\}\_\{p,\\text\{aug\}\}^\{i\}\)using augmented views of the same modality and negative pairs \(𝐡¯pi,𝐡¯pj\\bar\{\\mathbf\{h\}\}\_\{p\}^\{i\},\\bar\{\\mathbf\{h\}\}\_\{p\}^\{j\}\) wherep∈\{ts,img\}p\\in\\\{ts,\\text\{img\}\\\}\. Weak stochastic augmentations are applied to generate positive pairs: jittering and scaling for the temporal pathway, and random cropping and flipping for the visual pathway\. The total contrastive lossℒcl\\mathcal\{L\}\_\{cl\}is defined as: \(6\)ℒcl=ℒinter\(𝐡¯ts,𝐡¯img\)\+λ∑p∈\{ts,img\}ℒintra\(𝐡¯p,𝐡¯p,aug\)\\mathcal\{L\}\_\{cl\}=\\mathcal\{L\}\_\{\\text\{inter\}\}\\\!\\left\(\\bar\{\\mathbf\{h\}\}\_\{\\text\{ts\}\},\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}\\right\)\+\\lambda\\sum\_\{p\\in\\\{\\text\{ts\},\\text\{img\}\\\}\}\\mathcal\{L\}\_\{\\text\{intra\}\}\\\!\\left\(\\bar\{\\mathbf\{h\}\}\_\{p\},\\bar\{\\mathbf\{h\}\}\_\{p,\\text\{aug\}\}\\right\)where we employ the CLIP contrastive loss for both inter\- and intra\-modality objectives, andλ\\lambdais a trade\-off coefficient balancing the two objectives and set as 0\.5 in our paper\. The CLIP contrastive loss is defined as follows: \(7\)sim\(𝐡it,𝐡jv\)=𝐡it⋅𝐡jvτ\\mathrm\{sim\}\(\\mathbf\{h\}^\{t\}\_\{i\},\\mathbf\{h\}^\{v\}\_\{j\}\)=\\frac\{\\mathbf\{h\}^\{t\}\_\{i\}\\cdot\\mathbf\{h\}^\{v\}\_\{j\}\}\{\\tau\} \(8\)ℒt→v=−1B∑i=1Blogexp\(sim\(𝐡it,𝐡iv\)\)∑j=1Bexp\(sim\(𝐡it,𝐡jv\)\)\\mathcal\{L\}\_\{t\\rightarrow v\}=\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\log\\frac\{\\exp\(\\mathrm\{sim\}\(\\mathbf\{h\}^\{t\}\_\{i\},\\mathbf\{h\}^\{v\}\_\{i\}\)\)\}\{\\sum\_\{j=1\}^\{B\}\\exp\(\\mathrm\{sim\}\(\\mathbf\{h\}^\{t\}\_\{i\},\\mathbf\{h\}^\{v\}\_\{j\}\)\)\} \(9\)ℒv→t=−1B∑i=1Blogexp\(sim\(𝐡iv,𝐡it\)\)∑j=1Bexp\(sim\(𝐡iv,𝐡jt\)\)\\mathcal\{L\}\_\{v\\rightarrow t\}=\-\\frac\{1\}\{B\}\\sum\_\{i=1\}^\{B\}\\log\\frac\{\\exp\(\\mathrm\{sim\}\(\\mathbf\{h\}^\{v\}\_\{i\},\\mathbf\{h\}^\{t\}\_\{i\}\)\)\}\{\\sum\_\{j=1\}^\{B\}\\exp\(\\mathrm\{sim\}\(\\mathbf\{h\}^\{v\}\_\{i\},\\mathbf\{h\}^\{t\}\_\{j\}\)\)\} \(10\)ℒCLIP=12\(ℒt→v\+ℒv→t\)\\mathcal\{L\}\_\{\\mathrm\{CLIP\}\}=\\frac\{1\}\{2\}\\Big\(\\mathcal\{L\}\_\{t\\rightarrow v\}\+\\mathcal\{L\}\_\{v\\rightarrow t\}\\Big\) whereBBdenotes the mini\-batch size,𝐡t\\mathbf\{h\}^\{t\}and𝐡v\\mathbf\{h\}^\{v\}areL2L\_\{2\}\-normalized embeddings of the sample pair andτ\\tauis a temperature hyperparameter that scales the logits\. The functionsim\(⋅,⋅\)\\text\{sim\}\(\\cdot,\\cdot\)denotes the cosine similarity\. In this framework,ℒinter\\mathcal\{L\}\_\{\\text\{inter\}\}encourages the alignment between the temporal and visual modalities by pulling𝐡¯ts\\bar\{\\mathbf\{h\}\}\_\{\\text\{ts\}\}and𝐡¯img\\bar\{\\mathbf\{h\}\}\_\{\\text\{img\}\}closer in the joint embedding space\. Conversely,ℒintra\\mathcal\{L\}\_\{\\text\{intra\}\}facilitates modality\-specific robustness by ensuring that the representations of raw data and its augmented versions remain consistent\. Pre\-training and Fine\-tuning\.At the pre\-training stage, we perform contrastive learning on the text\-enhanced outputs of the dual pathways while discarding the MoE module; this design prevents the model from taking shortcuts \(e\.g\., relying solely on features from one pathway\) to complete the contrastive task\. During fine\-tuning, following AimTS\(Chenet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib89)\), we use a standard classifier on the temporal pathway representations to adapt them to the downstream task\. This ensures a fair comparison with baselines that also use single\-modality features \. ## 4\.Experiments ### 4\.1\.Experimental Setup Datasets\.We conduct experiments on datasets spanning three clinical diagnostic scenarios\. \(1\)Alzheimer’s Disease: APAVA\(Escuderoet al\.,[2006](https://arxiv.org/html/2608.28605#bib.bib45)\)and ADFTD\(Miltiadouset al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib46)\)are two EEG datasets for Alzheimer’s disease classification\. \(2\)Epilepsy: TUSZ v1\.5\.2\(Shahet al\.,[2018](https://arxiv.org/html/2608.28605#bib.bib49)\)is a large\-scale EEG dataset for epilepsy\. TUSZ \(2\-Classes\) provides a coarse\-grained seizure/non\-seizure setting, while TUSZ \(4\-Classes\) provides a fine\-grained four\-class seizure taxonomy\. \(3\)Cardiac Disease: PTB\(PhysioBank,[2000](https://arxiv.org/html/2608.28605#bib.bib47)\)and PTB\-XL\(Wagneret al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib48)\)are two large\-scale ECG databases for cardiac diagnosis\. PTB\-XL \(4\-Classes\) corresponds to a coarse\-grained 4\-class setting, while PTB\-XL \(5\-Classes\) corresponds to a fine\-grained 5\-class setting\. Table[1](https://arxiv.org/html/2608.28605#S4.T1)shows data details\. Table 1\.Datasets Statistics\.DatasetsTotal SamplesClassesChannelsStepsAPAVA \(2\-Classes\)5,967216256ADFTD \(3\-Classes\)69,752319256TUSZ \(2\-Classes\)22,0402196,000TUSZ \(4\-Classes\)2,8914196,000PTB \(2\-Classes\)64,356215300PTB\-XL \(4\-Classes\)17,1104121,000PTB\-XL \(5\-Classes\)17,1105121,000 Table 2\.Supervised Learning\.Red: best,Blue: second best\. Modalities are denoted as:Tfor time\-series signals,Vfor visual representations, andLfor language\.Methods\(Modality\)APAVA \(2\-Classes\)ADFTD \(3\-Classes\)TUSZ \(4\-Classes\)PTB\-XL \(4\-Classes\)PTB \(2\-Classes\)F1Acc\.F1Acc\.F1Acc\.F1Acc\.F1Acc\.Dlinear \(T\)56\.19±\\pm1\.2765\.48±\\pm0\.3339\.58±\\pm0\.8446\.96±\\pm2\.1170\.70±\\pm0\.3680\.14±\\pm0\.8833\.71±\\pm2\.7957\.01±\\pm2\.6862\.78±\\pm0\.6974\.42±\\pm0\.51MedGNN \(T\)80\.78±\\pm2\.9581\.91±\\pm1\.8145\.42±\\pm1\.0747\.71±\\pm2\.6283\.06±\\pm0\.7690\.16±\\pm2\.4666\.25±\\pm2\.2477\.21±\\pm0\.7480\.58±\\pm2\.3584\.36±\\pm2\.96MultiRocket \(T\)58\.81±\\pm1\.6358\.84±\\pm1\.5036\.59±\\pm2\.8351\.32±\\pm2\.3775\.55±\\pm2\.4684\.28±\\pm1\.0850\.93±\\pm2\.1963\.35±\\pm0\.8264\.11±\\pm1\.0674\.78±\\pm2\.40ResNet \(T\)69\.67±\\pm2\.3572\.26±\\pm1\.7640\.41±\\pm1\.0545\.91±\\pm2\.3078\.35±\\pm2\.9386\.36±\\pm2\.5863\.79±\\pm0\.6477\.25±\\pm0\.4968\.14±\\pm2\.4975\.97±\\pm2\.87InceptionTime \(T\)78\.36±\\pm2\.9080\.92±\\pm2\.0245\.79±\\pm1\.3553\.19±\\pm1\.9582\.95±\\pm2\.2688\.77±\\pm0\.4368\.71±\\pm2\.8477\.21±\\pm2\.6178\.91±\\pm2\.3783\.56±\\pm2\.29TimesNet \(T\)73\.12±\\pm2\.8775\.75±\\pm1\.7446\.09±\\pm0\.6750\.43±\\pm2\.3983\.86±\\pm2\.0388\.43±\\pm2\.6664\.13±\\pm2\.9475\.35±\\pm2\.0073\.11±\\pm2\.3279\.23±\\pm2\.83PatchTST \(T\)61\.93±\\pm2\.9467\.99±\\pm2\.0540\.41±\\pm1\.0943\.24±\\pm1\.8181\.03±\\pm2\.3281\.87±\\pm1\.3157\.97±\\pm1\.4174\.14±\\pm2\.4972\.77±\\pm1\.8280\.27±\\pm0\.78iTransformer \(T\)74\.31±\\pm1\.1776\.11±\\pm2\.1341\.57±\\pm2\.7345\.41±\\pm1\.9282\.27±\\pm0\.6687\.56±\\pm1\.0661\.41±\\pm2\.2874\.01±\\pm0\.3975\.52±\\pm2\.9082\.01±\\pm1\.40Medformer \(T\)72\.74±\\pm0\.6176\.59±\\pm2\.8146\.23±\\pm2\.8453\.79±\\pm1\.2884\.27±\\pm2\.5788\.77±\\pm2\.4862\.26±\\pm2\.1176\.88±\\pm2\.0977\.37±\\pm2\.3682\.41±\\pm2\.98ViTST \(V\)80\.93±\\pm0\.7881\.97±\\pm2\.3341\.57±\\pm2\.8945\.41±\\pm1\.3982\.74±\\pm2\.0588\.08±\\pm1\.0465\.22±\\pm2\.9975\.58±\\pm2\.6482\.62±\\pm2\.3185\.85±\\pm2\.48MedViA \(TV\)81\.26±\\pm2\.9683\.44±\\pm1\.8247\.85±\\pm1\.0851\.71±\\pm2\.6586\.51±\\pm0\.7889\.12±\\pm2\.4867\.89±\\pm2\.2579\.75±\\pm0\.7683\.93±\\pm2\.3686\.72±\\pm2\.99GPT4TS \(T\)77\.92±\\pm0\.9480\.92±\\pm1\.8343\.72±\\pm2\.5451\.51±\\pm2\.4483\.35±\\pm1\.8491\.19±\\pm1\.3363\.11±\\pm2\.8674\.98±\\pm0\.7976\.82±\\pm2\.3382\.32±\\pm1\.64MedTsLLM \(TL\)78\.24±\\pm2\.9978\.76±\\pm1\.8745\.83±\\pm1\.4152\.59±\\pm2\.6383\.95±\\pm2\.2089\.29±\\pm2\.4964\.21±\\pm2\.8776\.41±\\pm0\.7775\.66±\\pm2\.3681\.31±\\pm2\.97MedualTime \(TL\)79\.85±\\pm2\.9481\.62±\\pm1\.8346\.95±\\pm1\.1150\.77±\\pm2\.6485\.37±\\pm0\.8390\.51±\\pm2\.4772\.53±\\pm2\.2780\.58±\\pm0\.7577\.68±\\pm2\.3881\.78±\\pm2\.98TimeVLM \(TVL\)79\.93±\\pm2\.9882\.39±\\pm1\.8448\.35±\\pm1\.1353\.84±\\pm2\.6685\.72±\\pm0\.8190\.85±\\pm2\.4970\.29±\\pm2\.2880\.31±\\pm0\.7884\.71±\\pm2\.3787\.48±\\pm2\.99MedTVL \(TVL\)85\.52±\\pm2\.9785\.88±\\pm1\.8551\.27±\\pm1\.1554\.12±\\pm2\.6789\.26±\\pm0\.8291\.19±\\pm2\.5077\.93±\\pm2\.2984\.43±\\pm0\.7988\.22±\\pm2\.3990\.04±\\pm2\.99 Baselines\.For a comprehensive comparison, we choose representative baselines from diverse architectures and modalities as follows:\(1\) MLP: DLinear\(Zenget al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib51)\);\(2\) GNN: MedGNN\(Fanet al\.,[2025a](https://arxiv.org/html/2608.28605#bib.bib76)\);\(3\) CNN: ResNet\(Heet al\.,[2016](https://arxiv.org/html/2608.28605#bib.bib6)\), MultiRocket\(Tanet al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib62)\), InceptionTime\(Ismail Fawazet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib8)\)and TimesNet\(Wuet al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib53)\);\(4\) Transformer: PatchTST\(Nieet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib52)\), Medformer\(Wanget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib67)\), iTransformer\(Liuet al\.,[2024b](https://arxiv.org/html/2608.28605#bib.bib32)\), ViTST\(Liet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib116)\)and MedViA\(Fanet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib125)\);\(5\) LLM/VLM: GPT4TS\(Zhouet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib71)\), MedTsLLM\(Chanet al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib77)\), MedualTime\(Yeet al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib78)\), and Time\-VLM\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\)\. Our contrastive learning baselines include T\-Loss\(Franceschiet al\.,[2019](https://arxiv.org/html/2608.28605#bib.bib91)\), TS\-TCC\(Eldeleet al\.,[2021](https://arxiv.org/html/2608.28605#bib.bib92)\), TS2Vec\(Yueet al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib94)\), InfoTS\(Luoet al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib93)\), TimesURL\(Liu and Chen,[2024](https://arxiv.org/html/2608.28605#bib.bib90)\), and AimTS\(Chenet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib89)\)\. Implementation Details\.Following\(Fanet al\.,[2025b](https://arxiv.org/html/2608.28605#bib.bib125)\), we report macro\-averaged F1, AUROC, AUPRC, accuracy, precision, and recall\. We prioritize F1 score due to data imbalance while other metrics deferred to theAppendix\. For a fair comparison and to maintain consistent experimental conditions, all models are integrated into a unified implementation framework\. All models are trained with a batch size of 32 using the AdamW optimizer and a cosine scheduler for up to 50 epochs, supplemented by an early stopping patience of 7\. To ensure statistical reliability, we report the mean and standard deviation over 6 independent runs \(random seeds 42\-47\) on fixed dataset splits\. The optimal model is selected based on the highest validation F1\-score and subsequently evaluated on the test set\. For MedTVL, the Swin\-Base architecture \(patch size 4, window size 7, image size 224×\\times224\) pre\-trained on ImageNet\-21k serves as the visual backbone\. We fine\-tune only the first two encoder stagesS=\(0,1\)S=\(0,1\)to balance efficacy and efficiency, refining low\-level medical textures while preserving high\-level pre\-trained priors\. InceptionTime is adopted as the temporal backbone following the original configuration\(Ismail Fawazet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib8)\), featuring a depth of 6, multi\-scale kernels\[39,19,9\]\[39,19,9\], and filter channelsnf=32nf=32\. ClinicalBERT\(Wanget al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib129)\)is employed as a frozen language model\. As to MoE module, we set expert numberO=4O=4and topK=2K=2routed experts to handle the pathway heterogeneity\. The dimension of hidden embedding is128128and output embedding isD=64D=64\. We optimize baselines by tuning critical hyper\-parameters\. All experiments are using PyTorch on NVIDIA A60006000\(4848GB\) GPU\. ### 4\.2\.Supervised Learning Key findings in Table[2](https://arxiv.org/html/2608.28605#S4.T2)include: \(1\) Classical linear or shallow models such as DLinear and MultiRocket usually underperform across all datasets\. \(2\) Among CNN\-based methods, InceptionTime usually achieves the best or second\-best performance, particularly on ECG signals\. \(3\) Multi\-modal approaches generally outperform single\-modality baselines across most datasets, indicating that incorporating cross\-modal information improves robustness\. \(4\) MedTVL consistently outperforms competing methods across datasets and metrics\. Notably, it achieves an improvement of approximately3%–5%in average F1 over the second\-best models, highlighting the effectiveness of the proposed tri\-modal framework and the text\-assisted dual\-pathway design\. Figure 3\.Few\-shot performance in F1 score under different shot settings\. Other metrics are inAppendixLABEL:abl\_app\.Odenotes the source domain andTdenotes the target domain\. ### 4\.3\.Few\-shot Learning We conduct few\-shot learning to alleviate the label scarcity challenge by transferring knowledge from a label\-rich source domain to a target domain with restricted annotations\. Specifically, models are first pre\-trained on the source dataset and subsequently fine\-tuned on the target dataset under variouskk\-shot settings \(k∈\{5,15,25,35,45,55\}k\\in\\\{5,15,25,35,45,55\\\}\)\. Due to the rigid input dimension requirements of our baselines, we select source\-target dataset pairs with identical timesteps and channels, namely PTB\-XL \(4\-Classes\) and PTB\-XL \(5\-Classes\), TUSZ \(2\-Classes\) and TUSZ \(4\-Classes\)\. During the fine\-tuning phase, the pre\-trained backbones remain frozen, and only a newly initialized classification head is optimized\. As shown in Figure[3](https://arxiv.org/html/2608.28605#S4.F3): \(1\) Model performance generally scales positively with the number of shots, with our model achieving state\-of\-the\-art results in nearly all settings, demonstrating robust cross\-dataset transferability\. \(2\) Notably, our model exhibits sustained performance gains as the number of shots increases, suggesting that it can effectively internalize additional information without reaching an early plateau\. \(3\) Transferring from fine\-grained to coarse\-grained label sets yields superior results than the reverse\. This suggests that fine\-grained pre\-training compels the model to learn more nuanced, discriminative features for handling downstream tasks\. Table 3\.Contrastive learning via linear probing on 100% labeled set in F1 score\.APAVA\(2\-Classes\)ADFTD\(3\-Classes\)TUSZ\(4\-Classes\)PTB\-XL\(4\-Classes\)PTB\(2\-Classes\)T\-Loss \(T\)49\.44±\\pm1\.1236\.72±\\pm3\.8564\.03±\\pm1\.3433\.87±\\pm3\.6250\.88±\\pm1\.05TSTCC \(T\)57\.88±\\pm1\.0737\.55±\\pm3\.6871\.35±\\pm1\.2948\.15±\\pm3\.4763\.51±\\pm1\.13TS2Vec \(T\)62\.33±\\pm1\.0239\.26±\\pm3\.5373\.58±\\pm1\.2449\.01±\\pm3\.3264\.65±\\pm1\.18InfoTS \(T\)64\.73±\\pm0\.9743\.72±\\pm3\.3772\.22±\\pm1\.1953\.38±\\pm3\.1964\.11±\\pm1\.23TimesURL \(T\)62\.86±\\pm1\.0543\.40±\\pm3\.4175\.91±\\pm1\.2159\.29±\\pm3\.2566\.41±\\pm1\.16AimTS \(TV\)72\.65±\\pm0\.9144\.16±\\pm3\.2374\.77±\\pm1\.1460\.81±\\pm3\.0871\.32±\\pm1\.19MedTVL \(TVL\)76\.20±\\pm0\.8645\.16±\\pm3\.1179\.18±\\pm1\.0967\.53±\\pm2\.9774\.56±\\pm1\.14 Figure 4\.Contrastive learning via linear probing on 10% labeled set\. 1=T\-Loss, 2=TSTCC, 3=TS2Vec, 4=InfoTS, 5=TimesURL, 6=AimTS, 7=MedTVL\.Figure 5\.Contrastive learning via linear probing on 50% labeled set\. 1=T\-Loss, 2=TSTCC, 3=TS2Vec, 4=InfoTS, 5=TimesURL, 6=AimTS, 7=MedTVL\. ### 4\.4\.Contrastive Learning We first pre\-train our framework to produce unsupervised embeddings, and then evaluate their quality and label efficiency by training a linear classifier under both full\-resource \(100%100\\%labeled training data\) and data\-sparse \(5%5\\%and50%50\\%labeled training data\) settings, while keeping the validation and test sets fixed across all experiments\. The results for the 100% data setting are summarized in Table[3](https://arxiv.org/html/2608.28605#S4.T3)while Figure[4](https://arxiv.org/html/2608.28605#S4.F4)and[5](https://arxiv.org/html/2608.28605#S4.F5)show results under 10% and 50% labeled data settings\. Table[3](https://arxiv.org/html/2608.28605#S4.T3)shows that multimodal models consistently outperform time\-only baselines, highlighting the synergy between visual and temporal signal\. While AimTS shows strong transferability, our MedTVL achieves the best results across all metrics\. Notably, MedTVL surpasses the strongest baseline by 6\.72% in PTB\-XL \(4\-Classes\)\. This consistent lead demonstrates that tri\-modal fusion provides richer information density, enabling the backbone to learn more discriminative representations\. ### 4\.5\.More Experiments \(a\) Ablation Study\.To evaluate the impact of critical modules, we compare five ablation variants: ”w/o MoE \(MLP\)” replaces MoE with a standard MLP; ”w/o MoE \(Gated Fusion\)” replaces MoE with a gated fusion mechanism\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\); ”w/o Textual Guidance” removes the Adaptive Textual Guidance module; ”w/o Visual Pathway” and ”w/o Temporal Pathway” respectively exclude the visual and temporal branches\. Table[4](https://arxiv.org/html/2608.28605#S4.T4)shows that the dual pathways emerge as the most fundamental component\. The removal of the temporal pathway yields the steepest F1 decline on PTB\-XL \(about 10\.4%\), while excluding the visual pathway results in an average F1 drop of about 5% on TUSZ and APAVA\. This suggests that different medical datasets rely on distinct modal features\. When the MoE module is removed, gated fusion outperforms MLP\-based fusion, likely because it can assign different weights to modal features\. However, these weights are shared across samples rather than dynamically adapted to individual instances\. The MoE module and textual guidance further enhance the model, yielding approximately 3% and 2% F1 improvements, respectively\. These results validate the synergy of our tri\-modal design and the MoE\-based fusion in capturing sample\-wise physiological patterns\. Table 4\.Ablation study in F1 score on three datasets\.VariantsAPAVA\(2\-Classes\)TUSZ\(4\-Classes\)PTB\-XL\(4\-Classes\)w/o MoE \(MLP\)82\.60±\\pm1\.2386\.67±\\pm2\.5874\.18±\\pm3\.12w/o MoE \(Gated Fusion\)83\.33±\\pm1\.8787\.28±\\pm2\.4575\.77±\\pm2\.91w/o Textual Guidance83\.35±\\pm2\.1186\.56±\\pm3\.8775\.64±\\pm1\.45w/o Visual Pathway80\.43±\\pm1\.7884\.27±\\pm2\.9870\.58±\\pm3\.45w/o Temporal Pathway81\.11±\\pm2\.5685\.19±\\pm1\.2367\.53±\\pm0\.87MedTVL85\.52±\\pm2\.9789\.26±\\pm0\.8277\.93±\\pm2\.29 Figure 6\.Efficiency comparison on TUSZ \(4\-Classes\), whereMdenotes millions of trainable parameters\.\(b\) Efficiency Analysis\.In Figure[6](https://arxiv.org/html/2608.28605#S4.F6), we compare the efficiency of MedTVL with representative baselines on the TUSZ \(4\- Classes\), considering training time per epoch, F1 score, and the number of trainable parameters\. \(1\) Among all methods, InceptionTime is the most lightweight and fastest model, requiring only 6 seconds per epoch with 0\.46M parameters\. However, its F1 score \(82\.95%\) remains within the mid\-range of all evaluated methods\. \(2\) PatchTST, ResNet, and iTransformer are also parameter\-efficient, but they achieve the lowest F1 scores, indicating limited modeling capacity for this task\. \(3\) TimesNet incurs the highest computational cost, requiring 57 seconds per epoch, yet its performance does not scale proportionally with its model size\. \(4\) In general, multimodal approaches tend to be slower than unimodal models due to the additional cross\-modal processing\. \(5\) In contrast, MedTVL achieves the highest F1 score, significantly outperforming all competing methods\. The trainable parameters of MedTVL are 3\.23M, which is larger than those of MedualTime, while remaining substantially more compact than LLM\-based alternatives such as GPT4TS and MedTsLLM\. Moreover, it runs faster than MedTsLLM, TimeVLM, and TimesNet\. \(6\) Overall, MedTVL strikes a favorable balance between effectiveness and efficiency, achieving superior performance with a manageable computational footprint\. Table 5\.Backbone analysis in F1 score\.Blue: second best in each module\.Backbone of ModuleAPAVA\(2\-Classes\)TUSZ\(4\-Classes\)PTB\-XL\(4\-Classes\)TemporalResNet \(CNN\)81\.29±\\pm1\.5686\.09±\\pm1\.4574\.65±\\pm1\.67TCN\(Bai,[2018](https://arxiv.org/html/2608.28605#bib.bib101)\)\(CNN\)82\.03±\\pm0\.8985\.41±\\pm1\.9874\.37±\\pm1\.89Medformer \(Transformer\)80\.60±\\pm1\.9884\.66±\\pm2\.1372\.17±\\pm2\.34PatchTST \(Transformer\)79\.75±\\pm2\.4585\.19±\\pm1\.6773\.38±\\pm1\.56VisualViT\(Dosovitskiyet al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib103)\)\(Transformer\)82\.59±\\pm1\.2384\.17±\\pm2\.3475\.17±\\pm1\.12ConvNeXts\(Liuet al\.,[2022](https://arxiv.org/html/2608.28605#bib.bib102)\)\(CNN\)79\.71±\\pm2\.7880\.76±\\pm2\.5669\.65±\\pm2\.56ResNet \(CNN\)78\.53±\\pm1\.8781\.53±\\pm1\.8971\.64±\\pm1\.98TextLLaMA \(LLM\)\(Touvron and Lavril,[2023](https://arxiv.org/html/2608.28605#bib.bib126)\)84\.92±\\pm2\.1588\.04±\\pm1\.7877\.01±\\pm1\.67GPT\-2 \(LLM\)\(Radfordet al\.,[2019](https://arxiv.org/html/2608.28605#bib.bib127)\)83\.54±\\pm1\.8787\.55±\\pm2\.9976\.71±\\pm2\.45MedTVL85\.52±\\pm2\.9789\.26±\\pm0\.8277\.93±\\pm2\.29 \(c\) Backbone Analysis\.In theory, the temporal, visual, and textual modules allow for multiple backbone instantiations\. Nevertheless, in practice, the specific backbone choices play a critical role in shaping the framework’s overall effectiveness\. To examine the impact of backbone selection, we replace the default backbones with representative CNN, Transformer and LLM architectures\. Table[5](https://arxiv.org/html/2608.28605#S4.T5)shows that different backbone choices lead to different performance trends\. CNN\-based backbones outperform Transformers in the temporal pathway, whereas the opposite holds true for the visual pathway\. This phenomenon suggests that peak performance stems from a strategic heterogeneous architecture pairing\. The inductive bias of CNNs is uniquely suited for capturing local, shift\-invariant pathological motifs in temporal signals, while the Transformer’s global modeling capability better aligns with the spatial structural correlations of 2D images\. Within this optimized dual\-stream framework, the selection of InceptionTime and Swin Transformer further enhances efficacy\. Both models leverage multi\-scale receptive fields to capture multi\-granular medical signatures and exhibit superior computational efficiency—driven by Inception’s parallel convolutions and Swin’s shifted\-window mechanism\. This combination ensures a robust balance between diagnostic accuracy and inference latency\. Additionally, the results of textual backbones show that general\-purpose language models \(e\.g\., GPT\-2 and LLaMA\) consistently degrades performance\. This highlights the advantage of domain\-specific clinical model \(ClinicalBERT\) in providing more effective semantic textual guidance\. Figure 7\.An example of different visual representations\.Figure 8\.An example of grid construction of different visual representations on APAVA\)\.Table 6\.Visual representation analysis\.ImageRepresentationAPAVA \(2\-Classes\)TUSZ \(4\-Classes\)PTB\-XL \(4\-Classes\)F1Acc\.F1Acc\.F1Acc\.Line Plot \(grid\)87\.11±\\pm0\.2887\.77±\\pm0\.2285\.00±\\pm0\.3289\.64±\\pm0\.2672\.08±\\pm0\.2580\.72±\\pm0\.18STFT \(grid\)82\.74±\\pm0\.2984\.49±\\pm0\.3387\.12±\\pm0\.2089\.98±\\pm0\.2276\.07±\\pm0\.3483\.73±\\pm0\.26Heatmap \(grid\)83\.02±\\pm0\.2383\.79±\\pm0\.2686\.28±\\pm0\.3589\.98±\\pm0\.2975\.01±\\pm0\.2081\.79±\\pm0\.31TimeVLM \(non\-grid\)82\.48±\\pm1\.5683\.02±\\pm1\.8885\.79±\\pm0\.8990\.50±\\pm0\.4774\.85±\\pm1\.9281\.46±\\pm1\.29MedTVL \(CWT\)85\.52±\\pm2\.9785\.88±\\pm1\.8589\.26±\\pm0\.8291\.19±\\pm2\.5077\.93±\\pm2\.2984\.43±\\pm0\.79 \(d\) Visual Representation Analysis\.We investigate the impact of different time\-series\-to\-image representations on three datasets by replacing the CWT spectrogram in our grid structure with line plots, STFT spectrograms, and heatmaps \(Figure[7](https://arxiv.org/html/2608.28605#S4.F7)and[8](https://arxiv.org/html/2608.28605#S4.F8)\)\. We further evaluate the role of grid construction by substituting the CWT\-based grid with the Frequency–Periodicity–Multi\-scale Convolution Encoding from TimeVLM\(Zhonget al\.,[2025](https://arxiv.org/html/2608.28605#bib.bib123)\), which does not adopt a grid layout and is denoted as“TimeVLM \(non\-grid\)”\. As shown in Table[6](https://arxiv.org/html/2608.28605#S4.T6), although CWT\-based MedTVL does not achieve the best F1 score on every dataset, it delivers the strongest average performance across benchmarks, indicating greater robustness to diverse data characteristics\. Line plots perform best on APAVA, likely because the relatively small number of time steps allows temporal details to be clearly represented in a line plot\. In contrast, the larger strides in TUSZ \(6000\) and PTB\-XL \(1000\) lead to substantial loss of local patterns when using line plots\. By explicitly modeling time–frequency structures and multi\-scale dynamics, CWT remains effective across datasets\. Moreover, grid\-based representations consistently outperform the non\-grid variant, highlighting the importance of structured temporal layouts\. Figure 9\.Sensitivity Analysis in F1 of two datasets\.\(e\) Sensitivity Analysis\.We study the influence of three key hyperparameters on three datasets—ADFTD \(3\-Classes\), TUSZ \(4\-Classes\), and PTB \(2\-Classes\): \(1\) tuning stages \(S\) of the visual backbone, with values \(\[\(0,1\),\(2,3\),all\]\[\(0,1\),\(2,3\),\\text\{all\}\]\); \(2\) filter channelsnfnfof the temporal backbone, with values \(\[8,16,32,64\]\[8,16,32,64\]\); and \(3\) expert numberOOin the MoE module, with values \(\[2,4,6,8\]\[2,4,6,8\]\)\. Figure[9](https://arxiv.org/html/2608.28605#S4.F9)presents F1 on two datasets\. For tuning stages, ADFTD achieves its best performance when only the first two stages are fine\-tuned, whereas PTB reaches its optimum when all stages are tuned\. Since full tuning incurs higher computational cost, tuning only the first two stages provides a favorable trade\-off between performance and efficiency\. Regarding the expert numberOO, both datasets attain peak performance at \(O=4O=4\)\. IncreasingOOfurther leads to performance degradation, likely because too many experts dilute the training data per expert and introduce redundancy, harming generalization\. For the filter channel numbernfnf, a value of 32 consistently yields the best results across datasets, indicating an appropriate balance between temporal modeling capacity and model complexity\. Smaller values limit expressiveness, while larger ones tend to introduce redundancy without clear performance gains\. Figure 10\.Visualization analysis of PTB\-XL\(4\-Classes\)\.\(f\) Visualization\.To provide an intuitive visualization of the representations learned under supervised training, we apply t\-SNE\(Maaten and Hinton,[2008](https://arxiv.org/html/2608.28605#bib.bib130)\)to project the embeddings of representative models on the PTB\-XL \(4\-Classes\) dataset into a two\-dimensional space, as shown in Figure[10](https://arxiv.org/html/2608.28605#S4.F10)\. The results illustrate a progressive improvement in representation separability across models\. TimesNet primarily distinguishes the two largest classes, while MedViA and MedualTime further enhance class separability, particularly for the third major class\. In contrast, MedTVL produces the most structured embedding space, with all classes forming well\-separated clusters\. ## 5\.Conclusion This work presents MedTVL, a tri\-modal framework for medical time series classification that integrates numerical signals, visual representations, and clinical language to better reflect clinical practice\. By adopting a heterogeneous dual\-pathway design, MedTVL leverages complementary inductive biases from convolutional and transformer\-based architectures to capture both fine\-grained temporal dynamics in numerical signals and global morphological patterns in time\-series\-derived images\. Furthermore, the dual\-pathway structure naturally enables cross\-modal positive pairing, facilitating multimodal contrastive learning under limited clinical annotations\. Extensive experiments across diverse tasks demonstrate the robustness, adaptability, and generalizability of MedTVL for clinical decision support\. ## 6\.Acknowledgements This work is funded by National Natural Science Foundation of China Grant No\. 72371217, NSFC 62572418 and Guangdong Provincial Talent Program \(No\.2024TQ08X366\), the Guangzhou Industrial Informatics and Intelligence Key Laboratory No\. 2024A03J0628, the Nansha Key Area Science and Technology Project No\. 2023ZD003, and Project No\. 2021JC02X191\. ## Limitations and Ethical Considerations This study uses publicly available, de\-identified medical datasets from prior research and does not involve the collection of new data from human participants\. As the data are fully anonymized, no additional Institutional Review Board \(IRB\) approval is required under our institutional guidelines\. All experiments comply with ACM Publications Policies, including those on research involving human participants\. MedTVL is intended solely as a research and decision\-support framework rather than a standalone diagnostic system\. Its predictions should be interpreted by qualified professionals and must not replace clinical judgment, as misuse without appropriate domain expertise may lead to incorrect decisions\. Despite strong performance and transferability across multiple datasets and learning paradigms, MedTVL remains subject to limitations stemming from dataset quality, representativeness, and potential biases\. While developed for medical applications, its general design principles may extend to other domains; any such use should undergo domain\-specific ethical review and regulatory oversight\. We acknowledge that additional ethical concerns may be raised by the community\. ## References - O\. Almanza\-Conejo, D\. L\. Almanza\-Ojeda, J\. L\. Contreras\-Hernandez, and M\. A\. Ibarra\-Manzano \(2023\)Emotion recognition in eeg signals using the continuous wavelet transform and cnns\.Neural Computing and Applications35\(2\),pp\. 1409–1422\.Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p2.5)\. - S\. Bai \(2018\)An empirical evaluation of generic convolutional and recurrent networks for sequence modeling\.arXiv preprint arXiv:1803\.01271\.Cited by:[Table 5](https://arxiv.org/html/2608.28605#S4.T5.6.6.6.4)\. - N\. Chan, F\. Parker, W\. Bennett, T\. Wu, M\. Y\. Jia, J\. Fackler, and K\. Ghobadi \(2024\)MedTsLLM: leveraging llms for multimodal medical time series analysis\.arXiv preprint arXiv:2408\.07773\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p2.1),[§3\.3](https://arxiv.org/html/2608.28605#S3.SS3.p1.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - M\. Chen, L\. Shen, Z\. Li, X\. J\. Wang, J\. Sun, and C\. Liu \(2025a\)VisionTS: visual masked autoencoders are free\-lunch zero\-shot time series forecasters\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\),Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p2.5)\. - Y\. Chen, S\. Huang, Y\. Cheng, P\. Chen, Z\. Rao, Y\. Shu, B\. Yang, L\. Pan, and C\. Guo \(2025b\)AimTS: augmented series and image contrastive learning for time series classification\.In2025 IEEE 41st International Conference on Data Engineering \(ICDE\),pp\. 1952–1965\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1),[§3\.5](https://arxiv.org/html/2608.28605#S3.SS5.p8.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - M\. Cheng, Y\. Chen, Q\. Liu, Z\. Liu, and Y\. Luo \(2024\)Advancing time series classification with multimodal language modeling\.arXiv preprint arXiv:2403\.12371\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p2.1)\. - Z\. Cui, W\. Chen, and Y\. Chen \(2016\)Multi\-scale convolutional neural networks for time series classification\.arXiv preprint arXiv:1603\.06995\.Cited by:[§3\.1](https://arxiv.org/html/2608.28605#S3.SS1.p1.1)\. - C\. Ding, T\. Yao, C\. Wu, and J\. Ni \(2025\)Advances in deep learning for personalized ecg diagnostics: a systematic review addressing inter\-patient variability and generalization constraints\.Biosensors and Bioelectronics271,pp\. 117073\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1)\. - \[9\]S\. Dixit, L\. Heller, and C\. DonahueVision language models are few\-shot audio spectrogram classifiers\.InAudio Imagination: NeurIPS 2024 Workshop AI\-Driven Speech, Music, and Sound Generation,Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p2.5)\. - A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly,et al\.\(2020\)An image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations,Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p5.5),[Table 5](https://arxiv.org/html/2608.28605#S4.T5.15.15.15.5)\. - E\. Eldele, M\. Ragab, Z\. Chen, M\. Wu, C\. K\. Kwoh, X\. Li, and C\. Guan \(2021\)Time\-series representation learning via temporal and contextual contrasting\.InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence,pp\. 2352–2359\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - J\. Escudero, D\. Abásolo, R\. Hornero, P\. Espino, and M\. López \(2006\)Analysis of electroencephalograms in alzheimer’s disease patients with multiscale entropy\.Physiological measurement27\(11\),pp\. 1091\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p1.1),[§7\.1\.1](https://arxiv.org/html/2608.28605#S7.SS1.SSS1.p1.1)\. - W\. Fan, J\. Fei, D\. Guo, K\. Yi, X\. Song, H\. Xiang, H\. Ye, and M\. Li \(2025a\)Towards multi\-resolution spatiotemporal graph learning for medical time series classification\.InProceedings of the ACM on Web Conference 2025,pp\. 5054–5064\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - W\. Fan, J\. Fei, J\. Han, J\. Lian, H\. Ye, X\. Song, X\. Lv, K\. Yi, and M\. Li \(2025b\)MedViA: empowering medical time series classification with vision augmentation and multimodal fusion\.Information Fusion,pp\. 103659\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p3.1),[§2](https://arxiv.org/html/2608.28605#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p3.10)\. - M\. Fatourechi, A\. Bashashati, R\. K\. Ward, and G\. E\. Birch \(2007\)EMG and eog artifacts in brain computer interface systems: a survey\.Clinical neurophysiology118\(3\),pp\. 480–494\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1)\. - J\. Franceschi, A\. Dieuleveut, and M\. Jaggi \(2019\)Unsupervised scalable representation learning for multivariate time series\.Advances in neural information processing systems32\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - S\. Gao, T\. Koker, O\. Queen, T\. Hartvigsen, T\. Tsiligkaridis, and M\. Zitnik \(2024\)Units: a unified multi\-task time series model\.Advances in Neural Information Processing Systems37,pp\. 140589–140631\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p2.1)\. - X\. Gu, Y\. Shu, J\. Han, Y\. Liu, Z\. Liu, J\. Anibal, V\. Sangha, E\. Phillips, B\. Segal, H\. Yuan,et al\.\(2025\)Foundation models for biosignals: a survey\.Authorea Preprints\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1)\. - K\. Han, A\. Xiao, E\. Wu, J\. Guo, C\. Xu, and Y\. Wang \(2021\)Transformer in transformer\.Advances in neural information processing systems34,pp\. 15908–15919\.Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p5.5)\. - K\. He, X\. Zhang, S\. Ren, and J\. Sun \(2016\)Deep residual learning for image recognition\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 770–778\.Cited by:[§3\.1](https://arxiv.org/html/2608.28605#S3.SS1.p2.9),[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p5.5),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - G\. Huang, Y\. Sun, Z\. Liu, D\. Sedra, and K\. Q\. Weinberger \(2016\)Deep networks with stochastic depth\.InEuropean conference on computer vision,pp\. 646–661\.Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p5.5)\. - H\. Ismail Fawaz, B\. Lucas, G\. Forestier, C\. Pelletier, D\. F\. Schmidt, J\. Weber, G\. I\. Webb, L\. Idoumghar, P\. Muller, and F\. Petitjean \(2020\)Inceptiontime: finding alexnet for time series classification\.Data Mining and Knowledge Discovery34\(6\),pp\. 1936–1962\.Cited by:[§3\.1](https://arxiv.org/html/2608.28605#S3.SS1.p2.9),[§3](https://arxiv.org/html/2608.28605#S3.p2.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p3.10)\. - M\. Jin, S\. Wang, L\. Ma, Z\. Chu, J\. Y\. Zhang, X\. Shi, P\. Chen, Y\. Liang, Y\. Li, S\. Pan,et al\.\(2023\)Time\-llm: time series forecasting by reprogramming large language models\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p2.1)\. - J\. Klewer, J\. Springer, and J\. Morshedzadeh \(2022\)Premature ventricular contractions \(pvcs\): a narrative review\.The American Journal of Medicine135\(11\),pp\. 1300–1305\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p3.1)\. - X\. Lan, F\. Wu, K\. He, Q\. Zhao, S\. Hong, and M\. Feng \(2025\)Gem: empowering mllm for grounded ecg understanding with time series and images\.arXiv preprint arXiv:2503\.06073\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§1](https://arxiv.org/html/2608.28605#S1.p3.1),[§2](https://arxiv.org/html/2608.28605#S2.p3.1)\. - V\. J\. Lawhern, A\. J\. Solon, N\. R\. Waytowich, S\. M\. Gordon, C\. P\. Hung, and B\. J\. Lance \(2018\)EEGNet: a compact convolutional network for eeg\-based brain\-computer interfaces\.Journal of Neural Engineering,pp\. 056013\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p1.1)\. - J\. Li, C\. Liu, S\. Cheng, R\. Arcucci, and S\. Hong \(2024\)Frozen language model helps ecg zero\-shot learning\.InMedical Imaging with Deep Learning,pp\. 402–415\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1),[§2](https://arxiv.org/html/2608.28605#S2.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p4.1)\. - Z\. Li, S\. Li, and X\. Yan \(2023\)Time series as images: vision transformer for irregularly sampled time series\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 49187–49204\.Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p2.5),[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p3.4),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - C\. Liu, Z\. Wan, C\. Ouyang, A\. Shah, W\. Bai, and R\. Arcucci \(2024a\)Zero\-shot ecg classification with multimodal learning and test\-time clinical knowledge enhancement\.InForty\-first International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p4.1)\. - J\. Liu and S\. Chen \(2024\)Timesurl: self\-supervised contrastive learning for universal time series representation learning\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 13918–13926\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - X\. Liu, J\. Liu, G\. Woo, T\. Aksu, Y\. Liang, R\. Zimmermann, C\. Liu, J\. Li, S\. Savarese, C\. Xiong,et al\.\(2025a\)Moirai\-moe: empowering time series foundation models with sparse mixture of experts\.InForty\-second International Conference on Machine Learning,Cited by:[§3\.4](https://arxiv.org/html/2608.28605#S3.SS4.p1.1)\. - Y\. Liu, C\. Zhang, J\. Song, S\. Chen, S\. Yin, Z\. Wang, L\. Zeng, Y\. Cao, and J\. Jiao \(2025b\)Mofe\-time: mixture of frequency domain experts for time\-series forecasting models\.arXiv preprint arXiv:2507\.06502\.Cited by:[§3\.4](https://arxiv.org/html/2608.28605#S3.SS4.p1.1)\. - Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. Long \(2024b\)ITransformer: inverted transformers are effective for time series forecasting\.InInternational Conference on Learning Representations,Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - Z\. Liu, Y\. Lin, Y\. Cao, H\. Hu, Y\. Wei, Z\. Zhang, S\. Lin, and B\. Guo \(2021\)Swin transformer: hierarchical vision transformer using shifted windows\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 10012–10022\.Cited by:[§3\.2](https://arxiv.org/html/2608.28605#S3.SS2.p5.5)\. - Z\. Liu, H\. Mao, C\. Wu, C\. Feichtenhofer, T\. Darrell, and S\. Xie \(2022\)A convnet for the 2020s\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 11976–11986\.Cited by:[Table 5](https://arxiv.org/html/2608.28605#S4.T5.18.18.18.4)\. - Z\. Liu, A\. Alavi, M\. Li, and X\. Zhang \(2023\)Self\-supervised contrastive learning for medical time series: a systematic review\.Sensors23\(9\),pp\. 4221\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1),[§3\.1](https://arxiv.org/html/2608.28605#S3.SS1.p1.1)\. - D\. Luo, W\. Cheng, Y\. Wang, D\. Xu, J\. Ni, W\. Yu, X\. Zhang, Y\. Liu, Y\. Chen, H\. Chen,et al\.\(2023\)Time series contrastive learning with information\-aware augmentations\.InProceedings of the Thirty\-Seventh AAAI Conference on Artificial Intelligence and Thirty\-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence,pp\. 4534–4542\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - S\. Lyu, S\. Zhong, W\. Ruan, Q\. Liu, Q\. Wen, H\. Xiong, and Y\. Liang \(2025\)OccamVTS: distilling vision models to 1% parameters for time series forecasting\.arXiv preprint arXiv:2508\.01727\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p2.1)\. - L\. v\. d\. Maaten and G\. E\. Hinton \(2008\)Visualizing data using t\-sne\.Journal of Machine Learning Research9,pp\. 2579–2605\.Cited by:[§4\.5](https://arxiv.org/html/2608.28605#S4.SS5.p6.1)\. - J\. McLaren, J\. N\. de Alencar, E\. K\. Aslanger, H\. P\. Meyers, and S\. W\. Smith \(2024\)From st\-segment elevation mi to occlusion mi: the new paradigm shift in acute myocardial infarction\.JACC: Advances3\(11\),pp\. 101314\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p3.1)\. - A\. Miltiadous, K\. D\. Tzimourta, T\. Afrantou, P\. Ioannidis, N\. Grigoriadis, D\. G\. Tsalikakis, P\. Angelidis, M\. G\. Tsipouras, E\. Glavas, N\. Giannakeas,et al\.\(2023\)A dataset of scalp eeg recordings of alzheimer’s disease, frontotemporal dementia and healthy subjects from routine eeg\.Data8\(6\),pp\. 95\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p1.1),[§7\.1\.1](https://arxiv.org/html/2608.28605#S7.SS1.SSS1.p2.1)\. - S\. Mu and S\. Lin \(2025\)A comprehensive survey of mixture\-of\-experts: algorithms, theory, and applications\.arXiv preprint arXiv:2503\.07137\.Cited by:[§3\.4](https://arxiv.org/html/2608.28605#S3.SS4.p1.1),[§3\.4](https://arxiv.org/html/2608.28605#S3.SS4.p2.10)\. - Y\. Nie, N\. H\. Nguyen, P\. Sinthong, and J\. Kalagnanam \(2023\)A time series is worth 64 words: long\-term forecasting with transformers\.InInternational Conference on Learning Representations,Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - P\. PhysioBank \(2000\)Physionet: components of a new research resource for complex physiologic signals\.Circulation101\(23\),pp\. e215–e220\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p1.1),[§7\.1\.1](https://arxiv.org/html/2608.28605#S7.SS1.SSS1.p3.1)\. - S\. Pratiher, A\. Srivastava, Y\. B\. Priyatha, N\. Ghosh, and A\. Patra \(2022\)A dilated residual vision transformer for atrial fibrillation detection from stacked time\-frequency ecg representations\.InICASSP 2022\-2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 1121–1125\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p1.1)\. - A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, I\. Sutskever,et al\.\(2019\)Language models are unsupervised multitask learners\.OpenAI blog1\(8\),pp\. 9\.Cited by:[Table 5](https://arxiv.org/html/2608.28605#S4.T5.27.27.27.4)\. - B\. Rim, N\. Sung, S\. Min, and M\. Hong \(2020\)Deep learning in physiological signal data: a survey\.Sensors20\(4\),pp\. 969\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1)\. - R\. Salloum and C\.\-C\. J\. Kuo \(2017\)ECG\-based biometrics using recurrent neural networks\.InICASSP,pp\. 2062–2066\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p1.1)\. - V\. Shah, E\. Von Weltin, S\. Lopez, J\. R\. McHugh, L\. Veloso, M\. Golmohammadi, I\. Obeid, and J\. Picone \(2018\)The temple university hospital seizure detection corpus\.Frontiers in neuroinformatics12,pp\. 83\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p1.1),[§7\.1\.1](https://arxiv.org/html/2608.28605#S7.SS1.SSS1.p5.1)\. - R\. Sharma and H\. K\. Meena \(2024\)Emerging trends in eeg signal processing: a systematic review\.SN Computer Science5\(4\),pp\. 415\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1)\. - C\. Shen, W\. Yu, Z\. Zhao, D\. Song, W\. Cheng, H\. Chen, and J\. Ni \(2025\)Multi\-modal view enhanced large vision models for long\-term time series forecasting\.arXiv preprint arXiv:2505\.24003\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p2.1)\. - X\. Shi, S\. Wang, Y\. Nie, D\. Li, Z\. Ye, Q\. Wen, and M\. Jin \(2025\)Time\-moe: billion\-scale time series foundation models with mixture of experts\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§3\.4](https://arxiv.org/html/2608.28605#S3.SS4.p1.1)\. - C\. W\. Tan, A\. Dempster, C\. Bergmeir, and G\. I\. Webb \(2022\)MultiRocket: multiple pooling operators and transformations for fast and effective time series classification\.Data Mining and Knowledge Discovery36\(5\),pp\. 1623–1646\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - S\. Tang, J\. A\. Dunnmon, K\. Saab, X\. Zhang, Q\. Huang, F\. Dubost, D\. Rubin, and C\. Lee\-Messer \(2021\)Self\-supervised graph neural networks for improved electroencephalographic seizure analysis\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p1.1)\. - S\. Tonekaboni, D\. Eytan, and A\. Goldenberg \(2021\)Unsupervised representation learning for time series with temporal neighborhood coding\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1)\. - H\. Touvron and e\. Lavril \(2023\)Llama: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[Table 5](https://arxiv.org/html/2608.28605#S4.T5.24.24.24.5)\. - P\. Wagner, N\. Strodthoff, R\. Bousseljot, D\. Kreiseler, F\. I\. Lunze, W\. Samek, and T\. Schaeffter \(2020\)PTB\-xl, a large publicly available electrocardiography dataset\.Scientific data7\(1\),pp\. 1–15\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p1.1),[§7\.1\.1](https://arxiv.org/html/2608.28605#S7.SS1.SSS1.p4.1)\. - G\. Wang, X\. Liu, Z\. Ying, G\. Yang, Z\. Chen, Z\. Liu, M\. Zhang, H\. Yan, Y\. Lu, Y\. Gao,et al\.\(2023\)Optimized glycemic control of type 2 diabetes with reinforcement learning: a proof\-of\-concept trial\.Nature Medicine29\(10\),pp\. 2633–2642\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p3.10)\. - Y\. Wang, N\. Huang, T\. Li, Y\. Yan, and X\. Zhang \(2024\)Medformer: a multi\-granularity patching transformer for medical time\-series classification\.Advances in Neural Information Processing Systems37,pp\. 36314–36341\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1),[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p1.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1),[§7\.1\.2](https://arxiv.org/html/2608.28605#S7.SS1.SSS2.p1.1),[§7\.1\.3](https://arxiv.org/html/2608.28605#S7.SS1.SSS3.p1.1)\. - Z\. Wang, W\. Yan, and T\. Oates \(2017\)Time series classification from scratch with deep neural networks: a strong baseline\.In2017 International joint conference on neural networks \(IJCNN\),pp\. 1578–1585\.Cited by:[§3\.1](https://arxiv.org/html/2608.28605#S3.SS1.p2.9)\. - G\. Woo, C\. Liu, A\. Kumar, C\. Xiong, S\. Savarese, and D\. Sahoo \(2024\)Unified training of universal time series forecasting transformers\.Proceedings of Machine Learning Research235,pp\. 53140–53164\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p1.1)\. - G\. Woo, C\. Liu, D\. Sahoo, A\. Kumar, and S\. Hoi \(2021\)CoST: contrastive learning of disentangled seasonal\-trend representations for time series forecasting\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1)\. - H\. Wu, T\. Hu, Y\. Liu, H\. Zhou, J\. Wang, and M\. Long \(2022\)Timesnet: temporal 2d\-variation modeling for general time series analysis\.InInternational Conference on Learning Representations,Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - W\. Wu, Y\. Huang, and X\. Wu \(2024\)SRT: improved transformer\-based model for classification of 2d heartbeat images\.Biomedical Signal Processing and Control88,pp\. 105017\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p1.1)\. - J\. Ye, Y\. Yu, W\. Zhang, L\. Wang, J\. Li, and F\. Tsung \(2024\)Empowering time series analysis with foundation models: a comprehensive survey\.arXiv preprint arXiv:2405\.02358\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p2.1)\. - J\. Ye, W\. Zhang, Z\. Li, J\. Li, and F\. Tsung \(2026\)MedSpaformer: a transferable transformer with multi\-granularity token sparsification for medical time series classification\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 27791–27799\.Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p1.1)\. - J\. Ye, W\. Zhang, Z\. Li, J\. Li, M\. Zhao, and F\. Tsung \(2025\)MedualTime: A dual\-adapter language model for medical time series\-text multimodal learning\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2025, Montreal, Canada, August 16\-22, 2025,pp\. 7913–7921\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - Z\. Yue, Y\. Wang, J\. Duan, T\. Yang, C\. Huang, Y\. Tong, and B\. Xu \(2022\)Ts2vec: towards universal representation of time series\.InProceedings of the AAAI conference on artificial intelligence,Vol\.36,pp\. 8980–8987\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - J\. Yun, D\. Shin, E\. H\. Lee, J\. P\. Kim, H\. Ham, Y\. Gu, M\. Y\. Chun, S\. H\. Kang, H\. J\. Kim, D\. L\. Na,et al\.\(2025\)Temporal dynamics and biological variability of alzheimer biomarkers\.JAMA neurology82\(4\),pp\. 384–396\.Cited by:[§3\.4](https://arxiv.org/html/2608.28605#S3.SS4.p1.1)\. - A\. Zeng, M\. Chen, L\. Zhang, and Q\. Xu \(2023\)Are transformers effective for time series forecasting?\.InProceedings of the AAAI conference on artificial intelligence,Vol\.37,pp\. 11121–11128\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. - X\. Zhang, Z\. Zhao, T\. Tsiligkaridis, and M\. Zitnik \(2022\)Self\-supervised contrastive pre\-training for time series via time\-frequency consistency\.Advances in neural information processing systems35,pp\. 3988–4003\.Cited by:[§2](https://arxiv.org/html/2608.28605#S2.p4.1)\. - B\. Zhao, H\. Lu, S\. Chen, J\. Liu, and D\. Wu \(2017\)Convolutional neural networks for time series classification\.Journal of systems engineering and electronics28\(1\),pp\. 162–169\.Cited by:[§3\.1](https://arxiv.org/html/2608.28605#S3.SS1.p1.1)\. - S\. Zhong, W\. Ruan, M\. Jin, H\. Li, Q\. Wen, and Y\. Liang \(2025\)Time\-vlm: exploring multimodal vision\-language models for augmented time series forecasting\.InInternational Conference on Machine Learning \(ICML\), Poster,Cited by:[§1](https://arxiv.org/html/2608.28605#S1.p2.1),[§2](https://arxiv.org/html/2608.28605#S2.p3.1),[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1),[§4\.5](https://arxiv.org/html/2608.28605#S4.SS5.p1.1),[§4\.5](https://arxiv.org/html/2608.28605#S4.SS5.p4.1)\. - T\. Zhou, P\. Niu, L\. Sun, R\. Jin,et al\.\(2024\)One fits all: power general time series analysis by pretrained lm\.Advances in neural information processing systems36\.Cited by:[§4\.1](https://arxiv.org/html/2608.28605#S4.SS1.p2.1)\. ## 7\.Appendix Table 7\.Summary of benchmark datasets, including dataset statistics, train\-validation\-test splits, and data urls\.DatasetsDiseasesTotal SamplesClassesChannelsStepsTrainingValidationTestAPAVA \(2\-Classes\)Alzheimer5,9672162563,1231,4131,431ADFTD \(3\-Classes\)Alzheimer69,75231925640,44614,65814,648TUSZ \(2\-Classes\)Epilepsy22,0402196,00013,2244,4084,408TUSZ \(4\-Classes\)Epilepsy2,8914196,0001,734578579PTB \(2\-Classes\)Cardiopathy64,35621530041,99512,9939,368PTB\-XL \(4\-Classes\)Cardiopathy17,1104121,00010,2663,4223,422PTB\-XL \(5\-Classes\)Cardiopathy17,1105121,00010,2663,4223,422 - • ### 7\.1\.Datasets #### 7\.1\.1\.Datasets Details \(1\)APAVA \(2\-Classes\)\(Escuderoet al\.,[2006](https://arxiv.org/html/2608.28605#bib.bib45)\)is a public EEG dataset for Alzheimer’s disease \(AD\) classification\. It contains two classes: ”Healthy Person” and ”Alzheimer’s disease \(AD\)”\. Since the dataset does not provide one\-to\-one text pairs, we adopt its dataset description and task description as associated clinical semantics:Dataset description: The APAVA dataset comprises 16\-channel EEG recordings for distinguishing patients with AD from healthy control subjects\.Task description: Given the EEG signal, predict whether the sample belongs to a subject diagnosed with Alzheimer’s disease \(AD\) or a healthy control \(HC\) subject\. \(2\)ADFTD \(3\-Classes\)\(Miltiadouset al\.,[2023](https://arxiv.org/html/2608.28605#bib.bib46)\)is a public EEG dataset for Alzheimer’s disease classification\. It contains three classes: ”Healthy Person \(HC\)” , ”Frontotemporal Dementia \(FTD\)”, and ”Alzheimer’s disease \(AD\)” \. Since the dataset does not provide one\-to\-one text pairs, we adopt its dataset description and task description as associated clinical semantics:Dataset description: The ADFTD dataset comprises 16\-channel EEG recordings for differentiating patients with AD, FTD, and healthy control subjects\.Task description: Given the EEG signal, predict whether the sample belongs to a subject diagnosed with Alzheimer’s disease \(AD\), Frontotemporal Dementia \(FTD\), or a healthy control \(HC\) subject\. \(3\)PTB \(2\-Classes\)\(PhysioBank,[2000](https://arxiv.org/html/2608.28605#bib.bib47)\)is a public ECG dataset for heart disease classification\. It contains two classes: ”Healthy Person \(HC\)” and ”Myocardial infarction \(MI\)”\. Since the dataset does not provide one\-to\-one text pairs, we adopt its dataset description and task description as associated clinical semantics:Dataset description: The PTB dataset comprises 15\-lead ECG recordings for classifying and diagnosing cardiac conditions, including healthy controls and myocardial infarction patients\.Task description: Given the ECG signal, predict whether the sample belongs to a healthy control \(HC\) or a patient with myocardial infarction \(MI\)\. \(4\)PTB\-XL\(Wagneret al\.,[2020](https://arxiv.org/html/2608.28605#bib.bib48)\)is a large\-scale public 12\-lead ECG dataset for heart disease diagnosis\.PTB\-XL \(4\-Classes\)consists of coarse\-grained labels: ”Abnormal ECG” , ”Borderline ECG”, ”Normal ECG” , and ”Otherwise normal ECG” \.PTB\-XL \(5\-Classes\)provides fine\-grained labels: ”Conduction Disturbance” , ”Hypertrophy” , ”Myocardial Infarction” , ”ST\-T Changes”, and ”Normal ECG” \. PTB\-XL dataset provides clinical 12\-lead ECGs and their corresponding reports\. The clinical reports are automatically generated by the machine and have no diagnosis revealed\. Since the dataset provides one\-to\-one clinical report, we adopt these clinical records as associated textual semantics\. \(5\)TUSZ\(Shahet al\.,[2018](https://arxiv.org/html/2608.28605#bib.bib49)\)is a large\-scale EEG dataset capturing brain electrical activity across 19 channels for epilepsy diagnosis\.TUSZ \(2\-Classes\)provides coarse\-grained labels: ”Normal EEG” and ”Abnormal EEG” \.TUSZ \(4\-Classes\)provides fine\-grained labels, further categorizing abnormal EEG into four seizure types: ”Combined focal \(CF\) seizures” , ”Generalized non\-specific \(GN\) seizures” , ”Absence \(AB\) seizures” , and ”Combined tonic \(CT\) seizures”\. TUSZ contains 19\-channel EEG recordings along with patient\-level clinical records for each diagnostic session, including clinical history, medications, and other relevant notes\. Therefore, we adopt these records as the associated textual information for each EEG sample\. #### 7\.1\.2\.Datasets Split Table[7](https://arxiv.org/html/2608.28605#S7.T7)provides the train\-validation\-test splits for all the datasets\. The splits for the APAVA, ADFTD, and PTB datasets follow the previous work\(Wanget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib67)\), while the PTB\-XL and TUSZ datasets employ a 60%\-20%\-20% splitting strategy\. #### 7\.1\.3\.Data Pre\-processing For the data preprocessing of the APAVA, ADFTD, PTB, PTB\-XL dataset and TUSZ datasets, we follow previous work\(Wanget al\.,[2024](https://arxiv.org/html/2608.28605#bib.bib67)\)\.
Similar Articles
Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift
Introduces CALCoDe, a post-hoc reliability layer for frozen medical vision-language models that mitigates class-tail undercoverage under clinical shift, achieving strong worst-class accepted coverage across multiple dermatology shifts and VLM backbones.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
The paper proposes converting multimodal patient data (text, labs, vitals) into a single natural language sequence and fine-tuning LLMs for clinical prediction, achieving comparable or better performance than specialized fusion architectures across three tasks.
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Hulu-Med is a transparent medical vision-language model that unifies understanding across text, 2D/3D images, and video, achieving state-of-the-art performance on 30 benchmarks while being fully open-source.
LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
LLM4EHR proposes a clinical foundation model that temporally aligns Electronic Health Record time series with medical event sequences using a domain-adapted large language model and a regularized contrastive objective, improving downstream prediction tasks.