Tag
The PyTorch Foundation is hosting an Introduction Track and PyTorch Associate Training at PyTorch Conference North America 2026 to help developers, researchers, and engineers enhance their technical skills in deep learning and AI.
ROOSTER is a shared module that learns alignment between condition and target sequences for time-series forecasting and PPG-to-vital-sign reconstruction, achieving superior performance across multiple benchmarks.
The paper presents Neurogenesis Network (NGN), a differentiable parameterization for learning the optimal size of neural networks during training, applicable to various architectures like MLPs, CNNs, and Transformers.
This paper conducts a scaling study for fMRI foundation models, revealing that performance depends on the combination of pretraining data size, model size, and training duration, not just compute.
This paper introduces a dual-model masking metric to benchmark ten explainable methods for temporal attribution in sequential recommendation systems, finding that gradient-based methods like GradientSHAP and Integrated Gradients yield the most faithful and robust attributions.
The paper proposes DissipNet, a deep discrete-time dissipative recurrent neural network that explicitly enforces dissipativity through structural constraints to ensure stable modeling of dissipative systems, outperforming traditional RNNs and Physics-Informed Neural Networks.
CoRe-Stack+ is a meta-learning pipeline that improves deep stacking generalization by filtering redundancies and enhancing calibration, achieving better accuracy and efficiency on vision benchmarks.
Quartet introduces a graph transformer architecture with quad-branch cross-attention and a causal random walk sampler to enhance performance on relational graph tasks, outperforming state-of-the-art baselines like RelGT and HGT.
The paper introduces HARN, a hierarchical associative resonance network for event-driven multi-timeframe forecasting in financial time series, showing competitive results against baselines through evaluations on multiple assets.
The paper proposes ADNet, an adaptive decomposition network for multi-step traffic forecasting that learns to disentangle heterogeneous traffic dynamics into dominant and residual components via spectral decomposition, achieving superior performance on the TraffiDent dataset.
MT-ProtBERT is a multi-task learning model for classifying intrinsically disordered proteins under data scarcity, integrating self-supervised and biochemistry-informed tasks to outperform existing methods like PARROT.
This paper identifies blind spots in evaluating deep imbalanced regression, proposing balanced metrics and showing high tail-region instability across random seeds.
This paper introduces R-GEAN, an asymmetric candidate-scoring network for predicting medication changes in hospital admissions, and establishes a leakage-controlled benchmark to accurately evaluate edit-level performance.
This paper designs a four-tier experimental teaching system for multimodal medical image intelligent diagnosis, translating research into undergraduate labs to address gaps in education for clinical AI.
Hapi is a U-Net Swin Transformer that provides medium-range hydrological forecasts at continental scale, outperforming physics-based and AI models in flood detection across the contiguous United States.
KVMEM enhances AI agent memory by preserving old KV cache states, improving task performance and efficiency over compaction methods in long-running agents.
This paper proposes a method using Gaussian Process decorrelation to remove spatial correlations before training an LSTM for predicting snow water equivalent, improving predictive accuracy and incorporating conformal prediction for uncertainty quantification.
Contrastive World Models propose a new approach for learning latent dynamics without pixel reconstruction, using a contrastive objective to improve robustness and efficiency in visually complex environments for model-based reinforcement learning.
This paper presents a predictive maintenance framework that uses deep learning and uncertainty estimation to improve remaining useful life estimation for semiconductor manufacturing, reducing maintenance costs.
StationPDE is a station-oriented surface PDE learning model for multi-station multivariate weather forecasting that constructs terrain-aware continuous fields and models physical dynamics to outperform state-of-the-art baselines with a 9.6% MSE reduction.