Tag
SMILESGNN introduces a multimodal architecture combining SMILES transformers and graph neural networks with cross-attention for interpretable drug toxicity prediction, achieving competitive performance on benchmarks like ClinTox and Tox21 with minimal parameters.
ROOSTER is a shared module that learns alignment between condition and target sequences for time-series forecasting and PPG-to-vital-sign reconstruction, achieving superior performance across multiple benchmarks.
The paper introduces RecCAR, a regularization method to address the reciprocal correspondence gap in joint multimodal diffusion transformers, improving performance in video generation tasks.
EnSol is an environment-aware graph neural network that predicts molecular solubility by representing solutes and solvents as graphs and using cross-attention to model interactions, with probabilistic outputs to capture temperature effects and experimental uncertainty. It achieves state-of-the-art performance on benchmark datasets, validated experimentally.
An interactive diagram comparing self-attention and cross-attention mechanisms in AI models, published as part of an educational library on attention.
QueryFormer is a unified transformer architecture that won the KDD Cup 2026 Tencent UniRec Challenge for post-click conversion rate prediction, focusing on query generation and efficient scaling.
This paper proposes Time–Frequency Geometric Cross-Attention (TFGCA), a drop-in module for chunked vision-language-action models that improves action trajectory prediction by decomposing chunks into time-frequency representations and capturing geometric relationships, resulting in significant performance gains on benchmarks and real-robot tasks.
RelightFormer introduces a feed-forward generative Transformer for direct single- and multi-view image relighting, using cross-attention for illumination injection and permutation-invariant encodings for unordered views, trained on a massive synthetic dataset to achieve state-of-the-art visual quality.
This paper investigates semantic leakage in audio-video diffusion models through the 'attention triangle' of cross-attention mechanisms, and presents methods to enhance semantic grounding during generation.
This paper introduces addressable and cardinality-preserving global memory for message-passing neural networks via cross-attention slots, addressing the finite-capacity bottleneck of virtual nodes and improving performance on multiplicity-aware tasks.
This paper proposes a dual-modal Vision Cross-Attention architecture for reaction yield prediction, fusing tabular physical-organic data with 2D molecular topologies, and demonstrates that a generic computer vision backbone can outperform purely quantum-based baselines.
Introduces SyRuP, a decoding-time framework that trains a cross-attention reward head to produce token-level adherence scores for system prompts, improving LLM following of complex prompts without model tuning.
TokenMem injects knowledge into frozen LLMs via a dedicated cross-attention channel, training a thin gating adapter through two-phase curriculum to improve knowledge compliance under counterfactual knowledge, achieving 69-70% KC compared to 20-52% for vanilla RAG.
MOSS-VL-Realtime is a realtime streaming vision-language model that processes continuous video frames, supports interruptible interaction, proactive silence, and dynamic correction, with timestamp-aware encoding and a 256K context window.
RaysUp is an ultra-lightweight, task-agnostic feature upsampling framework that uses geometry-aware ray domain techniques to reconstruct high-resolution features from low-resolution VFM outputs, achieving state-of-the-art performance with 84% fewer parameters than prior work and 7x faster inference.
KaLM-Reranker-V1 is a fast reranker that decouples query and passage computation using an encoder-decoder architecture with Matryoshka embedding pooling and cross-attention, achieving state-of-the-art reranking performance on BEIR and competitive results on multilingual benchmarks.
This paper proposes ST-Merge, a steerable model merging framework that uses a gated cross-attention mechanism to adaptively modulate contributions of a multilingual model and a reasoning model, outperforming fixed merging approaches on multilingual reasoning benchmarks across 21 languages.
This paper introduces Multi-Adapter PPO, a reinforcement learning framework with cross-attention for wavelength selection in LIBS quantitative analysis, achieving 28.4% better composite scores and 45.2% improvement in prediction accuracy over traditional methods on steel and coal datasets.
This paper extends optimal transport-based hallucination detection to all decoder layers in NMT and abstractive summarization, finding that detection is concentrated in early layers and that the geometric signal transfers poorly to summarization due to faithfulness failures not detectable via attention concentration.
This paper proposes SCALE, a deep reinforcement learning scheduler for agentic LLM workflow DAGs that generalizes to unseen cluster sizes using cross-attention and structured representation regularization, reducing response time without retraining.