Tag
This paper proposes Time–Frequency Geometric Cross-Attention (TFGCA), a drop-in module for chunked vision-language-action models that improves action trajectory prediction by decomposing chunks into time-frequency representations and capturing geometric relationships, resulting in significant performance gains on benchmarks and real-robot tasks.
RelightFormer introduces a feed-forward generative Transformer for direct single- and multi-view image relighting, using cross-attention for illumination injection and permutation-invariant encodings for unordered views, trained on a massive synthetic dataset to achieve state-of-the-art visual quality.
This paper investigates semantic leakage in audio-video diffusion models through the 'attention triangle' of cross-attention mechanisms, and presents methods to enhance semantic grounding during generation.
This paper introduces addressable and cardinality-preserving global memory for message-passing neural networks via cross-attention slots, addressing the finite-capacity bottleneck of virtual nodes and improving performance on multiplicity-aware tasks.
This paper proposes a dual-modal Vision Cross-Attention architecture for reaction yield prediction, fusing tabular physical-organic data with 2D molecular topologies, and demonstrates that a generic computer vision backbone can outperform purely quantum-based baselines.
Introduces SyRuP, a decoding-time framework that trains a cross-attention reward head to produce token-level adherence scores for system prompts, improving LLM following of complex prompts without model tuning.
TokenMem injects knowledge into frozen LLMs via a dedicated cross-attention channel, training a thin gating adapter through two-phase curriculum to improve knowledge compliance under counterfactual knowledge, achieving 69-70% KC compared to 20-52% for vanilla RAG.
MOSS-VL-Realtime is a realtime streaming vision-language model that processes continuous video frames, supports interruptible interaction, proactive silence, and dynamic correction, with timestamp-aware encoding and a 256K context window.
RaysUp is an ultra-lightweight, task-agnostic feature upsampling framework that uses geometry-aware ray domain techniques to reconstruct high-resolution features from low-resolution VFM outputs, achieving state-of-the-art performance with 84% fewer parameters than prior work and 7x faster inference.
KaLM-Reranker-V1 is a fast reranker that decouples query and passage computation using an encoder-decoder architecture with Matryoshka embedding pooling and cross-attention, achieving state-of-the-art reranking performance on BEIR and competitive results on multilingual benchmarks.
This paper proposes ST-Merge, a steerable model merging framework that uses a gated cross-attention mechanism to adaptively modulate contributions of a multilingual model and a reasoning model, outperforming fixed merging approaches on multilingual reasoning benchmarks across 21 languages.
This paper introduces Multi-Adapter PPO, a reinforcement learning framework with cross-attention for wavelength selection in LIBS quantitative analysis, achieving 28.4% better composite scores and 45.2% improvement in prediction accuracy over traditional methods on steel and coal datasets.
This paper extends optimal transport-based hallucination detection to all decoder layers in NMT and abstractive summarization, finding that detection is concentrated in early layers and that the geometric signal transfers poorly to summarization due to faithfulness failures not detectable via attention concentration.
This paper proposes SCALE, a deep reinforcement learning scheduler for agentic LLM workflow DAGs that generalizes to unseen cluster sizes using cross-attention and structured representation regularization, reducing response time without retraining.
This paper proposes a query-based cross-modal projector that compresses visual tokens via cross-attention to improve Mamba-based multimodal LLMs, boosting both performance and throughput on vision-language benchmarks while eliminating the need for manual 2D scan order design.
Introduces ERP-XTTN, a cross-attention architecture for interpretable ERP classification across subjects without calibration. Evaluated on multiple datasets, it achieves competitive performance with black-box models while providing transparent routing insights.
ReactiveGWM is a reactive game world model that enables dynamic player-NPC interactions by decoupling player controls from NPC behaviors using diffusion models and cross-attention modules, achieving zero-shot strategy transfer across different games.
Open_MOSS released MOSS-VL, an 11B Apache 2.0 vision-language model using cross-attention and XRoPE that outperforms Qwen3-VL-8B by 8.3 points on VSI-bench.
Motif-Video 2B is a 2B parameter text-to-video generation model that achieves 83.76% on VBench, surpassing Wan2.1 14B while using 7x fewer parameters and trained on fewer than 10M clips with less than 100,000 H200 GPU hours. The model uses a specialized architecture with shared cross-attention and a three-part backbone to separate prompt alignment, temporal consistency, and detail refinement.