Tag
ReVA introduces a region-aware visual assistant that enhances visually grounded question answering by integrating whole-image and region-level representations, reducing hallucinations in multimodal large language models.
TRACE-BN introduces a curriculum-guided dataset for structured Bangla-English tutoring and demonstrates transferring this behavior to a sub-1B language model using LoRA, achieving significant improvements in tutoring quality for resource-constrained offline environments.
This paper presents a system for CLEF 2026 CheckThat! Task 2 that uses LLM-based trace ranking and grouped reward modeling for verifying numerical claims in English and Arabic, comparing fine-tuned verifiers with lightweight reward models.
This paper proposes that AI memory consolidation should recombine knowledge across domains (like dreaming) rather than merely replaying experiences, and demonstrates that cross-domain consolidation improves performance in both neural (LoRA fine-tuning) and symbolic systems.
This paper presents a strategy to adapt an open-source spoken language model to the Singaporean Home Team context using LoRA fine-tuning, a surrogate text-QA dataset, and a multi-task objective, achieving competitive performance across five speech tasks in Singapore's four official languages.
This paper systematically evaluates time series foundation models (TSFMs) for forecasting extreme PM2.5 concentrations from wildfire smoke using a 12-year dataset from California. Results show that fully-trained recurrent baselines like BiLSTM outperform TSFMs, challenging the assumption that larger pretrained models dominate in environmental forecasting.
LayerRoute is a lightweight adapter that selectively skips transformer blocks during inference based on input type, achieving compute savings while maintaining or improving model quality through gated routing and LoRA adaptation. It achieves a 12.91% skip differential on agentic language models.
Warp-as-History proposes a novel interface that transforms camera-induced warps into pseudo-history representations, enabling a frozen video generation model to follow camera trajectories without training or test-time optimization. A lightweight LoRA fine-tuning on a single video further improves camera adherence and generalizes to unseen videos.
TrackCraft3R repurposes video diffusion transformers for dense 3D tracking from monocular video, using dual-latent representation and temporal RoPE alignment to achieve state-of-the-art performance with 1.3x faster speed and 4.6x less peak memory than prior methods.
This paper proposes CAP-TTA, a test-time adaptation framework that uses preconditioned LoRA updates triggered by bias-risk scores to mitigate toxicity and bias in large language models during narrative generation, achieving faster optimization and better fluency than standard baselines.
LiconStudio releases a LoRA adapter for LTX-2.3 fine-tuned on the VBVR dataset to enhance video generation with improved prompt understanding, motion dynamics, and temporal consistency for complex video reasoning tasks.