parameter-free

Tag

Cards List
#parameter-free

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization

arXiv cs.LG ↗ · 2026-08-06 Cached

This paper introduces SpecDrop, a parameter-free category-conditioned routing scheme for modular networks, showing that on vision tasks it achieves competitive accuracy while on fuzzy language partitions it reduces to no-routing baselines, suggesting granularity alignment matters more than router design.

0 favorites 0 likes
#parameter-free

Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection

arXiv cs.LG ↗ · 2026-07-27 Cached

This paper proposes a parameter-free adaptive sparse attention method that uses gzip compression ratios to dynamically select non-redundant blocks for long-range attention, achieving significant perplexity improvements over fixed and learned sparse attention baselines on PG-19 language modeling.

0 favorites 0 likes
#parameter-free

Zero-order Parameter-free Optimization for LMO-based Methods: Novel Approach for Efficient Fine-tuning

arXiv cs.LG ↗ · 2026-06-16 Cached

This paper introduces AdaNAGED, a method that combines zero-order optimization, parameter-free adaptation, and non-Euclidean update geometry for memory-efficient fine-tuning of large language models, with theoretical convergence guarantees and validation on the OPT-1.3B model.

0 favorites 0 likes
#parameter-free

Simply Stabilizing the Loop via Fully Looped Transformer

arXiv cs.LG ↗ · 2026-05-20 Cached

This paper identifies gradient oscillation and residual explosion as causes of training instability in Looped Transformers, and proposes Fully Looped Transformer with two parameter-free modifications (Fully Looped Architecture and Attention Injection) to stabilize training up to 12 loop iterations, achieving up to 13.2% improvement in downstream performance.

0 favorites 0 likes
← Back to home

Submit Feedback