parameter-efficiency

Tag

Cards List
#parameter-efficiency

@TechByMarkandey: What if progress is no longer about adding parameters?

X AI KOLs Timeline ↗ · 2026-09-07 Cached

MiniCPM5-2B is a 2B-parameter open-source language model that achieves high intelligence density for edge deployment and ranks highly in several AI performance benchmarks.

0 favorites 0 likes
#parameter-efficiency

NanoSleep: A Parameter-Efficient Hybrid Temporal Convolutional Network for Single-Channel Sleep Stage Classification

arXiv cs.LG ↗ · 2026-08-20 Cached

The paper presents NanoSleep, a parameter-efficient hybrid temporal convolutional network for automatic sleep stage classification using single-channel EEG, designed for wearable and home-based monitoring on resource-constrained devices.

0 favorites 0 likes
#parameter-efficiency

Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper presents Wiola, a 13M parameter decoder-only language model with novel components like Spiral Rotary Positional Encoding and Gated Spiral Attention to enhance parameter efficiency for small-scale, on-device language models.

0 favorites 0 likes
#parameter-efficiency

Recursive transformers for semiconductor thermo-mechanical reliability

arXiv cs.LG ↗ · 2026-07-31 Cached

This paper evaluates recursive weight-sharing transformer architectures as parameter-efficient surrogate models for semiconductor thermo-mechanical reliability prediction, comparing performance, parameter count, and computational cost on small engineering datasets.

0 favorites 0 likes
#parameter-efficiency

Between Gradient and Natural Gradient: A Continuum of LoRA Initializations

arXiv cs.LG ↗ · 2026-07-30 Cached

This paper proposes Unified LoRA (ULoRA), a two-parameter family of preconditioned gradient initializations for low-rank adaptation, showing that existing LoRA initialization methods are points on a continuum. The authors demonstrate that a tuned ULoRA matches or exceeds full fine-tuning on GLUE tasks with RoBERTa-base and is competitive on GSM8K with LLaMA 2-7B, and introduce ULoRA-Auto for zero-search deployment.

0 favorites 0 likes
#parameter-efficiency

Scaling Closed-Loop Feature Channel Configuration with LLMs

arXiv cs.LG ↗ · 2026-07-24 Cached

This paper scales a closed-loop LLM-based channel configuration search to 250 candidates per cycle, showing positive accuracy trends and improved parameter efficiency on CIFAR-100, and revealing architectural regularities in LLM-generated channel priors.

0 favorites 0 likes
#parameter-efficiency

SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions

arXiv cs.LG ↗ · 2026-07-22 Cached

SechKAN is a novel Kolmogorov-Arnold Network architecture that uses hyperbolic secant functions as basis functions, achieving competitive performance in function fitting, PDE problems, and image classification tasks while maintaining parameter efficiency comparable to MLPs.

0 favorites 0 likes
#parameter-efficiency

AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling

Hugging Face Daily Papers ↗ · 2026-07-13 Cached

AsySplat proposes an asymmetric architecture that decouples geometry and appearance modeling in 3D Gaussian Splatting, achieving high efficiency for long-sequence scene modeling with nearly 800x speedup over optimization-based methods.

0 favorites 0 likes
#parameter-efficiency

MultiHashFormer: Hash-based Generative Language Models

arXiv cs.CL ↗ · 2026-06-29 Cached

MultiHashFormer is a hash-based generative language model that represents each token as a unique hash signature, enabling parameter-efficient autoregression. It outperforms standard Transformer LMs at 100M, 1B, and 3B scales and supports multilingual vocabulary expansion without increasing parameters.

0 favorites 0 likes
#parameter-efficiency

@Fenng: Saw this piece written by a self-media account — 'The latest fourth-generation WeLM-80B now has only 80 billion total parameters, with 3 billion activated, an activation rate of just 3.75%. For comparison — DeepSeek-V4-Flash, the domestic representative of extreme cost-performance, has 284 billion total parameters, 13 billion activated, activation rate of 4.6%...'

X AI KOLs Timeline ↗ · 2026-06-23 Cached

Fenng shares a self-media comparison between the fourth-generation WeLM-80B (80B total params, 3B activated, 3.75% activation rate) and DeepSeek-V4-Flash (284B total, 13B activated, 4.6% activation rate), with a humorous comment.

0 favorites 0 likes
#parameter-efficiency

Operator Boosting Produces Pareto-Efficient PDE Surrogates

arXiv cs.LG ↗ · 2026-06-17 Cached

Operator Boosting is a stagewise residual-learning framework that constructs compact neural operator surrogates for PDEs by training tiny models on residual fields. It achieves accuracy comparable to or better than full-size models while reducing parameters by up to 95%, demonstrating Pareto improvements on several benchmarks.

0 favorites 0 likes
#parameter-efficiency

Looped World Models

Hugging Face Daily Papers ↗ · 2026-06-16 Cached

Looped World Models introduce iterative latent state refinement through shared transformer blocks, achieving 100x parameter efficiency while adapting computational depth to prediction complexity.

0 favorites 0 likes
#parameter-efficiency

Tiny Scale Is All I Can Spare To Play With Transformer

Reddit r/LocalLLaMA ↗ · 2026-06-11

A student introduces Silia, a novel transformer architecture that combines attention and FFN into a unified operation to save parameters at scales ≤10M, achieving comparable performance to GPT-2 with fewer parameters despite limited compute resources.

0 favorites 0 likes
#parameter-efficiency

Communication Dynamics Neural Networks: FFT-Diagonalized Layers for Improved Hessian Conditioning at Reduced Parameter Count

arXiv cs.LG ↗ · 2026-05-12 Cached

This paper introduces CDLinear, a block-circulant neural network layer that reduces parameter count and improves Hessian conditioning via FFT diagonalization, validated on MNIST with theoretical proofs.

0 favorites 0 likes
#parameter-efficiency

Forgive my ignorance but how is a 27B model better than 397B?

Reddit r/LocalLLaMA ↗ · 2026-04-22

User questions how Qwen's 27B dense model can outperform its 397B MoE variant, sparking discussion on MoE efficiency versus dense model quality.

0 favorites 0 likes
#parameter-efficiency

@aakashgupta: Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1…

X AI KOLs Timeline ↗ · 2026-04-22 Cached

Andrej Karpathy claimed to Dwarkesh Patel that a 1B-parameter model trained on ultra-clean data could match today's 1.8T-parameter frontier models, implying 1,800× effective compression.

0 favorites 0 likes
#parameter-efficiency

ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning

arXiv cs.CL ↗ · 2026-04-22 Cached

ShadowPEFT introduces a centralized parameter-efficient fine-tuning method that uses a depth-shared shadow module to refine transformer layer representations, matching or outperforming LoRA/DoRA with comparable trainable parameters.

0 favorites 0 likes
← Back to home

Submit Feedback