reasoning-efficiency

Tag

Cards List
#reasoning-efficiency

ukisai/Swift-Qwen3.8-27b

Hugging Face Models Trending · 2026-09-08 Cached

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, reducing thinking tokens by 58.3% while maintaining near-identical performance.

0 favorites 0 likes
#reasoning-efficiency

Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

Reddit r/LocalLLaMA · 2026-09-03

A paper introduces an inference-time optimization for sparse MoE models by adjusting expert selection in late transformer layers, reducing reasoning tokens by 8.5% and latency by 10.9% without retraining, while maintaining accuracy.

0 favorites 0 likes
#reasoning-efficiency

Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)

Reddit r/LocalLLaMA · 2026-07-21

Gemma-4-26B-a4B, with an updated chat template, outperforms Qwen3.6-MoE and Qwen3.5-MoE in fine-tuned instruct mode and reasoning efficiency.

0 favorites 0 likes
#reasoning-efficiency

@kaiostephens: anoooother one! Introducing Grug-30b-a3b! Grug is an open-source experimental model build on top of Qwen-3.6-30b-a3b. D…

X AI KOLs Timeline · 2026-07-03 Cached

Grug-35B-A3B is an open-source experimental model based on Qwen-3.6-35B-A3B that reduces reasoning token usage by approximately 69.8% while maintaining answer quality within 2% of the base model, resulting in faster generation and reduced context usage.

0 favorites 0 likes
#reasoning-efficiency

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling

arXiv cs.AI · 2026-06-08 Cached

This paper introduces DyCon, a training-free framework that uses step-level embeddings to model evolving task difficulty and dynamically control reasoning depth in Large Reasoning Models, effectively reducing overthinking and improving efficiency without sacrificing accuracy.

0 favorites 0 likes
#reasoning-efficiency

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

arXiv cs.AI · 2026-06-03 Cached

ThoughtFold proposes a framework using introspective preference learning to reduce redundant explorations in Chain-of-Thought reasoning for Large Reasoning Models, achieving ~56% token reduction on DeepSeek-R1-Distill-Qwen-7B without accuracy loss.

0 favorites 0 likes
#reasoning-efficiency

Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness

arXiv cs.LG · 2026-05-13 Cached

This paper introduces the VPG-EA framework, which uses variational inference and posterior guidance to improve the reasoning efficiency of large language models by addressing the 'overthinking' phenomenon in chain-of-thought generation.

0 favorites 0 likes
#reasoning-efficiency

Hint Tuning: Less Data Makes Better Reasoners

arXiv cs.CL · 2026-05-12 Cached

This paper introduces 'Hint Tuning,' a data-efficient method that reduces token usage in reasoning models by calibrating reasoning depth based on problem difficulty. It achieves significant token reduction (24–66%) on models like Qwen3-Thinking and DeepSeek-R1-Distill using only 1K self-annotated samples.

0 favorites 0 likes
← Back to home

Submit Feedback