efficiency

Tag

Cards List
#efficiency

Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows

arXiv cs.AI ↗ · 6h ago Cached

The paper introduces Learning What to Skip (LW2S), a method that uses counterfactual credit assignment to optimize multi-agent LLM workflows by selectively skipping components, reducing token cost while maintaining or improving accuracy.

0 favorites 0 likes
#efficiency

As A.I. makes law firms more efficient, clients ask: 'Where's my discount?'

Hacker News Top ↗ · 8h ago

Artificial intelligence is making law firms more efficient, leading clients to request discounts. The article discusses the impact of AI on legal services and pricing expectations.

0 favorites 0 likes
#efficiency

Over-reasoning in LLMs.

Reddit r/ArtificialInteligence ↗ · 13h ago Cached

The article critiques the trend of over-reasoning in large language models, arguing that excessive computation and longer answers often reduce efficiency and user experience without proportional benefits.

0 favorites 0 likes
#efficiency

What are chinese labs doing differently?

Reddit r/artificial ↗ · 19h ago

The post questions why Chinese AI labs are achieving better results with less investment compared to American labs, considering factors like efficiency, post-training techniques, and data acquisition.

0 favorites 0 likes
#efficiency

@AnatoliKopadze: Jev + Opus 5.5 is the most underrated setup in AI right now, and i don't get why nobody's talking about it. I stopped l…

X AI KOLs Timeline ↗ · 21h ago Cached

A user shares a cost-saving method for AI agents by using Jev for routine decisions and Opus 5.5 only for hard reasoning, significantly reducing costs while maintaining results.

0 favorites 0 likes
#efficiency

@flitternie: Jev has attracted many attention by treating decision-making as a first-class model problem. I wrote about where it fai…

X AI KOLs Timeline ↗ · 2d ago Cached

Jev redefines AI decision-making by using predefined output spaces instead of generation, offering efficiency gains for enterprises and long-horizon tasks, as analyzed in a detailed discussion.

0 favorites 0 likes
#efficiency

Writing Efficient C++ Code

Hacker News Top ↗ · 2d ago Cached

This article discusses techniques for writing efficient C++ code, emphasizing data-oriented design and performance optimization for applications like games and real-time processing.

0 favorites 0 likes
#efficiency

Swift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash!

Reddit r/LocalLLaMA ↗ · 2d ago

The article compares Swift1.5-Qwen3.8-Flash-Next with the base Qwen3.8-Flash-Next, showing that the Swift model maintains similar quality while drastically reducing token usage and time for coding tasks.

0 favorites 0 likes
#efficiency

@BottleCapAI: Same model size. Same GPU. ~4.7x more work done! ThinkingCap: Qwen 3.8 reaches answers with about half the reasoning. O…

X AI KOLs Timeline ↗ · 2d ago Cached

ThinkingCap is a finetuned model based on Qwen 3.6 27B that reduces reasoning tokens by about 50% while maintaining performance, leading to significant efficiency gains in inference.

0 favorites 0 likes
#efficiency

An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model

arXiv cs.CL ↗ · 3d ago Cached

This paper presents an exploratory ablation study of TALH, a hybrid language model combining MLA and SSM, showing that SSM integration is more critical for validation performance than MLA in the tested setup, with insights on memory usage and timing on consumer hardware.

0 favorites 0 likes
#efficiency

No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow

arXiv cs.CL ↗ · 3d ago Cached

This paper introduces Corpus Task Complexity (CTC) to characterize how task difficulty scales with corpus size, presents high-CTC tasks, and releases CTC-Bench, showing that high-CTC tasks are more challenging for long-context language models.

0 favorites 0 likes
#efficiency

Opus 5.5 summary: 66.4% Terminal-Bench, ~40% cheaper to run than Opus 5, cache reads down 60%

Reddit r/artificial ↗ · 3d ago

Opus 5.5, launched by Anthropic, achieves 66.4% on Terminal-Bench, reduces costs by ~40% compared to Opus 5, and features a 1M context window with cache reads down 60%.

0 favorites 0 likes
#efficiency

Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

arXiv cs.LG ↗ · 4d ago Cached

This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.

0 favorites 0 likes
#efficiency

StateComp: Learning When to Compress History in Long Horizon Agents

arXiv cs.AI ↗ · 4d ago Cached

StateComp is a framework that compresses historical interactions in long-horizon agents based on the current state, reducing token usage by 52.27% and achieving a 12.67× speedup in representation extraction while maintaining task performance.

0 favorites 0 likes
#efficiency

Realize What Matters: Principled Context Representation for Large-Scale Reasoning

arXiv cs.CL ↗ · 4d ago Cached

The paper proposes principled methods for context representation in large-scale AI reasoning, introducing R3Con which outperforms baselines and enables smaller models to achieve performance comparable to larger ones at lower cost.

0 favorites 0 likes
#efficiency

@omarsar0: Agents are coming to healthcare. Doctors spend much of their day on admin and research. @almanac_health lets them deleg…

X AI KOLs Following ↗ · 4d ago Cached

Almanac Health is launching an AI agent platform that helps doctors delegate administrative and research tasks, thereby increasing time for patient care.

0 favorites 0 likes
#efficiency

Flux 3 Action: open weights 7B World Action Model

Reddit r/singularity ↗ · 4d ago

FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.

0 favorites 0 likes
#efficiency

@Lonely__MH: All rise! The GGUF quantized version of Qwen-Image-2.1 is here! Thanks to the @UnslothAI team for stepping in! Based on…

X AI KOLs Timeline ↗ · 5d ago Cached

The GGUF quantized version of the Qwen-Image-2.1 AI model is released, featuring Dynamic 2.0 technology for efficient 4-bit quantization, supporting text-to-image and transparent image generation in a 4.2GB size suitable for Mac users.

0 favorites 0 likes
#efficiency

Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices

arXiv cs.LG ↗ · 5d ago Cached

This paper proposes GittinsEval, a cost-aware Bayesian bandit framework for efficient LLM evaluation that significantly reduces costs while maintaining high performance by adaptively selecting configurations.

0 favorites 0 likes
#efficiency

When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used

arXiv cs.AI ↗ · 5d ago Cached

The paper introduces CounterCredit, a training method for vision-language agents that ensures visual calls are both needed and used, leading to higher performance and fewer spurious calls on benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback