Tag
The paper introduces Learning What to Skip (LW2S), a method that uses counterfactual credit assignment to optimize multi-agent LLM workflows by selectively skipping components, reducing token cost while maintaining or improving accuracy.
Artificial intelligence is making law firms more efficient, leading clients to request discounts. The article discusses the impact of AI on legal services and pricing expectations.
The article critiques the trend of over-reasoning in large language models, arguing that excessive computation and longer answers often reduce efficiency and user experience without proportional benefits.
The post questions why Chinese AI labs are achieving better results with less investment compared to American labs, considering factors like efficiency, post-training techniques, and data acquisition.
A user shares a cost-saving method for AI agents by using Jev for routine decisions and Opus 5.5 only for hard reasoning, significantly reducing costs while maintaining results.
Jev redefines AI decision-making by using predefined output spaces instead of generation, offering efficiency gains for enterprises and long-horizon tasks, as analyzed in a detailed discussion.
This article discusses techniques for writing efficient C++ code, emphasizing data-oriented design and performance optimization for applications like games and real-time processing.
The article compares Swift1.5-Qwen3.8-Flash-Next with the base Qwen3.8-Flash-Next, showing that the Swift model maintains similar quality while drastically reducing token usage and time for coding tasks.
ThinkingCap is a finetuned model based on Qwen 3.6 27B that reduces reasoning tokens by about 50% while maintaining performance, leading to significant efficiency gains in inference.
This paper presents an exploratory ablation study of TALH, a hybrid language model combining MLA and SSM, showing that SSM integration is more critical for validation performance than MLA in the tested setup, with insights on memory usage and timing on consumer hardware.
This paper introduces Corpus Task Complexity (CTC) to characterize how task difficulty scales with corpus size, presents high-CTC tasks, and releases CTC-Bench, showing that high-CTC tasks are more challenging for long-context language models.
Opus 5.5, launched by Anthropic, achieves 66.4% on Terminal-Bench, reduces costs by ~40% compared to Opus 5, and features a 1M context window with cache reads down 60%.
This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.
StateComp is a framework that compresses historical interactions in long-horizon agents based on the current state, reducing token usage by 52.27% and achieving a 12.67× speedup in representation extraction while maintaining task performance.
The paper proposes principled methods for context representation in large-scale AI reasoning, introducing R3Con which outperforms baselines and enables smaller models to achieve performance comparable to larger ones at lower cost.
Almanac Health is launching an AI agent platform that helps doctors delegate administrative and research tasks, thereby increasing time for patient care.
FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.
The GGUF quantized version of the Qwen-Image-2.1 AI model is released, featuring Dynamic 2.0 technology for efficient 4-bit quantization, supporting text-to-image and transparent image generation in a 4.2GB size suitable for Mac users.
This paper proposes GittinsEval, a cost-aware Bayesian bandit framework for efficient LLM evaluation that significantly reduces costs while maintaining high performance by adaptively selecting configurations.
The paper introduces CounterCredit, a training method for vision-language agents that ensures visual calls are both needed and used, leading to higher performance and fewer spurious calls on benchmarks.