Tag
The article outlines nine prompt discipline rules that reduced wasted thinking in coding agents by up to 70% based on A/B tests on GLM 5.3 and GLM 5.3 Flash models, with a self-testable exam available on GitHub.
A tweet suggests that small specialized classifiers can be deployed without hosting and trained directly on mobile devices, highlighting advancements in lightweight AI models.
The author suggests that heavy AI users may inefficiently allocate model capacity by not optimizing model selection, proposing that work should be routed to the least expensive capable model to save costs and improve efficiency.
An article critiques the overuse of AI agents for deterministic tasks, sharing a case study where a multi-agent customer support system was refactored with simpler code, resulting in lower latency and costs.
Model Grafting technique modifies Qwen3.5-4B into a causal encoder-decoder, creating variants with up to 3.7x speedup in prompt processing and minimal accuracy loss.
Jev-Mem introduces an agentic memory architecture inspired by System-One/System-Two cognition, enhancing efficiency and effectiveness for long-horizon AI agents with improved scores and faster operations.
The dedicated evaluation model JEV demonstrated high efficiency and low-cost potential in processing interview transcripts, emphasizing the advantages of using large language models as logical judgment layers rather than content generators.
Bending Spoons, a Milan-based private equity firm, has acquired multiple software companies including Miro and Airtable, using AI to cut costs and improve profitability after a Nasdaq IPO.
NVIDIA explains how optimizing power management for AI factories using DSX Flex and partner platforms can boost efficiency, with Lambda validating a 24% throughput increase on a fixed power budget.
ByteShape released ShapeLearn GGUFs for Qwen 3.8 27B, achieving high accuracy on benchmarks, and discussed the role of KLD in quantization, with a related paper accepted at EMNLP 2026.
A tweet asking why FlashAttention achieves both speed and GPU memory efficiency in AI computations.
The article argues that OpenAI's claimed 3x AI productivity gain comes from AI systems working continuously like extra shifts, rather than making humans more efficient, and highlights the high costs and defect rates involved.
This paper explores how prompt properties like cognitive load and phrasing pattern influence energy usage in on-device LLM inference, showing that cognitive load affects energy per token while phrasing impacts token usage, highlighting the need for model-aware prompt design for energy efficiency.
AWS Cloud promotes the Twitch premiere of 'In The Field', a show exploring a lab where nature is teaching AI to be more efficient, in collaboration with BioComputingCo.
The author shares their experience running a 30B parameter model with EXL3 quantization on a 12GB VRAM GPU, achieving efficient performance and speed for coding and agent tasks.
The article explores how relational buffering—extra tokens from misaligned intentions—might be a significant source of waste in AI interactions, proposing 'tokens per resolved intention' as a metric to reduce computational cost while preserving fidelity.
The article proposes AdaThinking-E, a reinforcement learning framework that uses one-token entropy regulation to enable adaptive thinking in multimodal large language models, improving accuracy on complex tasks and efficiency on simple ones.
IQ Routing is a trajectory-aware LLM routing system designed to reduce the cost of AI agents by optimizing task routing based on trajectory data.
The author describes a token routing system to optimize LLM usage, reducing API costs by up to 70% through techniques like cost-based routing, prompt caching, and output discipline.
This paper conducts the first systematic study of local AI inference efficiency across models and hardware, measuring intelligence per watt and showing a 5.3x improvement from 2023 to 2025, indicating potential for redistributing demand from centralized infrastructure.