Tag
This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.
StateComp is a framework that compresses historical interactions in long-horizon agents based on the current state, reducing token usage by 52.27% and achieving a 12.67× speedup in representation extraction while maintaining task performance.
The paper proposes principled methods for context representation in large-scale AI reasoning, introducing R3Con which outperforms baselines and enables smaller models to achieve performance comparable to larger ones at lower cost.
FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.
The GGUF quantized version of the Qwen-Image-2.1 AI model is released, featuring Dynamic 2.0 technology for efficient 4-bit quantization, supporting text-to-image and transparent image generation in a 4.2GB size suitable for Mac users.
This paper proposes GittinsEval, a cost-aware Bayesian bandit framework for efficient LLM evaluation that significantly reduces costs while maintaining high performance by adaptively selecting configurations.
The paper introduces CounterCredit, a training method for vision-language agents that ensures visual calls are both needed and used, leading to higher performance and fewer spurious calls on benchmarks.
AIBuildAI-2.5 introduces an autonomous AI model development system using LLM-guided tree search to enhance efficiency, ranking first on MLE-Bench with a 73.3% medal rate and outperforming baselines on multiple tasks.
FLEET introduces a memory mechanism to text generation in large language models, using logits entropy to enhance trajectories, resulting in improved accuracy and a 3x speedup, particularly on coding tasks.
The Biological Computing Co. partners with AWS to commercialize a neuron-derived AI video model that optimizes text-to-video generation using biological data, offering faster and cheaper inference on standard hardware.
Anthropic's Opus 5.5 and OpenAI's GPT-6 Sol and Luna models promise similar performance with significant cost reductions, making advanced AI more accessible.
Anthropic reports that their Claude Opus 5.5 model may detect evaluation scenarios, complicating the generalization of observed behavior to real deployments. The model offers performance comparable to Fable 5.1 with a 40% cost reduction and faster output.
Diogo Amogo, founder of Jev, released a PDF blueprint for building a Jev Harness to enhance coding agents, claiming to make them 200× faster and 400× cheaper.
AgentJev-0.6B is a small open-source AI decision model that outperforms Laya on Typed Decisions with improved accuracy and significant computation reduction, enabling faster local inference.
The paper proposes Independent Token Sampling (ITS), a query-efficient framework for detecting training data in diffusion large language models by selecting weakly dependent tokens to reduce approximation errors, achieving improved performance over state-of-the-art baselines.
Yohei Nakajima presents a method to read typed visual judgements from a frozen open vision-language model using logits, achieving similar accuracy to hosted models with reduced time and GPU cost.
This tweet recommends a Stanford course on efficient generative language models, covering techniques from pre-training to inference to balance performance and cost with limited compute.
This paper presents a decomposition framework for quantization error in large language models, separating it into activation-guided weight compensation and orthogonal residual, and derives practical guidelines for improving W4A4 quantization through techniques like Hadamard rotation and sign selection.
RBS-Attention introduces a training-free sparse-prefill method with dual-branch selection to mitigate mean dilution in long-context LLM inference, achieving up to 20.65× speedup on H100 GPUs while maintaining near-dense quality on benchmarks.
The paper proposes optimizations to the Viterbi algorithm using the Hirschberg algorithm and constrained random walk, reducing memory usage from 140 GB to 5 MB and improving speed, enabling forced alignment to run on end-user devices for better scalability in speech processing.