All articles, most recently crawled first.
CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.
MMDiff uses multimodal sparse autoencoders to isolate, detect, and control features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.
MobileMem introduces a benchmark and framework for on-device AI systems to learn from year-long mobile experiences, focusing on temporal reasoning, knowledge updating, and preference inference.
PRM-as-a-Judge 1.5 is a toolkit that provides fine-grained metrics and reliability tools for evaluating embodied robotic models, moving beyond binary success rates to assess process progress and execution quality.
This paper introduces Latent On-Policy Self-Distillation (LOPD), a method that makes the teacher's privileged context learnable end-to-end from experience, providing dense token-level supervision to enhance agent performance and efficiency in agentic tool use and code generation.
Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during LLM agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy.
The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.
UniProbe is a lightweight detector that uses graph-based multi-structural internal representations to identify hallucinated tokens in large vision-language models, reducing object hallucinations by up to 55% with minimal latency increase.
This paper introduces LittleLearner, a 5B-parameter language model trained on a curated elementary-grade corpus, to study knowledge acquisition in a controlled sandbox environment.
Vocal Slice is a desktop application for audio editing that uses on-device Whisper transcription to allow users to select and export clips by highlighting text in the transcript, designed for podcasters and voice professionals.
The article reports that PJM's modeling errors in electricity market auctions have led to $12 billion in unnecessary costs for ratepayers between 2025 and 2027, due to underestimating existing power plant capacities and flawed auction design.
This 2024 paper examines the effects of strong gravitational lensing and microlensing on supernovae, providing potential insights into cosmological phenomena.
nurb enables users to design 3D-printable parts through natural language conversation with an AI, handling design, checks, and revisions without traditional CAD software.
The fourth edition of Sheldon Axler's 'Linear Algebra Done Right' textbook is now available for free as Open Access in multiple languages, featuring new exercises and improvements, with print versions also available.
Tesla Model Y was China's best-selling midsize SUV last month with 25,158 retail sales, and ranked third overall in vehicle sales.
A Li Auto L8 owner was frustrated when the vehicle scraped the adjacent car while using voice command for automatic parking. The article also notes that Li Auto was formerly Lifan Motor, which began by manufacturing motorcycles and three-wheelers.
The article critiques the personality-driven approach to AI safety and advocates for a rules-based, government-backed bureaucracy to address systemic issues in the AI industry.
DeepSeek Harness is an open-source GitHub project with over 114k stars, offering a plugin-based architecture where models, tools, and configurations can be swapped at the config layer.
CZ from Binance admitted misjudging a feature in TrustWallet, and the team informed him of an existing 'Ignore coins' feature that may be updated soon.
MARGINAL is an open-source governance layer for coding agents that monitors agent trajectories to prevent inefficient actions, with features like shadow mode and earned enforcement to improve reliability.