Newest

All articles, most recently crawled first.

Cards List

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

Hugging Face Daily Papers · 3d ago Cached

CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.

0 favorites 0 likes

Multimodal Model Diffing for Feature Discovery and Control

Hugging Face Daily Papers · 2026-08-10 Cached

MMDiff uses multimodal sparse autoencoders to isolate, detect, and control features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.

0 favorites 0 likes

MobileMem: Learning from a Year of Mobile Experiences

Hugging Face Daily Papers · 6d ago Cached

MobileMem introduces a benchmark and framework for on-device AI systems to learn from year-long mobile experiences, focusing on temporal reasoning, knowledge updating, and preference inference.

0 favorites 0 likes

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Hugging Face Daily Papers · 3d ago Cached

PRM-as-a-Judge 1.5 is a toolkit that provides fine-grained metrics and reliability tools for evaluating embodied robotic models, moving beyond binary success rates to assess process progress and execution quality.

0 favorites 0 likes

Latent On-Policy Self-Distillation

Hugging Face Daily Papers · 4d ago Cached

This paper introduces Latent On-Policy Self-Distillation (LOPD), a method that makes the teacher's privileged context learnable end-to-end from experience, providing dense token-level supervision to enhance agent performance and efficiency in agentic tool use and code generation.

0 favorites 0 likes

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

Hugging Face Daily Papers · 3d ago Cached

Second Thought is a training-free framework that runs auxiliary reasoning branches in parallel during LLM agent action-observation waits to reduce sequential decoding and turn counts without harming accuracy.

0 favorites 0 likes

Forecast Collapse in Time-Series Foundation Models

Hugging Face Daily Papers · 3d ago Cached

The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.

0 favorites 0 likes

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

Hugging Face Daily Papers · 5d ago Cached

UniProbe is a lightweight detector that uses graph-based multi-structural internal representations to identify hallucinated tokens in large vision-language models, reducing object hallucinations by up to 55% with minimal latency increase.

0 favorites 0 likes

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Hugging Face Daily Papers · 4d ago Cached

This paper introduces LittleLearner, a 5B-parameter language model trained on a curated elementary-grade corpus, to study knowledge acquisition in a controlled sandbox environment.

0 favorites 0 likes

Show HN: Vocal Slice – Cut audio by selecting text, fully on-device

Hacker News Top · 6d ago Cached

Vocal Slice is a desktop application for audio editing that uses on-device Whisper transcription to allow users to select and export clips by highlighting text in the transcript, designed for podcasters and voice professionals.

0 favorites 0 likes

$12B of US ratepayers' money wasted on a modeling mistake in PJM

Hacker News Top · 3h ago Cached

The article reports that PJM's modeling errors in electricity market auctions have led to $12 billion in unnecessary costs for ratepayers between 2025 and 2027, due to underestimating existing power plant capacities and flawed auction design.

0 favorites 0 likes

Strong gravitational lensing and microlensing of supernovae (2024)

Hacker News Top · 6h ago

This 2024 paper examines the effects of strong gravitational lensing and microlensing on supernovae, providing potential insights into cosmological phenomena.

0 favorites 0 likes

Design 3D-printable parts by talking

Hacker News Top · 2d ago Cached

nurb enables users to design 3D-printable parts through natural language conversation with an AI, handling design, checks, and revisions without traditional CAD software.

0 favorites 0 likes

Linear Algebra Done Right – Sheldon Axler

Hacker News Top · 3h ago Cached

The fourth edition of Sheldon Axler's 'Linear Algebra Done Right' textbook is now available for free as Open Access in multiple languages, featuring new exercises and improvements, with print versions also available.

0 favorites 0 likes

@cb_doge: Tesla Model Y was China’s #1 best-selling midsize SUV last month. • 25,158 retail sales • #3 best-selling vehicle overa…

X AI KOLs Following · 19h ago Cached

Tesla Model Y was China's best-selling midsize SUV last month with 25,158 retail sales, and ranked third overall in vehicle sales.

0 favorites 0 likes

@cuichenghao: Today, a Li Auto L8 owner in mainland China used a voice command to automatically park out, but ended up scraping the adjacent car, leaving the owner extremely frustrated. Li Auto was formerly Lifan Motor, which started by manufacturing motorcycles and three-wheelers.

X AI KOLs Following · 6h ago Cached

A Li Auto L8 owner was frustrated when the vehicle scraped the adjacent car while using voice command for automatic parking. The article also notes that Li Auto was formerly Lifan Motor, which began by manufacturing motorcycles and three-wheelers.

0 favorites 0 likes

@joshua_saxe: This deserves far more attention; it's a happy, path dependent quirk of history that the AI industry is led by people w…

X AI KOLs Following · 7h ago Cached

The article critiques the personality-driven approach to AI safety and advocates for a rules-based, government-backed bureaucracy to address systemic issues in the AI industry.

0 favorites 0 likes

@Saboo_Shubham_: Crazy...DeepSeek Harness is already at 114k+ stars on GitHub. Everything is a plugin. Model, tools, skills, sessions, s…

X AI KOLs Following · yesterday Cached

DeepSeek Harness is an open-source GitHub project with over 114k stars, offering a plugin-based architecture where models, tools, and configurations can be swapped at the config layer.

0 favorites 0 likes

@cz_binance: I mis-judged that. A few people mentioned it is a useful feature. @TrustWallet team saw this too. They told me there is…

X AI KOLs Following · 4h ago Cached

CZ from Binance admitted misjudging a feature in TrustWallet, and the team informed him of an existing 'Ignore coins' feature that may be updated soon.

0 favorites 0 likes

Evidence-based governor for coding agents — looking for people to try it and constructive feedback

Reddit r/AI_Agents · 3h ago

MARGINAL is an open-source governance layer for coding agents that monitors agent trajectories to prevent inefficient actions, with features like shadow mode and earned enforcement to improve reliability.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback