Newest

All articles, most recently crawled first.

Cards List

A digestion of the proof of Sendov's conjecture

Hacker News Top · 3d ago Cached

This article discusses the proof of Sendov's conjecture and its variants, focusing on polynomials with zeros in the unit disk and their critical points.

0 favorites 0 likes

California's new tire efficiency rules could save drivers $1B a year

Hacker News Top · 1h ago Cached

California has approved new tire efficiency standards to save drivers money and reduce emissions, with phased implementation starting in 2029 and stricter standards in 2033.

0 favorites 0 likes

Israel creates fake think tank in likely attempt to dupe AI chatbots

Hacker News Top · 7h ago Cached

Israel created a fake think tank to produce reports designed to influence AI chatbots like Claude and Gemini, a practice referred to as LLM poisoning.

0 favorites 0 likes

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Hugging Face Daily Papers · 5d ago Cached

NaviDC-OCR is a unified vision-language framework that improves document parsing accuracy through deformation-aware learning and adaptive sampling, achieving state-of-the-art results on multiple benchmarks.

0 favorites 0 likes

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

Hugging Face Daily Papers · 4d ago Cached

MegaParts introduces a scalable framework for part-aware 3D object generation using token-efficient vector-quantized tokens and autoregressive modeling, enabling generation of objects with up to 300 parts.

0 favorites 0 likes

Understanding Cognition-Induced Risks in Agentic AI Systems

Hugging Face Daily Papers · 3d ago Cached

The paper systematically analyzes risks in agentic AI systems induced by expanding cognitive capabilities across physical, social, and self-referential levels, and proposes mitigation strategies to ensure their safe development.

0 favorites 0 likes

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Hugging Face Daily Papers · 3d ago Cached

This paper proposes VibeWorlding, a framework for benchmarking and training multimodal agents to construct 3D open worlds from user queries, showing that reinforcement learning improves open-source models to compete with closed-source frontiers.

0 favorites 0 likes

GenRouter: Unified Workflow Routing for Agentic Image Generation

Hugging Face Daily Papers · yesterday Cached

GenRouter is a unified routing framework for agentic image generation that adaptively directs prompts to optimal workflows, significantly reducing costs and latency while improving visual alignment through demand profiling and self-evolution.

0 favorites 0 likes

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

Hugging Face Daily Papers · yesterday Cached

R^3-Bench introduces a benchmark showing that LLMs struggle to maintain performance when allocated shared computation budgets across multiple reasoning tasks like math, coding, and abstract reasoning.

0 favorites 0 likes

ClawGym II: Exploring Black-Box RL on Agent Harness

Hugging Face Daily Papers · yesterday Cached

This paper introduces a unified black-box reinforcement learning framework for stable and scalable optimization of agents through complex harnesses, using sandbox execution and trajectory reconstruction with improvements on benchmarks like ClawGym-Bench.

0 favorites 0 likes

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

Hugging Face Daily Papers · 2026-08-11 Cached

ENTLORE is a graph-grounded benchmark framework for enterprise question answering that evaluates latent organizational reasoning, showing that even with gold document sources, many implicit relation questions remain unanswered.

0 favorites 0 likes

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

Hugging Face Daily Papers · yesterday Cached

Ventor-QTest proposes a black-box audit method for vendor-hosted LLM APIs, using repeated and long-sequence probes to measure fidelity loss and detect degradation in long-horizon agentic tasks.

0 favorites 0 likes

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

Hugging Face Daily Papers · 2d ago Cached

ConceptFormer learns continuous latent concept representations for visual document retrieval, bridging visual evidence and semantic relevance without text intermediates, achieving significant improvements over baselines.

0 favorites 0 likes

When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse

Hugging Face Daily Papers · 2026-08-07 Cached

This paper proposes D-SCAN, a detection framework for RAG poisoning that monitors attention collapse dynamics in language model generations, showing effectiveness against adversarial attacks even when outputs appear benign.

0 favorites 0 likes

Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

Hugging Face Daily Papers · 2026-08-02 Cached

This position paper argues that AI agents in scientific teams should be studied as human-agent systems to enhance collaboration and mitigate risks such as reduced diversity in scientific inquiry.

0 favorites 0 likes

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

Hugging Face Daily Papers · 2d ago Cached

UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstrations to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks.

0 favorites 0 likes

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

Hugging Face Daily Papers · 6d ago Cached

VideoGAIA introduces a benchmark for assessing agentic video understanding in multimodal models through complex, multi-turn tasks, revealing that even frontier models like GPT-5.5 achieve less than 60% accuracy.

0 favorites 0 likes

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Hugging Face Daily Papers · 4d ago Cached

PACE-Bench introduces a simulator-grounded benchmark for evaluating self-evolving agents on physics adaptation tasks involving iterative code redesign after environmental mutations, revealing that mechanism redesign is a major bottleneck compared to parameter inference.

0 favorites 0 likes

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

Simon Willison's Blog · 4h ago Cached

The Qwen 3.8 27B AI model achieved a score of 52 on the Artificial Analysis Intelligence Index, as highlighted in a blog post by Simon Willison.

0 favorites 0 likes

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve

Hugging Face Daily Papers · yesterday Cached

The paper proposes improvements to the optimization problem in combination loss analysis using modern techniques and AlphaEvolve, yielding an improved upper bound on the matrix multiplication exponent.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback