Trending

Trending stories ranked by heat, importance and recency.

Cards List
#41

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

arXiv cs.AI · 5h ago Cached

This paper introduces SEE, a multimodal benchmark of expert-curated questions for scientific discovery in chemistry, biology, and materials science. Evaluation of 19 MLLMs shows the best model reaches only 48.7% accuracy, and even with tool use only 52.7%, revealing that current models lack reliable evidence-bounded scientific reasoning.

0 favorites 0 likes
#42

Fast LapSum: Exact Differentiable Top-k at Million Scale

arXiv cs.AI · 5h ago Cached

Fast LapSum introduces an exact differentiable top-k operator that runs efficiently at million scale on GPUs, enabling practical use in sparse routing, retrieval, and large-scale optimization. The method preserves exact selection mass while remaining fully differentiable and demonstrates order-of-magnitude speedups in applications like megapixel sparse adversarial examples.

0 favorites 0 likes
#43

From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning

arXiv cs.AI · 5h ago Cached

This arXiv paper introduces the Edge-Conditioned Spectral Operator (ESO), a spectral neural operator that uses local edge-wise variations to adapt global spectral mixing, improving performance on physics-sensitive PDE benchmarks.

0 favorites 0 likes
#44

CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems

arXiv cs.AI · 5h ago Cached

Introduces CEDAR, an autonomous method that uses LLM agents with Monte Carlo Tree Search to discover complex systems satisfying user-specified behavioral goals, reducing human effort and enabling goal-directed design.

0 favorites 0 likes
#45

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

arXiv cs.AI · 5h ago Cached

Surg-UniWorld is a unified surgical world model with multimodal control experts, enabling controllable generation of coherent instrument-tissue interaction videos using edge, depth, and optical-flow inputs. It introduces a new benchmark (Cholec80-SurgWAM) and outperforms existing controllable video generation methods.

0 favorites 0 likes
#46

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

arXiv cs.AI · 5h ago Cached

This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.

0 favorites 0 likes
#47

From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

arXiv cs.AI · 5h ago Cached

This paper addresses the growing threat of pure-synthesis fake news videos generated by text-to-video models, introducing a new ternary classification task and the first pure-synthesis fake news video dataset (PS-FNVD), along with a Reasoning-guided framework (R-T2V) that achieves state-of-the-art detection accuracy.

0 favorites 0 likes
#48

WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance

arXiv cs.AI · 5h ago Cached

WebRider is a hierarchical framework that formalizes delegated web tasks as intent contracts, preserving persona-conditioned policies through every browsing step. It includes RiderBench, a benchmark of 4,096 live-web contracts, and an 8B action-policy model trained through its guarded interface.

0 favorites 0 likes
#49

A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers

arXiv cs.AI · 5h ago Cached

This paper introduces CGMas, a multi-agent LLM framework that automates coarse-grained molecular dynamics for polymers, including topology construction, equilibration, mapping, potential derivation, and validation. It completed 27 polymer tasks and matched atomistic densities within 5% in most cases, drastically reducing simulation time.

0 favorites 0 likes
#50

CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

arXiv cs.AI · 5h ago Cached

CellWorld introduces a latent-space predictive pretraining approach for spatial transcriptomics foundation models, predicting latent representations of masked cells instead of reconstructing gene measurements. Across held-out datasets, even small variants outperform existing baselines on all benchmarks, showing that scaling and broad biological diversity improve transferability.

0 favorites 0 likes
#51

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

arXiv cs.AI · 5h ago Cached

TRACE is a new multi-layer benchmark for diagnosing drift and failures in human-AI controller coordination, built from ALFRED traces with 1,918 drifted samples annotated across five execution layers. Baseline results show drift identification and attribution well above random baselines across classical, recurrent, and attention-based model families.

0 favorites 0 likes
#52

ADIAS: Automated Design of Interactive Agentic Systems

arXiv cs.AI · 5h ago Cached

ADIAS is a framework for automated design of agentic systems that uses issue-centric optimization, maintaining a persistent issue state across repair rounds. It outperforms the strongest baseline by 25.2% on average across five interactive benchmarks and shows consistent gains with four backbone models.

0 favorites 0 likes
#53

Auto mode is now the default in Claude Code

Hacker News Top · 6h ago Cached

Anthropic is making auto mode the default in Claude Code for Pro, Max, and Team plans, citing safety research and productivity gains, with classifier overhead now free for those plans.

0 favorites 0 likes
#54

Long-Run Effects of H-1B Immigration on the U.S. Economy (July 2026)

Hacker News Top · 5h ago Cached

NBER working paper examining the long-run economic effects of H-1B immigration on the U.S. economy, released July 2026.

0 favorites 0 likes
#55

@Tesla: FSD Supervised can drive you for over 25,000 miles without you having to touch the wheel

X AI KOLs Timeline · 5h ago Cached

Tesla celebrates the first owner to complete 25,000 miles on the FSD Supervised streak counter without touching the wheel, highlighting the system's reliability.

0 favorites 0 likes
#56

@NVIDIAAP: AI agents work in loops. The model reasons → the CPU executes → the result comes back → repeat. Every step depends on t…

X AI KOLs Timeline · 5h ago Cached

NVIDIA highlights that AI agents operate in sequential loops where single-threaded speed matters more than core count, and introduces the Vera CPU designed to deliver maximum single-threaded performance across all 88 cores at scale.

0 favorites 0 likes
#57

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

arXiv cs.CL · 5h ago Cached

CoinRAG is a new method for long-context RAG that reuses fine-grained contextualized information nugget KV caches instead of full chunks, improving efficiency and answering quality. It achieves a new Pareto frontier with 5.3% relative F1 improvement on LongBench multi-hop QA tasks.

0 favorites 0 likes
#58

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

arXiv cs.CL · 5h ago Cached

Presents LitTraceQA, a benchmark for scientific question answering that requires systems to retrieve relevant papers, locate supporting evidence, and produce verified answers in multiple formats.

0 favorites 0 likes
#59

Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

arXiv cs.CL · 5h ago Cached

This paper introduces a concept-centric benchmark to probe LLMs' understanding of geo-spatial concepts like direction, distance, and topology, testing abstraction, compositionality, and grounding across various model architectures and scales. Findings reveal clear limitations in current LLMs' conceptual understanding.

0 favorites 0 likes
#60

Natural Language Processing Psychometrics

arXiv cs.CL · 5h ago Cached

This paper introduces NLP Psychometrics, a framework that treats psychological prediction from text as a psychometric problem. Using LLM personas, emotional profiles, and syntactic-semantic networks with random forest regressors, it explains up to 76% of variance in mental health scores and shows promise and limits of synthetic data for psychometric prediction.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback