Trending stories ranked by heat, importance and recency.
This paper introduces SEE, a multimodal benchmark of expert-curated questions for scientific discovery in chemistry, biology, and materials science. Evaluation of 19 MLLMs shows the best model reaches only 48.7% accuracy, and even with tool use only 52.7%, revealing that current models lack reliable evidence-bounded scientific reasoning.
Fast LapSum introduces an exact differentiable top-k operator that runs efficiently at million scale on GPUs, enabling practical use in sparse routing, retrieval, and large-scale optimization. The method preserves exact selection mass while remaining fully differentiable and demonstrates order-of-magnitude speedups in applications like megapixel sparse adversarial examples.
This arXiv paper introduces the Edge-Conditioned Spectral Operator (ESO), a spectral neural operator that uses local edge-wise variations to adapt global spectral mixing, improving performance on physics-sensitive PDE benchmarks.
Introduces CEDAR, an autonomous method that uses LLM agents with Monte Carlo Tree Search to discover complex systems satisfying user-specified behavioral goals, reducing human effort and enabling goal-directed design.
Surg-UniWorld is a unified surgical world model with multimodal control experts, enabling controllable generation of coherent instrument-tissue interaction videos using edge, depth, and optical-flow inputs. It introduces a new benchmark (Cholec80-SurgWAM) and outperforms existing controllable video generation methods.
This paper presents Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy, using specialist training via reinforcement learning followed by weight-space merging and routed policy-space distillation. It introduces a new state verification benchmark and demonstrates strong performance across embodied and general capabilities at 2B and 35B-A3B scales.
This paper addresses the growing threat of pure-synthesis fake news videos generated by text-to-video models, introducing a new ternary classification task and the first pure-synthesis fake news video dataset (PS-FNVD), along with a Reasoning-guided framework (R-T2V) that achieves state-of-the-art detection accuracy.
WebRider is a hierarchical framework that formalizes delegated web tasks as intent contracts, preserving persona-conditioned policies through every browsing step. It includes RiderBench, a benchmark of 4,096 live-web contracts, and an 8B action-policy model trained through its guarded interface.
This paper introduces CGMas, a multi-agent LLM framework that automates coarse-grained molecular dynamics for polymers, including topology construction, equilibration, mapping, potential derivation, and validation. It completed 27 polymer tasks and matched atomistic densities within 5% in most cases, drastically reducing simulation time.
CellWorld introduces a latent-space predictive pretraining approach for spatial transcriptomics foundation models, predicting latent representations of masked cells instead of reconstructing gene measurements. Across held-out datasets, even small variants outperform existing baselines on all benchmarks, showing that scaling and broad biological diversity improve transferability.
TRACE is a new multi-layer benchmark for diagnosing drift and failures in human-AI controller coordination, built from ALFRED traces with 1,918 drifted samples annotated across five execution layers. Baseline results show drift identification and attribution well above random baselines across classical, recurrent, and attention-based model families.
ADIAS is a framework for automated design of agentic systems that uses issue-centric optimization, maintaining a persistent issue state across repair rounds. It outperforms the strongest baseline by 25.2% on average across five interactive benchmarks and shows consistent gains with four backbone models.
Anthropic is making auto mode the default in Claude Code for Pro, Max, and Team plans, citing safety research and productivity gains, with classifier overhead now free for those plans.
NBER working paper examining the long-run economic effects of H-1B immigration on the U.S. economy, released July 2026.
Tesla celebrates the first owner to complete 25,000 miles on the FSD Supervised streak counter without touching the wheel, highlighting the system's reliability.
NVIDIA highlights that AI agents operate in sequential loops where single-threaded speed matters more than core count, and introduces the Vera CPU designed to deliver maximum single-threaded performance across all 88 cores at scale.
CoinRAG is a new method for long-context RAG that reuses fine-grained contextualized information nugget KV caches instead of full chunks, improving efficiency and answering quality. It achieves a new Pareto frontier with 5.3% relative F1 improvement on LongBench multi-hop QA tasks.
Presents LitTraceQA, a benchmark for scientific question answering that requires systems to retrieve relevant papers, locate supporting evidence, and produce verified answers in multiple formats.
This paper introduces a concept-centric benchmark to probe LLMs' understanding of geo-spatial concepts like direction, distance, and topology, testing abstraction, compositionality, and grounding across various model architectures and scales. Findings reveal clear limitations in current LLMs' conceptual understanding.
This paper introduces NLP Psychometrics, a framework that treats psychological prediction from text as a psychometric problem. Using LLM personas, emotional profiles, and syntactic-semantic networks with random forest regressors, it explains up to 76% of variance in mental health scores and shows promise and limits of synthetic data for psychometric prediction.