Tag
This paper introduces HalluPeer, a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews, providing annotated data to evaluate and improve detection methods.
Astronomers have discovered a 10-sided polygonal structure in Saturn's southern atmosphere, similar to the known hexagon at the north pole, indicating that such polygonal waves can form in both polar regions.
The paper introduces Sci-ZSEL, a cost-efficient zero-shot scientific entity linking framework that selectively uses LLMs and an ontology-aware filter to enhance performance on benchmarks with low lexical overlap.
This paper introduces Scientific Agent Skills, an open-source library of 163 procedural knowledge skills across 16 scientific domains to help AI research agents perform defensible analyses.
Qwen3.8-Max has been upgraded to version 0902 with further post-training on coding and coworking, resulting in improved performance for complex enterprise tasks, scientific research, and long horizon workflows.
A private group led by Philip Johnston announces the Fermi Explorer mission to send a low-cost spacecraft to Alpha Centauri to investigate the Fermi paradox, aiming for a launch by 2029 using existing technology.
AutoSciRub is an evaluation-first framework that improves autonomous scientific agents by generating task-specific executable rubrics to guide experiments and verification, achieving consistent performance gains on benchmarks.
The Nancy Grace Roman Space Telescope has launched successfully, set to survey the universe 1,000 times faster than Hubble to study dark matter, dark energy, and exoplanets using its advanced infrared camera and coronagraph system.
Google DeepMind publishes a paper on using AI agents for real-world scientific research, showing they can outperform frontier models in computer science tasks like HealthBench Hard.
Terminal-Bench-Science is a benchmark developed by Stanford University researchers to evaluate AI agents on real scientific research workflows, aiming to drive AI capabilities in science.
Terminal-Bench-Science is a new benchmark released as version 0.1.0 for evaluating AI agents on scientific research workflows, led by a Stanford community effort.
The article reports on a deadly flood in Nepal likely triggered by rising temperatures causing glacier collapse, highlighting the broader scientific links between climate change and glacial instability.
Tried the new update of Apodex 1.1 and found significant improvements in task processing, 3D molecular view rendering, and incremental computation, making it very suitable for researchers.
This paper presents a framework for ensuring epistemic legitimacy and accountability in research assisted by large language models, emphasizing the importance of human verification and ownership.
Scientists are using environmental DNA collected from spiderwebs to monitor biodiversity and conservation, with artificial webs showing promise as a cost-effective alternative for eco-surveillance.
The article discusses how AI tools like ChatGPT are reinforcing the Matthew effect in scientific citations by repeatedly referencing popular papers, similar to human behavior, which may hinder innovation in research.
Inherent, a startup founded by DeepMind alumni, claims its AI agent Faraday outperformed Anthropic and OpenAI models in replicating scientific research, using a smaller model trained with reinforcement learning to develop 'research taste'.
This paper proposes a Rust-based multi-agent architecture that uses LLM hallucinations as a feature to generate and evaluate scientific hypotheses, comparing its performance against direct prompting and other methods.
Scientists have released the largest 2D map of the universe, a 5.6-trillion-pixel image covering 75% of the sky, to facilitate research on dark energy and other cosmic phenomena.
The article speculates on the potential impact of concentrating the global compute used for millions of AI prompts onto a single scientific problem, such as curing cancer, and questions whether this could lead to deeper scientific intelligence.