scientific-research

Tag

Cards List
#scientific-research

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

arXiv cs.AI ↗ · 2026-09-04 Cached

This paper introduces HalluPeer, a taxonomy-driven benchmark for detecting hallucinations in scientific peer reviews, providing annotated data to evaluate and improve detection methods.

0 favorites 0 likes
#scientific-research

Astronomers Detect a 10-Sided Structure in Saturn's Atmosphere

Hacker News Top ↗ · 2026-09-03 Cached

Astronomers have discovered a 10-sided polygonal structure in Saturn's southern atmosphere, similar to the known hexagon at the north pole, indicating that such polygonal waves can form in both polar regions.

0 favorites 0 likes
#scientific-research

Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking

arXiv cs.CL ↗ · 2026-09-02 Cached

The paper introduces Sci-ZSEL, a cost-efficient zero-shot scientific entity linking framework that selectively uses LLMs and an ontology-aware filter to enhance performance on benchmarks with low lexical overlap.

0 favorites 0 likes
#scientific-research

Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents

arXiv cs.CL ↗ · 2026-09-02 Cached

This paper introduces Scientific Agent Skills, an open-source library of 163 procedural knowledge skills across 16 scientific domains to help AI research agents perform defensible analyses.

0 favorites 0 likes
#scientific-research

Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902!

Reddit r/singularity ↗ · 2026-09-02

Qwen3.8-Max has been upgraded to version 0902 with further post-training on coding and coworking, resulting in improved performance for complex enterprise tasks, scientific research, and long horizon workflows.

0 favorites 0 likes
#scientific-research

Private group wants to launch "cheapest possible" mission to Alpha Centauri

Ars Technica ↗ · 2026-09-01 Cached

A private group led by Philip Johnston announces the Fermi Explorer mission to send a low-cost spacecraft to Alpha Centauri to investigate the Fermi paradox, aiming for a launch by 2029 using existing technology.

0 favorites 0 likes
#scientific-research

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

AutoSciRub is an evaluation-first framework that improves autonomous scientific agents by generating task-specific executable rubrics to guide experiments and verification, achieving consistent performance gains on benchmarks.

0 favorites 0 likes
#scientific-research

The Nancy Grace Roman Space Telescope launches to study dark matter and dark energy

The Verge ↗ · 2026-08-30 Cached

The Nancy Grace Roman Space Telescope has launched successfully, set to survey the universe 1,000 times faster than Hubble to study dark matter, dark energy, and exoplanets using its advanced infrared camera and coronagraph system.

0 favorites 0 likes
#scientific-research

@dair_ai: Impressive new work from Google DeepMind showing real-world applications of agents for scientific research.

X AI KOLs Timeline ↗ · 2026-08-28 Cached

Google DeepMind publishes a paper on using AI agents for real-world scientific research, showing they can outperform frontier models in computer science tasks like HealthBench Hard.

0 favorites 0 likes
#scientific-research

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Hacker News Top ↗ · 2026-08-28 Cached

Terminal-Bench-Science is a benchmark developed by Stanford University researchers to evaluate AI agents on real scientific research workflows, aiming to drive AI capabilities in science.

1 favorites 1 likes
#scientific-research

@himanshustwts: i was trying to understand task design philosophy of TB-science. they have explicitly rejected tasks like - textbook di…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

Terminal-Bench-Science is a new benchmark released as version 0.1.0 for evaluating AI agents on scientific research workflows, led by a Stanford community effort.

0 favorites 0 likes
#scientific-research

How Rising Temperatures Likely Contributed to Nepal’s Deadly Flood

Wired ↗ · 2026-08-26 Cached

The article reports on a deadly flood in Nepal likely triggered by rising temperatures causing glacier collapse, highlighting the broader scientific links between climate change and glacial instability.

0 favorites 0 likes
#scientific-research

@dingyi: Tried out the new update of Apodex 1.1, it has evolved a lot, for example, I directly used the official case 'KRAS G12D mutant combined with small molecule inhibitor conformation' to test. Although I don't understand biomedicine, through the whole process, I can see that many details of this product are indeed very thoughtful. For example: • Automatically split tasks and…

X AI KOLs Timeline ↗ · 2026-08-26 Cached

Tried the new update of Apodex 1.1 and found significant improvements in task processing, 3D molecular view rendering, and incremental computation, making it very suitable for researchers.

0 favorites 0 likes
#scientific-research

Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper presents a framework for ensuring epistemic legitimacy and accountability in research assisted by large language models, emphasizing the importance of human verification and ownership.

0 favorites 0 likes
#scientific-research

Welcome to the spiderverse, a world measured through webs

MIT Technology Review ↗ · 2026-08-25 Cached

Scientists are using environmental DNA collected from spiderwebs to monitor biodiversity and conservation, with artificial webs showing promise as a cost-effective alternative for eco-surveillance.

0 favorites 0 likes
#scientific-research

AI Cites the Same Papers Over and Over Again – Just Like Humans

Reddit r/ArtificialInteligence ↗ · 2026-08-24 Cached

The article discusses how AI tools like ChatGPT are reinforcing the Matthew effect in scientific citations by repeatedly referencing popular papers, similar to human behavior, which may hinder innovation in research.

0 favorites 0 likes
#scientific-research

Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research

TechCrunch AI ↗ · 2026-08-22 Cached

Inherent, a startup founded by DeepMind alumni, claims its AI agent Faraday outperformed Anthropic and OpenAI models in replicating scientific research, using a smaller model trained with reinforcement learning to develop 'research taste'.

0 favorites 0 likes
#scientific-research

Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper proposes a Rust-based multi-agent architecture that uses LLM hallucinations as a feature to generate and evaluate scientific hypotheses, comparing its performance against direct prompting and other methods.

0 favorites 0 likes
#scientific-research

Scientists Release Biggest 2D Map of the Universe

Hacker News Top ↗ · 2026-08-20 Cached

Scientists have released the largest 2D map of the universe, a 5.6-trillion-pixel image covering 75% of the sky, to facilitate research on dark energy and other cosmic phenomena.

0 favorites 0 likes
#scientific-research

What would happen if we gave a single ai problem the compute currently used for millions of prompts?

Reddit r/singularity ↗ · 2026-08-19

The article speculates on the potential impact of concentrating the global compute used for millions of AI prompts onto a single scientific problem, such as curing cancer, and questions whether this could lead to deeper scientific intelligence.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback