scientific-reasoning

Tag

Cards List
#scientific-reasoning

@levie: We’re going to see an increasing divergence between what AI does in our personal lives and in daily productivity vs. wh…

X AI KOLs Following · 2026-08-01 Cached

A commentary on AI capabilities diverging between consumer use and deep domains, citing OpenAI's next model family Astra allegedly solving major open problems in math and science.

0 favorites 0 likes
#scientific-reasoning

@polynoamial: The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices. We’re …

X AI KOLs Following · 2026-08-01 Cached

OpenAI's upcoming Astra model family solved 10 major open problems in mathematics and theoretical computer science, with proof generation costing under $2,000. The tweet highlights Astra's potential for scientific reasoning.

0 favorites 0 likes
#scientific-reasoning

@polynoamial: An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum …

X AI KOLs Following · 2026-08-01 Cached

OpenAI's internal version of its next major model family Astra reportedly solved ten major open problems in mathematics and theoretical computer science, marking a major step for scientific reasoning.

0 favorites 0 likes
#scientific-reasoning

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination

arXiv cs.AI · 2026-07-10 Cached

Introduces G-Frame, a game theory-driven multi-agent framework that reduces hallucinations in lightweight LLMs by internalizing domain constraints, achieving a 79.46% reduction in hallucinations and performance parity with GPT-4o mini on chemistry benchmarks.

0 favorites 0 likes
#scientific-reasoning

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

arXiv cs.LG · 2026-07-02 Cached

Active-GRPO introduces an adaptive imitation and self-improving reasoning framework that dynamically decides when to imitate references and when to reinforce the model's own discoveries for molecular optimization, achieving statistically significant improvements over previous methods on the TOMG-Bench-MolOpt benchmark.

0 favorites 0 likes
#scientific-reasoning

InternScience/Agents-A1 · Hugging Face

Reddit r/LocalLLaMA · 2026-06-30 Cached

Agents-A1 is a 35B Mixture-of-Experts agentic model from InternScience that achieves competitive performance against frontier-scale systems like GPT-5.5 and DeepSeek-V4-pro using long-horizon trajectory scaling and multi-teacher multi-domain distillation.

0 favorites 0 likes
#scientific-reasoning

@ProfBuehlerMIT: For science, AI sovereignty and physics-grounded reasoning are non-negotiable. But how can we teach a small LLM like Ge…

X AI KOLs Timeline · 2026-06-18 Cached

mistral.rs now natively supports Agent Skills, enabling locally-run small LLMs to perform complex agentic workflows for scientific tasks, with full control over models, data, and execution.

0 favorites 0 likes
#scientific-reasoning

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark

Hugging Face Daily Papers · 2026-06-17 Cached

This paper introduces PhySciBench, a benchmark of 200 expert-curated questions for physical sciences, and DelveAgent, a multi-agent framework that improves accuracy and reduces inference costs compared to baselines like Gemini Deep Research.

0 favorites 0 likes
#scientific-reasoning

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

Hugging Face Daily Papers · 2026-06-14 Cached

SciOrch presents an 8B vision-language model trained with MCTS to coordinate multiple expert LLMs for multimodal scientific reasoning, achieving superior performance while reducing API costs.

0 favorites 0 likes
#scientific-reasoning

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

arXiv cs.AI · 2026-06-12 Cached

SciR is a new controllable benchmark for evaluating LLMs on scientific reasoning including deduction, induction, and causal abduction, with parametric control over extraction and inference difficulty. Tests show both axes degrade performance across models, with reasoning models like DeepSeek-R1 outperforming instruct models on inference.

0 favorites 0 likes
#scientific-reasoning

@ChenHenryWu: Self-improvement depends on whether a model can judge its own work. We usually train models to generate better - why no…

X AI KOLs Timeline · 2026-06-05 Cached

This tweet thread introduces research showing that training models to verify their own work can nearly double accuracy on hard math problems and improve scientific reasoning by 14x.

0 favorites 0 likes
#scientific-reasoning

FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

arXiv cs.AI · 2026-06-04 Cached

FalsifyBench is a new evaluation framework for assessing inductive reasoning in LLMs, inspired by the Wason 2-4-6 task, where agents discover hidden semantic rules by proposing examples and receiving feedback. Evaluation of 12 LLMs shows reasoning models outperform instruction-tuned models, with negative testing (hypothesis falsification) being the key driver of success.

0 favorites 0 likes
#scientific-reasoning

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

arXiv cs.AI · 2026-06-04 Cached

SCI-PRM introduces a tool-aware Process Reward Model for scientific reasoning, trained on the SCIPRM70K dataset featuring 'Chain-of-Tool' trajectories that interleave reasoning with scientific tool execution. It enables effective test-time scaling and serves as a dense reward signal in reinforcement learning, outperforming proprietary models like GPT-5-Mini on tool-calling steps across scientific benchmarks.

0 favorites 0 likes
#scientific-reasoning

Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

arXiv cs.AI · 2026-06-04 Cached

Researchers from the University of Michigan introduce MechSim, a mechanism-grounded neuro-symbolic reasoning framework that enables LLM agents to reason about the internal assumptions, dependencies, and execution behavior of scientific simulators rather than treating them as black boxes. The framework improves explanation quality and decision-making reliability across high-stakes domains like healthcare, finance, and public policy.

0 favorites 0 likes
#scientific-reasoning

Domain Adaptation and Reasoning Frameworks in Language Models: A Controlled Experiment with Historical Cosmology

arXiv cs.CL · 2026-06-01 Cached

This paper investigates how domain adaptation reshapes explanatory behavior in language models by training on a pre-Copernican corpus, finding that fine-tuning shifts explanatory framing more than cosmological stance.

0 favorites 0 likes
#scientific-reasoning

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning

arXiv cs.LG · 2026-05-20 Cached

ReCrit introduces a transition-aware reinforcement learning framework for scientific critic reasoning, decomposing initial-to-critic behavior into four quadrants (Correction, Sycophancy, Robustness, Boundary) and using dynamic asynchronous rollout. It improves critic accuracy significantly on Qwen models across multiple scientific benchmarks.

0 favorites 0 likes
#scientific-reasoning

Learning from Language Feedback via Variational Policy Distillation

Hugging Face Daily Papers · 2026-05-18 Cached

Variational Policy Distillation (VPD) formalizes learning from language feedback as a variational EM problem, co-training teacher and student networks to improve policy learning in reinforcement learning from verifiable rewards. It shows consistent improvements over baselines on code generation and scientific reasoning tasks.

0 favorites 0 likes
#scientific-reasoning

@ClementDelangue: Paper of the day! https://huggingface.co/papers/2605.13301…

X AI KOLs Following · 2026-05-15 Cached

A paper introduces a unified recipe (SU-01) that combines reverse-perplexity curriculum, two-stage reinforcement learning, and test-time scaling to achieve gold-medal-level performance on IMO and IPhO problems using a 30B-A3B backbone.

0 favorites 0 likes
#scientific-reasoning

Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling

arXiv cs.AI · 2026-05-14 Cached

This paper presents a simple and unified recipe combining supervised fine-tuning, two-stage reinforcement learning, and test-time scaling to train a reasoning model (SU-01) that achieves gold-medal-level performance on International Mathematical and Physics Olympiad problems.

0 favorites 0 likes
#scientific-reasoning

NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models

arXiv cs.CL · 2026-05-11 Cached

This paper introduces NSMQ Riddles, a novel benchmark using scientific and mathematical riddles from Ghana's National Science and Maths Quiz to evaluate Large Language Models, addressing the underrepresentation of Global South datasets in AI research.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback