@ModelScope2022: SciJudge-30B and 4B learn to predict which scientific work will carry stronger citation impact. License: Apache-2.0 30B…
Summary
ModelScope releases SciJudge-30B and SciJudge-4B, models trained on millions of arXiv papers to predict citation impact, achieving state-of-the-art accuracy surpassing larger models like GPT-5.2 and Gemini 3 Pro.
View Cached Full Text
Cached at: 07/03/26, 02:29 AM
SciJudge-30B and 4B learn to predict which scientific work will carry stronger citation impact. License: Apache-2.0
30B https://modelscope.ai/models/openmoss/SciJudge-30B-2605… 4B https://modelscope.ai/models/openmoss/SciJudge-4B-2605… https://modelscope.ai/papers/2603.14473…
Scientific Judge accuracy: SciJudge-30B reaches 80.6 in-domain, surpassing GPT-5.2, GLM-5, and Gemini 3 Pro; SciJudge-4B also outperforms much larger baseline models
Data signal: built from 2.1M arXiv papers and 696,758 field- and time-matched citation-based preference pairs
Training: GRPO with DAPO loss and citation-based pairwise rewards
Similar Articles
@AlphaSignalAI: A 4B model can now anticipate scientific breakthroughs before scientists do. Researchers often build breakthroughs by c…
A new paper introduces GIANTS-4B, a 4-billion-parameter model trained with reinforcement learning to predict scientific insights by combining ideas from foundational papers, achieving higher similarity and citation potential than larger models like Gemini 3 Pro.
MOOSE-Star (ICML 2026): 7B model + 108K-paper dataset for scientific hypothesis discovery
MOOSE-Star presents a 7B model fine-tuned from DeepSeek-R1-Distill-Qwen-7B for scientific hypothesis discovery, along with a dataset of 108K NCBI papers. The model achieves state-of-the-art inspiration retrieval accuracy, outperforming larger models like GPT-5.4 and Gemini-3 Pro.
Reward Modeling for Scientific Writing Evaluation
This paper proposes SciRM, cost-efficient open-source reward models tailored for evaluating scientific writing through a two-stage training framework that optimizes evaluation preferences and reasoning capabilities. The models generalize across diverse scientific writing tasks without requiring task-specific retraining, addressing limitations of existing LLM-based judges on domain-specific evaluation criteria.
@jinyuhou0: On popular benchmarks, our 30B model matches systems 20-30x its size (gpt-5.4-xhigh, DeepSeek-V3.2, Kimi-K2.5), while u…
A new 30B model matches systems 20-30x its size on popular benchmarks while using up to 95% fewer reasoning tokens than comparable agentic LLMs, achieved through a learned configurator that decides when and how to reason. Model and code are openly available.
@VikParuchuri: We're open sourcing a 9B model that extracts structured data from documents at near-frontier performance. - 90.2% on ou…
Vik Paruchuri is open-sourcing a 9B model that extracts structured data from documents with near-frontier performance (90.2% on their benchmark, vs Gemini 3.5 Flash at 91.3%).