SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
Summary
This paper introduces SciGen-Verifier, a multimodal reasoning model for explainable verification of scientific image generation, along with a new benchmark and a reinforcement learning training pipeline.
View Cached Full Text
Cached at: 09/29/26, 08:18 AM
Paper page - SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
Source: https://huggingface.co/papers/2609.33399
Abstract
Inrealisticeducation,asolutionisoftenexpressednotonlyinwordsbutinadrawing--acircuit,ageometricconstruction,afunctionplot--andateachermustgradethedrawingascarefullyasthetext.Recentadvancesinunifiedmultimodalmodelshaveenabledscientificimagegeneration,yetverifyingthecorrectnessofthesespecializedvisualoutputsremainsacriticalbottleneck:errorsoftenarisefromintricatedomainknowledge,structuralreasoning,andmulti-stepinstructionratherthansurface-levelartifacts.Existingverifiersmainlytargetnaturalimagesandcompressjudgementintoscalarscores,leavingscientificcoverageandexplainablefeedbackforerrorcorrectionunderexplored.Tobridgethisgap,wemakethreemaincontributions.(1)WeconstructSciGen-Verify,abenchmarkdedicatedtoexplainableverificationofscientificimagegeneration,spanninginstructionfollowing,multidisciplinaryreasoning,andworldknowledgedomains.Itcontainsathree-tierhierarchicalprotocoloverthebinaryjudgement,supportingexplanation,andcorrectiveeditinginstruction.(2)WedevelopSciGen-Verifier,areasoning-drivenmultimodalverifiertrainedviacold-startsupervisedfine-tuningfollowedbyacurriculum-basedtwo-stagereinforcementlearningpipeline.Therubric-guidedprocessrewardsfirststrengthenscientificreasoningexplorationandoutcomerewardssubsequentlyalignoutputwithground-truthannotation.(3)OnSciGen-Verify,SciGen-Verifierachievescompetitiveperformanceagainstmuchlargerproprietarymodels.Itfurtherservesasapracticalonlinecriticforiterativeimagerectification.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.33399
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33399 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33399 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33399 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
This paper introduces ToolSciVer, the first tool-augmented framework for multimodal scientific claim verification (MSCV), which equips a VLM with type-aware visual tools and trains the policy using GRPO to achieve superior performance on SciVer and MuSciClaims datasets across multiple model families.
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Introduces Sci-VBench, a benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains, finding that visual realism has not translated into reliable scientific and causal correctness.
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Sci-MMR is a benchmark for evaluating multi-step evidence-grounded scientific reasoning in multimodal agents, revealing gaps where answer accuracy exceeds evidence recovery by over 20%.
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
Introduces SVR-R1, a multi-turn reinforcement learning framework that uses the model's own verification as a learning signal for multi-modal reasoning, achieving significant accuracy improvements over standard GRPO baselines on vision-language reasoning benchmarks.
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
SciIR introduces a large-scale dataset (SciIR-82k) and benchmark (SciIR-Bench) to enhance scientific reasoning in text-to-image models, with fine-tuning on Qwen leading to a significant performance improvement.