SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

Hugging Face Daily Papers Papers

Summary

This paper introduces SciGen-Verifier, a multimodal reasoning model for explainable verification of scientific image generation, along with a new benchmark and a reinforcement learning training pipeline.

In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.
Original Article
View Cached Full Text

Cached at: 09/29/26, 08:18 AM

Paper page - SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation

Source: https://huggingface.co/papers/2609.33399

Abstract

Inrealisticeducation,asolutionisoftenexpressednotonlyinwordsbutinadrawing--acircuit,ageometricconstruction,afunctionplot--andateachermustgradethedrawingascarefullyasthetext.Recentadvancesinunifiedmultimodalmodelshaveenabledscientificimagegeneration,yetverifyingthecorrectnessofthesespecializedvisualoutputsremainsacriticalbottleneck:errorsoftenarisefromintricatedomainknowledge,structuralreasoning,andmulti-stepinstructionratherthansurface-levelartifacts.Existingverifiersmainlytargetnaturalimagesandcompressjudgementintoscalarscores,leavingscientificcoverageandexplainablefeedbackforerrorcorrectionunderexplored.Tobridgethisgap,wemakethreemaincontributions.(1)WeconstructSciGen-Verify,abenchmarkdedicatedtoexplainableverificationofscientificimagegeneration,spanninginstructionfollowing,multidisciplinaryreasoning,andworldknowledgedomains.Itcontainsathree-tierhierarchicalprotocoloverthebinaryjudgement,supportingexplanation,andcorrectiveeditinginstruction.(2)WedevelopSciGen-Verifier,areasoning-drivenmultimodalverifiertrainedviacold-startsupervisedfine-tuningfollowedbyacurriculum-basedtwo-stagereinforcementlearningpipeline.Therubric-guidedprocessrewardsfirststrengthenscientificreasoningexplorationandoutcomerewardssubsequentlyalignoutputwithground-truthannotation.(3)OnSciGen-Verify,SciGen-Verifierachievescompetitiveperformanceagainstmuchlargerproprietarymodels.Itfurtherservesasapracticalonlinecriticforiterativeimagerectification.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.33399

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33399 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33399 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33399 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles