Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Summary
Apple-π is a benchmark that evaluates video generation models on their ability to reason about physical laws through a three-stage protocol: perception, formulation, deduction. It includes 400 videos covering classical mechanics tasks and reveals current models fall short of reliable law-grounded world simulation.
View Cached Full Text
Cached at: 07/21/26, 10:36 AM
Paper page - Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Source: https://huggingface.co/papers/2607.16401 Published on Jul 17
·
Submitted byhttps://huggingface.co/lifuguan
leolion Jul 21
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Modernvideogenerationmodelsareincreasinglyhailedasemergingworldmodelswithaninternalizedgraspofphysicallaw.Yetexistingbenchmarkslargelyevaluatephysicalplausibilityonlyattheoutputlevel,withoutverifyingwhetherthemodelarrivestherethroughafaithful,law-groundedreasoningprocess.WeintroduceApple-PI,thefirstbenchmarkthatanchorsvideo-modelevaluationexplicitlyinphysicallaws.Apple-PIcomprisesthreecomponents.1)Orchard:adatasetof400videoscoveringtencanonicaltasksinclassicalmechanics.Itseparatessingle-lawtasksforconfounder-freediagnosisfrommulti-lawtasksforprobinggeneralization.2)BenchmarkProtocol:athree-stageprotocolbasedonscientificreasoning,includingPerception,Formulation,andDeduction.Ituseschain-of-framespromptingoninfographic-annotatedfirstframes,treatingthegeneratedvideoasthemodel’svisiblereasoningtrace.3)EvaluationSuite:ahybridevaluationsuitethatcombinesMLLM-basedsubjectivescoringwithphysics-law-groundedobjectivemeasures.Thisenablesstage-resolveddiagnosisofnotonlywhetheramodelfails,butwhereitfails.Benchmarking11modelsshowsthatcurrentvideomodelsremainfarfromreliablelaw-groundedworldsimulators,withthebestvideomodelscoringonly0.473.Ourstage-,pillar-,andsource-resolvedanalysesfurtherexposeaPerception-to-Formulation-to-Deductionbottleneck,weakmulti-lawstatetransfer,andapersistentSim-to-Realgap.ThesefindingspositionApple-PIasadiagnosticfoundationforguidingfuturevideomodelstowardworldmodelswithlaw-groundedphysicalintelligence.
View arXiv pageView PDFProject pageGitHub10Add to collection
Get this paper in your agent:
hf papers read 2607\.16401
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.16401 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.16401 in a dataset README.md to link it from this page.
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Principia: Relational Physics Tests for Video Models
Principia is a benchmark that evaluates video models on Newtonian physics using relational consistency between paired objects, revealing significant gaps in current models' physical reasoning.
Physics-IQ Verified
This paper presents a systematic audit of the Physics-IQ benchmark for evaluating physical understanding in video generative models, proposing improvements to prompts and scoring to enhance reliability.
PhysBrain 1.0 Technical Report
PhysBrain 1.0 is a technical report presenting a method that uses human egocentric video to generate physical commonsense supervision for vision-language-action models, achieving state-of-the-art results on embodied control benchmarks including ERQA, PhysBench, SimplerEnv-WidowX, LIBERO, and RoboCasa.
PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
PhysMent introduces an interactive benchmark that evaluates LLM physical reasoning via iterative tool-mediated experimentation with a MuJoCo physics simulator, revealing model weaknesses in multi-step procedural tasks.
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Introduces Video-MME-Logical, a controlled benchmark for evaluating video temporal-logical reasoning in multimodal large language models, revealing a substantial human-model gap.