Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Summary
Apple-π is a benchmark that evaluates video generation models on their ability to reason about physical laws through a three-stage protocol: perception, formulation, deduction. It includes 400 videos covering classical mechanics tasks and reveals current models fall short of reliable law-grounded world simulation.
View Cached Full Text
Cached at: 07/21/26, 10:36 AM
Paper page - Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Source: https://huggingface.co/papers/2607.16401 Published on Jul 17
·
Submitted byhttps://huggingface.co/lifuguan
leolion Jul 21
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Modernvideogenerationmodelsareincreasinglyhailedasemergingworldmodelswithaninternalizedgraspofphysicallaw.Yetexistingbenchmarkslargelyevaluatephysicalplausibilityonlyattheoutputlevel,withoutverifyingwhetherthemodelarrivestherethroughafaithful,law-groundedreasoningprocess.WeintroduceApple-PI,thefirstbenchmarkthatanchorsvideo-modelevaluationexplicitlyinphysicallaws.Apple-PIcomprisesthreecomponents.1)Orchard:adatasetof400videoscoveringtencanonicaltasksinclassicalmechanics.Itseparatessingle-lawtasksforconfounder-freediagnosisfrommulti-lawtasksforprobinggeneralization.2)BenchmarkProtocol:athree-stageprotocolbasedonscientificreasoning,includingPerception,Formulation,andDeduction.Ituseschain-of-framespromptingoninfographic-annotatedfirstframes,treatingthegeneratedvideoasthemodel’svisiblereasoningtrace.3)EvaluationSuite:ahybridevaluationsuitethatcombinesMLLM-basedsubjectivescoringwithphysics-law-groundedobjectivemeasures.Thisenablesstage-resolveddiagnosisofnotonlywhetheramodelfails,butwhereitfails.Benchmarking11modelsshowsthatcurrentvideomodelsremainfarfromreliablelaw-groundedworldsimulators,withthebestvideomodelscoringonly0.473.Ourstage-,pillar-,andsource-resolvedanalysesfurtherexposeaPerception-to-Formulation-to-Deductionbottleneck,weakmulti-lawstatetransfer,andapersistentSim-to-Realgap.ThesefindingspositionApple-PIasadiagnosticfoundationforguidingfuturevideomodelstowardworldmodelswithlaw-groundedphysicalintelligence.
View arXiv pageView PDFProject pageGitHub10Add to collection
Get this paper in your agent:
hf papers read 2607\.16401
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.16401 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.16401 in a dataset README.md to link it from this page.
Spaces citing this paper1
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Physics-IQ Verified
This paper presents a systematic audit of the Physics-IQ benchmark for evaluating physical understanding in video generative models, proposing improvements to prompts and scoring to enhance reliability.
PhysBrain 1.0 Technical Report
PhysBrain 1.0 is a technical report presenting a method that uses human egocentric video to generate physical commonsense supervision for vision-language-action models, achieving state-of-the-art results on embodied control benchmarks including ERQA, PhysBench, SimplerEnv-WidowX, LIBERO, and RoboCasa.
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Introduces Video-MME-Logical, a controlled benchmark for evaluating video temporal-logical reasoning in multimodal large language models, revealing a substantial human-model gap.
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
Introduces Pegasus, a low-resource framework that translates human demonstration videos into robot-executable data using graph-based task representation, hierarchical affordance latent space, and closed-loop physics verification, aiming to turn hardware data collection into scalable knowledge transfer.
@rohanpaul_ai: Most video models look better than they understand and Video quality is only the easiest thing to notice. LongCat just …
LongCat released WBench, a benchmark for video world models that tests control, memory, instruction-following, and physical plausibility across 289 cases and 20 models, finding that no model excels in all dimensions, highlighting the gap between video quality and true world simulation.