Benchmarking Video Models on Physical Mechanics
Listen to the summary
Uses a voice available on your device
Audio options
On this page 4 sections
Related concepts 1 concepts
Key Takeaways
- The Orchard dataset provides 400 videos across ten canonical tasks in classical mechanics for standardized model evaluation.
- Unified understanding-generation models like GPT Image 2 and Nano Banana 2 currently lead with scores of 0.704 and 0.699.
- Dedicated video generation models show limited performance, with the best model achieving an average score of only 0.473.
- The evaluation reveals that current models remain far from reliable simulators that follow physical laws.
Summary & Methodology Analysis
The researchers developed a benchmark protocol that organizes model evaluation into three stages: Perception, Formulation, and Deduction. This approach uses the Orchard dataset, which consists of 400 videos covering ten tasks in classical mechanics. To evaluate performance, the team tested 11 models, including six unified understanding-generation models like GPT Image 2 and Nano Banana 2, against standard video generation architectures. The workflow prompts models with an infographic-annotated first frame to generate a reasoning trace, which is then measured using a hybrid suite of subjective scores and objective metrics like mask intersection over union and velocity estimation.
Interactive System Flowchart
Cross-Examination & FAQs
A deeper dive clarifying mechanics, constraints, and baseline evaluations.
Q1. What is the primary goal of the Apple-π benchmark?
It serves as a diagnostic tool to evaluate how well video models can perform reasoning based on classical mechanics.
Q2. What does the Orchard dataset contain?
It contains 400 videos focused on ten canonical tasks in classical mechanics.
Q3. Are current video models capable of reliable physical reasoning?
No, the research shows that current models remain far from reliable law-grounded simulators.
Q4. Which models achieved the highest scores in the benchmark?
GPT Image 2 and Nano Banana 2 achieved the highest overall scores of 0.704 and 0.699 respectively.
Q5. What was the performance of the best specialized video generation model?
The best-performing video generation model achieved an average score of 0.473.
Q6. What physical phenomena does the benchmark exclude?
It does not evaluate fluids, thermodynamics, electromagnetism, deformable bodies, fracture, granular media, biological motion, or quantum phenomena.
Q7. How does the camera setup impact the benchmark's scope?
The benchmark uses a single fixed camera for all cases, meaning it does not test multi-view consistency or novel-view generation.
Q8. What are the limitations of the objective metrics used?
Metrics can be noisy because mask intersection over union depends on segmentation quality, pixel metrics may penalize harmless appearance differences, and velocity estimation can be inaccurate for real-world data.
Q9. Which specific unified understanding-generation models were tested?
The tested models were BAGEL, OmniGen2, SenseNova-U1-8B-MoT, SenseNova-U1-8B-MoT-Think, GPT Image 2, and Nano Banana 2.