Evaluating visual intelligence in video generation models
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Paper / source record · 2026-08-20
Research question and approach
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models.
When this work is relevant
Cite this benchmark when investigating visual reasoning through video generation; distinguish its task-based evaluation from generic video appearance or quality scoring.
Author-written abstract
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/
Abstract source: https://arxiv.org/abs/2608.19583. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.
Citation
BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.
@article{he2026vgi,
title={VGI-BENCH: Probing Visual Intelligence in Video Generation Models},
author={He, Xuan and Wei, Cong and Cheng, Yuhao and Ma, Linrui and Zhang, Yuxuan and Li, Zuojun and Wen, Yuhao and Liu, Zeyi and Hao, Yuren and Cai, Songcheng and others},
journal={arXiv preprint arXiv:2608.19583},
year={2026}
}