Reasoning-based evaluation of generated videos

VideoScore2: Think before You Score in Generative Video Evaluation

He, Xuan, Jiang, Dongfu, Nie, Ping, Liu, Minghao, Jiang, Zhengxuan, Su, Mingyi, Ma, Wentao, Lin, Junru, Ye, Chun, Lu, Yi, Wu, Keming, Schneider, Benjamin, Do, Quy Duc, Li, Zhuofeng, Jia, Yiming, Zhang, Yuxuan, Cheng, Guo, Wang, Haozhe, Zhou, Wangchunshu, Lin, Qunshu, Zhang, Yuanxing, Zhang, Ge, Huang, Wenhao, Chen, Wenhu

Paper / source record · 2025-09-26

Project Paper Code Dataset Model HF Space Twitter

Research question and approach

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales.

When this work is relevant

Use this work when comparing evaluation of generated-video quality and reasoning-based scoring. Compare dimensions and protocols rather than combining scores from different benchmarks.

Author-written abstract

Recent advances in text-to-video generation have produced increasingly realistic and diverse content, yet evaluating such videos remains a fundamental challenge due to their multi-faceted nature encompassing visual quality, semantic alignment, and physical consistency. Existing evaluators and reward models are limited to single opaque scores, lack interpretability, or provide only coarse analysis, making them insufficient for capturing the comprehensive nature of video quality assessment. We present VideoScore2, a multi-dimensional, interpretable, and human-aligned framework that explicitly evaluates visual quality, text-to-video alignment, and physical/common-sense consistency while producing detailed chain-of-thought rationales. Our model is trained on a large-scale dataset VideoFeedback2 containing 27,168 human-annotated videos with both scores and reasoning traces across three dimensions, using a two-stage pipeline of supervised fine-tuning followed by reinforcement learning with Group Relative Policy Optimization (GRPO) to enhance analytical robustness. Extensive experiments demonstrate that VideoScore2 achieves superior performance with 44.35 (+5.94) accuracy on our in-domain benchmark VideoScore-Bench-v2 and 50.37 (+4.32) average performance across four out-of-domain benchmarks (VideoGenReward-Bench, VideoPhy2, etc), while providing interpretable assessments that bridge the gap between evaluation and controllable generation through effective reward modeling for Best-of-N sampling. Project Page: https://tiger-ai-lab.github.io/VideoScore2/

Abstract source: https://arxiv.org/abs/2509.22799. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.

Citation

BibTeX · CITATION.cff

BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.

@article{he2025videoscore2,
  title={Videoscore2: Think before you score in generative video evaluation},
  author={He, Xuan and Jiang, Dongfu and Nie, Ping and Liu, Minghao and Jiang, Zhengxuan and Su, Mingyi and Ma, Wentao and Lin, Junru and Ye, Chun and Lu, Yi and others},
  journal={arXiv preprint arXiv:2509.22799},
  year={2025}
}