Video post-training that depends on visual evidence

Watch Before You Answer: Learning from Visually Grounded Post-Training

Zhang, Yuxuan, Hwang, EunJeong, Zhang, Huaisong, Du, Penghui, Jia, Yiming, Jiang, Dongfu, He, Xuan, Zhang, Shenhui, Nie, Ping, West, Peter, Allen, Kelsey R.

Paper / source record · 2026-04-06

Project arXiv Code HuggingFace X Video overview Probe log mini-lab

Research question and approach

Video question answering should depend on what the frames show. Language-only shortcuts can make benchmark accuracy look stronger than visual understanding. This study investigates visually grounded post-training and the role of data selection.

When this work is relevant

Cite this study when discussing linguistic shortcuts in video question answering or selecting post-training data for visual dependence. Check the paper's experimental scope before generalizing.

Author-written abstract

It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com.

Abstract source: https://arxiv.org/abs/2604.05117. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.

Citation

BibTeX · CITATION.cff

BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.

@article{zhang2026watch,
  title={Watch before you answer: Learning from visually grounded post-training},
  author={Zhang, Yuxuan and Hwang, EunJeong and Zhang, Huaisong and Du, Penghui and Jia, Yiming and Jiang, Dongfu and He, Xuan and Zhang, Shenhui and Nie, Ping and West, Peter and others},
  journal={arXiv preprint arXiv:2604.05117},
  year={2026}
}