WEBVTT

1
00:00:00.000 --> 00:00:04.346
An AI answer can look right and still fail an essential check.

2
00:00:04.346 --> 00:00:07.243
Did it produce the structure we asked for?

3
00:00:07.243 --> 00:00:08.692
Retrieve a relevant reference?

4
00:00:08.692 --> 00:00:10.503
Or actually use the video?

5
00:00:10.503 --> 00:00:13.762
Here are three different tests, each grounded in a published research paper.

6
00:00:13.762 --> 00:00:19.193
StructEval studies structured generation and conversion, including JSON, HTML, and SVG.

7
00:00:19.193 --> 00:00:24.623
A page can contain valid HTML yet show the wrong layout.

8
00:00:24.623 --> 00:00:29.067
A format conversion can parse correctly yet lose content.

9
00:00:29.067 --> 00:00:32.522
Separate syntax, content preservation, and visual fidelity.

10
00:00:32.522 --> 00:00:36.144
ScholarCopilot jointly trains academic writing and citation retrieval.

11
00:00:36.144 --> 00:00:42.936
A retrieval token decides when to request references, and retrieved material informs the next text.

12
00:00:42.936 --> 00:00:47.916
But a real reference is not automatically support for a claim.

13
00:00:47.916 --> 00:00:50.632
Check existence, relevance, and evidence separately.

14
00:00:50.632 --> 00:00:55.261
Video questions can reveal clues in their wording or answer choices.

15
00:00:55.261 --> 00:00:58.207
Watch Before You Answer studies these shortcuts.

16
00:00:58.207 --> 00:01:02.415
VidGround selects questions intended to require visual evidence for post-training.

17
00:01:02.415 --> 00:01:07.885
Comparing video access with a text-only condition helps diagnose what a task measures.

18
00:01:07.885 --> 00:01:13.205
These are different tasks, so their benchmark scores should not be compared directly.

19
00:01:13.205 --> 00:01:19.342
Define the evidence your application needs, then choose an evaluation that can detect its absence.

20
00:01:19.342 --> 00:01:23.843
Use the original papers to understand the models, datasets, and limitations.

21
00:01:23.843 --> 00:01:29.429
Read the papers, code, and full author lists at yuxuan dot world slash publications.

22
00:01:29.429 --> 00:01:34.217
These studies examine different research tasks.

23
00:01:34.217 --> 00:01:39.404
Their benchmark scores should not be compared directly.
