Measuring distinct capabilities
IoU is 0.6: recognizing the class alone would not measure this localization error.
Links containing ?t= open the video at a specific second.
The ideas to retain
1. What it means to evaluate grounding
In multimodal systems, grounding is the degree to which the model's answer is supported by the actual content of the image or audio rather than by statistical expectations derived from the…
2. The problem of benchmark contamination
The second systematic problem is contamination. Foundation models are pretrained on massive amounts of internet data, and there is no guarantee that image-description pairs or evaluation…
3. Language bias: answering from probability, not evidence
A language prior is the tendency of a model to generate answers that are statistically likely given the question text regardless of the image content. It is one of the subtlest forms of…
Jump directly to a section
Read the reviewed transcript
This video has no narration. This transcript reproduces its on-screen text; it does not invent a spoken track.
Naming and locating are different successes
The target box measures four by two units. The prediction is shifted by one unit.
The intersection has area six and the union area ten.
IoU is 0.6: recognizing the class alone would not measure this localization error.
The same words can change who is where
Two images are paired with two captions using the same object names.
Reversing above and below changes the correct matching.
Compositional evaluation requires the crossed relations, not only object recognition.
Checking a value requires checking the cell
The answer claims that order B totals twenty.
Twenty appears in the table but belongs to order A; B totals thirty.
The test must validate row, column and value to detect a plausible but false answer.
Order can be the whole question
Sequence A turns on a light and then opens a door.
Sequence B contains the same events in the opposite order.
A temporal test must distinguish them even when their object and action inventories match.
A test should require visual evidence
The question asks how many dots are in the image; its text stays identical.
Only the image changes from two dots to three, while other conditions are held fixed.
The answer must change; always answering two can reveal reliance on text rather than the image.
Answering more cases can admit more errors
Five cases have confidence scores, but some answers are incorrect.
A high threshold accepts two with no errors here; lowering it accepts four with one error.
Report coverage and error together: two successes do not prove reliability beyond this sample.


