HOST: Why should someone checking video descriptions at work care about this? EXPERT: It could help them spot a description that sounds right but loses track of a person or attaches a sound to the wrong moment. That's a possible use, not something the authors tested at work. HOST: How can a test see that when an ordinary caption is just a paragraph? EXPERT: OmniCapBench asks models for linked records instead: who or what keeps showing up, what each camera shot includes, and when different sounds happen. HOST: So, say a dog appears and then barks after a cut. What happens to those records? EXPERT: Rules check identifiers, time spans, and links. The language model has a restricted matching job for certain local identity descriptions and brief visual actions, while time shots and events are matched by specified measures. Only then does a local judge compare corresponding descriptions. HOST: So, what did the authors actually find in their study? EXPERT: On their annotated videos, Gemini 2.5 Pro received 37.81% on a measure of keeping identities consistent across cuts. The authors used separate measures to show other errors, including sound-to-shot links. HOST: Does that tell me how well a captioning tool will work on any video? EXPERT: No, the authors tested videos under five minutes, their annotations can't cover every true detail, and some drafts actually came from the families of models they later tested. The practical lesson is to inspect timing, identity, and links separately when you review a caption, not to treat this benchmark as a deployment guarantee.