HOST: If I'm choosing an AI assistant for research work, why should I care about this? Okay, why should I care about this? EXPERT: You might want to see whether it can follow clues through different languages, videos, and documents. This paper offers questions for testing that possibility, not a workplace guarantee. HOST: So, what makes a question different from a normal web search? EXPERT: The short answer may sit several clues away. One example starts with a German question about a video, points to a place on a map, and ends in a bank report. HOST: So, did the authors test the AI models on their own? EXPERT: No, they tested models together with search setups. Those setups are basically the tools and rules the assistant uses to look things up. HOST: What are the strongest tested setup manage? EXPERT: Gemini 3.7 Flash with its built-in search got 31.68% accuracy on these questions. That counts answers judged correct against the author's short reference answers. HOST: Does that tell me how it would do on my everyday tasks? EXPERT: Not directly. The authors deliberately made this a hard stress test, and live search can change. My practical takeaway is to ask which model and search tools produced a score, then check the evidence trail yourself.