HOST: So why should someone starting a research project care about this? EXPERT: A search tool might someday help them spot an older idea they could build on. This paper tests how well tools find papers that researchers themselves say helped or could have helped their projects. HOST: Isn't that kind of what searching for related papers already does? EXPERT: Not quite. Imagine studying how to correct a machine while it works. A paper about recovering from mistakes might offer a useful idea, even if its title sounds less related than other search hits. That's a hypothetical example, not a result from this study. HOST: So how did the authors decide which hits actually counted? EXPERT: They asked the people who led completed computer science projects to check their early research questions and judge earlier papers. Those authors also explained their choices. HOST: So what happened when the search systems actually tried the questions? EXPERT: On the broad project questions, one retriever, Qwen3 embedding 8B, had a recall at 20 score of 0.37. That score counts the share of author-credited papers in its first 20 results. A GPT-4.1 search agent scored 0.37 in the same setting. HOST: Does that tell us how a researcher would fare using one at work? EXPERT: No, this was a test using a restricted paper collection, not an everyday search deployment. The authors also note that people judge their own projects after completion, when hindsight could affect their choices. The practical lesson is to check whether a promising paper offers an idea you can use, not only whether it shares your keywords.