HOST: If I work with specialist research software, why should I care about this? EXPERT: It could help you ask whether an assistant produced the right saved result, not just whether it looked busy in the software. HOST: So what made that distinction so clear? EXPERT: In an astronomy task, the authors describe a GPT-5.6 Terra run that saved a wrong measurement and later checked the saved value against itself. The check confirmed the entry, not the science. HOST: So how did the researchers test assistance more broadly? EXPERT: They gave them tasks in scientific programs, recorded what they did, and then checked what was left behind. They task-specific evaluator, that's a check written for that particular task, can inspect a file or a result inside the program. HOST: So does the headline score mean most tasks were completely finished? EXPERT: No, Claude Fable 5.1 had the highest overall mean score in the authors' comparison, but that average includes partial credit. It's not a count of fully completed tasks. HOST: Could giving an assistant more screenshots fix the problem? EXPERT: On the QPath pathology tasks, the authors tried different history windows, basically how many recent screenshots the assistant could consult. The scores did change, but after correcting for multiple comparisons, they didn't see a significant difference from the default. HOST: So with that in mind, what would you say are the key takeaways someone could apply to their own work? EXPERT: As a practical possibility, define a result you need and an independent way to check it. The authors test concerned benchmark tasks, and some detailed analyses cover only part of the evaluated work. They do not show a ready-to-deploy assistant.