HOST: If I'm choosing an AI tool for office work, why should I care about this paper? EXPERT: It offers a possible way to test the thing you need back, a finished file or other deliverable, rather than a convincing answer about the job. HOST: So what does the test ask an agent to do? EXPERT: The authors give it a professional task, input files and access to relevant software. It works through the steps, and the test checks what it leaves behind. HOST: So what would that look like in everyday terms? EXPERT: Imagine, just as an illustration, asking for a completed financial workbook. The important question is whether the saved workbook has the required contents, not whether the agent says it made one. HOST: And what happened when the authors ran agents through the test? EXPERT: Codex paired with GPT-5.5 and desktop tools fully passed 38.1% of the near-term tasks, but 0.0% of the last exam tasks. A full pass means the task received full credit. An agent could still get partial credit without passing. HOST: Does that tell me the same tool would manage my team's work? EXPERT: No, the authors flagged the risk of agents encountering public tasks in training or being tuned to them, and only some configurations had repeated runs. The practical lesson is to define your own finished output and checks before trusting a promising demonstration.