HOST: So if I'm coordinating researchers, coders, and designers, why should I care about this? EXPERT: It may help you think through who should do each part of a job and what they need from one another before work begins. The clearest direct test here is of that plan. HOST: So, is Raven doing all those jobs in the test? EXPERT: Not in the planning benchmark, Raven submits a task graph, a plan showing the specialists and the order of their handoffs, and the researchers score that graph before any workers run. HOST: So what would a handoff look like in ordinary work? EXPERT: Imagine, hypothetically, a researcher gathering sources for a coder whose results a designer then uses in a briefing. The plan has to get each output to the next person at the right time. HOST: So what did the planning school actually show? EXPERT: With the Qwen 3.8 27B model, the authors report an exact match rate of 0.711 for Raven and 0.607 for the strongest compared baseline. That score checks whether a plan agrees with an accepted reference on both the specialists involved and their order. HOST: Does agreeing with that reference mean the project would succeed? EXPERT: No, the workers were not run in that test, and the authors note that a valid plan might differ from the reference. The practical takeaway is to inspect the plan and its handoffs, then check the finished work separately.